# Tokens, Tallied: The Economics of Inference for Agentic Workloads

> Token economics for agentic workloads pit a 280× collapse in LLM inference cost against quadrillion-token volumes; here are the prices, the spending forecasts and the FinOps levers, dated to September 2026.

- Canonical: https://aiagentinfra.com/articles/token-economics-agentic-inference
- Author: Ryan Elliott Dennis
- Category: Models & Reasoning
- Kind: Reference article
- Last verified: 2026-09-04
- Keywords: token economics, LLM inference cost, cost per million tokens, prompt caching, inference spending, AI-optimized IaaS, agent FinOps, GPT-5.6 pricing, Claude Fable 5.1 pricing

> "Now, compute is revenue." — Jensen Huang, founder and CEO of Nvidia (Nvidia Q2 FY2027 earnings release, Aug. 26, 2026)

Nvidia reported revenue of $96.2 billion for its second fiscal quarter of 2027, up 106% from a year earlier, with data-center revenue of $89.0 billion and third-quarter guidance of $108.0 billion, in an earnings release dated Aug. 26, 2026 in which Nvidia CEO Jensen Huang declared that compute had become revenue. The demand behind that revenue is denominated in tokens. Token economics for agentic workloads now rest on two opposing curves: the price of a million tokens has fallen by orders of magnitude since 2022, while the number of tokens an agent consumes per task has risen with every reasoning budget, context window and tool call added to the stack. Which curve wins decides whether LLM inference cost is a cost center or a margin engine. This article tallies both, with the price list as of Sept. 4, 2026, the spending forecasts and the FinOps levers that move the bill.

## Price Deflation: The Token Economics of a 280× Fall in LLM Inference Cost

Stanford's AI Index 2025, published in April 2025, tracked the cost of querying a model at GPT-3.5's level of 64.8% on MMLU from $20 per million tokens in November 2022 to $0.07 in October 2024, a 280× decline in about 18 months, and it put the annual rate of decline between 9× and 900× depending on the task. Andreessen Horowitz had named the phenomenon "LLMflation" in November 2024, estimating that the cost of constant capability was falling about 10× per year over the prior three years, with GPT-3-quality output dropping from $60 to $0.06 per million tokens between late 2021 and late 2024. The curve has yet to flatten. Crypto Briefing reported in August 2026 that its frontier-model token price index stood at $1.16 to $1.18 per million tokens, down 43% from $2.04 at the end of May 2026 and about 12% of March 2023 levels, a ten-week move that annualizes to roughly 18× per year by this journal's arithmetic, above the a16z rate. VoxBooster, an aggregator whose figures this journal treats as secondary, cites Epoch AI for a median decline of 50× per year and 200× per year since January 2024. Deflation of that speed changes procurement behavior. A price negotiated in the spring is a bad price by the fall.

## Volume Inflation: A Trillion Tokens per Customer, Quadrillions per Platform

Microsoft said on its April 29, 2026 earnings call for the third quarter of fiscal 2026 that more than 300 Azure AI Foundry customers were on track to process over a trillion tokens each this year, that token processing rose about 30% quarter over quarter, and that its AI business had reached a $37 billion annualized run rate, up 123%. Databricks, according to PointFive's coverage of the 2026 Data + AI Summit, reported more than 100,000 agents built on its platform and more than a quadrillion tokens a year flowing through it; the figure is the vendor's, relayed by a secondary source. Google processed roughly 3.2 quadrillion tokens a month by mid-2026, about seven times the prior year, according to the VoxBooster aggregation, which this journal flags as secondary pending a first-party Google disclosure.

Huang's release put the supply-side reading in one line: "Its tokens are productive and profitable." Agentic workloads drive the volume because a single agent task compounds tokens. System prompts and tool schemas are re-sent each turn, reasoning budgets add thinking tokens in proportion to difficulty, and multi-step tasks multiply turns. Price per token falls; tokens per task rise; the product of the two is the number a CFO sees.

## Inference Spending: Gartner's $42 Billion AI-Optimized IaaS Market and the 55% Inference Share

Gartner said in an Aug. 10, 2026 press release that worldwide spending on AI-optimized infrastructure as a service will reach $42 billion in 2026, up 96%, and $66 billion in 2027, up 56.5%, after $21.5 billion in 2025, itself up 180%. Inference accounts for 55% of the 2026 figure, $23.3 billion against $19 billion for training, and Gartner expects the inference share to reach 59% in 2027. This is the year inference overtook training as the larger line item. That crossover matters for agent infrastructure because inference is the line that scales with usage: training spend is episodic and concentrated in a few labs, while inference spend is continuous and distributed across every application that calls a model. VoxBooster's aggregation places OpenAI's 2025 inference spend near $8.4 billion and Anthropic's near $2.7 billion; both figures are estimates from a secondary source and are recorded here as such.

## The Price List, Recorded Twice: GPT-5.6 Sol, Terra, Luna, Sonnet 5, Opus 5 and Fable 5.1

| Model | Input / output per MTok | Cached input | Source and date |
|---|---|---|---|
| GPT-5.6 Sol | $5.00 / $30.00 | $0.50 | OpenAI pricing page, Sept. 4, 2026; TechCrunch, July 9, 2026 |
| GPT-5.6 Terra (launch) | $2.50 / $15.00 | — | TechCrunch, July 9, 2026 |
| GPT-5.6 Terra (current) | $2.00 / $12.00 | $0.20 | OpenAI pricing page, Sept. 4, 2026 |
| GPT-5.6 Luna (launch) | $1.00 / $6.00 | — | TechCrunch, July 9, 2026 |
| GPT-5.6 Luna (current) | $0.20 / $1.20 | $0.02 | OpenAI pricing page, Sept. 4, 2026 |
| Claude Fable 5.1 / Mythos 5.1 | $10.00 / $50.00 | $0.25 cache read | Anthropic, Sept. 1, 2026 |
| Claude Opus 5 | $5.00 / $25.00 | — | Anthropic pricing page, Sept. 4, 2026 |
| Claude Sonnet 5 | $2.00 / $10.00 | — | Anthropic pricing page, Sept. 4, 2026 |
| Claude Haiku 4.5 | $1.00 / $5.00 | — | Anthropic pricing page, Sept. 4, 2026 |

The two rows each for Terra and Luna record a conflict. TechCrunch reported at the July 9, 2026 launch that Terra cost $2.50 and $15 and Luna $1 and $6 per million input and output tokens; OpenAI's pricing page as viewed Sept. 4, 2026 lists Terra at $2.00 and $12.00 and Luna at $0.20 and $1.20, with cached input at a tenth of list and a 50% batch discount. OpenAI's own channels have yet to explain the difference in the research base for this article, so both figures stand, and a Luna price that fell 80% within two months would be the sharpest in-family cut on record. Anthropic held Fable 5.1 at $10 and $50 at the Sept. 1 release while cutting cache reads 75% to $0.25, and the company claims workload costs around 25% lower on typical use and up to around 45% lower on highly agentic use relative to Fable 5; those percentages are the vendor's. Sonnet 5 lists at $2 and $10, Opus 5 at $5 and $25, and the retired Opus 4 and 4.1 had listed at $15 and $75, which means Anthropic's mid tier now costs roughly a seventh of the list price of its retired top tier.

## Agent FinOps: Prompt Caching, Batching, Routing and Active-CPU Billing

Four levers move an agent's inference bill, and each now has a published price. Prompt caching is the largest: Anthropic's $0.25 cache-read price is 2.5% of Fable 5.1's $10 input rate, and OpenAI's cached input for Sol is $0.50, a tenth of list, which rewards agents that keep long system prompts, tool schemas and retrieved context stable across turns. Batching is the second: OpenAI's batch tier halves both input and output prices for workloads that tolerate latency, and agent evaluation runs, backfills and nightly report generation qualify. Routing is the third: Sonnet 5 and Luna price at a fifth of their flagship siblings or below on input, and an orchestrator that classifies task difficulty before choosing a model captures the difference on every easy turn. Active-CPU billing is the fourth, and it lives in the runtime layer: MarkTechPost's Aug. 27, 2026 sandbox benchmark priced Cloudflare at $0.072 and Vercel at $0.128 per vCPU-hour billed on active CPU alone, against Northflank at $0.0167, Daytona and E2B at $0.0504 and Modal at $0.0710 billed on wall-clock time, with the cost of 1,000 executions ranging from $1.67 to $52.80 across providers. An agent that spends most of its life waiting on a model pays for the wait under wall-clock billing and pays for the compute alone under active-CPU billing.

Visibility lags the levers. KPMG's Q2 2026 AI Pulse, which surveyed 204 US C-suite leaders at companies with revenue above $1 billion between April 28 and May 25, 2026, found that 26% have full real-time visibility into AI operating costs, while the average planned AI investment over the next 12 months is $202 million. A company spending $202 million on a bill it sees quarterly is buying tokens on faith.

## What to Watch

Three prices and one disclosure will define token economics through early 2027. Gartner's 2027 forecast update will show whether inference reaches its projected 59% share of AI-optimized IaaS or overshoots as agent volumes compound. OpenAI's explanation of the Terra and Luna price cuts, or a correction to the launch-day reporting, will settle the largest pricing conflict in this ledger. Anthropic's 45% claim for highly agentic workloads will meet its first independent test once observability vendors publish cache-hit statistics from production traces. The disclosure to watch is Google's: a first-party token count would either confirm the 3.2 quadrillion-per-month aggregate figure or retire it, and it would give the industry its first audited denominator for price per token at the scale where agents run.

## By the numbers

- Nvidia revenue, Q2 FY2027: $96.2 billion — +106% year over year; data center $89.0 billion; Q3 guidance $108.0 billion ±2% [1]
- Cost to query a GPT-3.5-level model: $20 to $0.07 per MTok — Nov. 2022 to Oct. 2024, a 280× decline (Stanford AI Index 2025) [2]
- Inference share of AI-optimized IaaS spending, 2026: 55% — $23.3 billion of $42 billion; 59% in 2027 (Gartner, Aug. 10, 2026) [5]
- Frontier-model token price index, Aug. 2026: $1.16 to $1.18 per MTok — Down 43% from $2.04 at end-May 2026 (Crypto Briefing) [4]
- C-suite leaders with full real-time visibility into AI operating costs: 26% — KPMG Q2 2026 AI Pulse, 204 US leaders at $1B+ companies [14]

## Sources

1. Nvidia, "NVIDIA Announces Financial Results for Second Quarter Fiscal 2027," Nvidia newsroom, Aug. 26, 2026. https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-second-quarter-fiscal-2027
2. Stanford HAI, "AI Index 2025: State of AI in 10 Charts," Stanford Institute for Human-Centered AI, April 2025. https://hai.stanford.edu/news/ai-index-2025-state-of-ai-in-10-charts
3. Andreessen Horowitz, "Welcome to LLMflation: LLM Inference Cost Is Going Down Fast," a16z, November 2024. https://a16z.com/llmflation-llm-inference-cost/
4. Crypto Briefing, "AI Token Prices Hit New Record Lows as Inference Costs Plunge 43% in Ten Weeks," Crypto Briefing, August 2026. https://cryptobriefing.com/ai-token-prices-record-lows/
5. Gartner, "Gartner Forecasts Worldwide AI-Optimized IaaS Spending to Grow 96% in 2026," Gartner press release, Aug. 10, 2026. https://www.gartner.com/en/newsroom/press-releases/2026-08-10-gartner-forecasts-worldwide-artificial-intelligence-optimized-iaas-spending-to-grow-96-percent-in-2026
6. Microsoft, "Earnings Release FY26 Q3," Microsoft Investor Relations, April 29, 2026. https://www.microsoft.com/en-us/investor/events/fy-2026/earnings-fy-2026-q3
7. PointFive, "Snowflake and Databricks Summits 2026: What Actually Matters," PointFive, June 2026. https://www.pointfive.co/blog/snowflake-and-databricks-summits-2026-what-actually-matters
8. VoxBooster, "AI Inference Cost Statistics (2026)," VoxBooster (aggregator), 2026. https://voxbooster.com/blog/ai-inference-cost-statistics-2026/
9. TechCrunch, "OpenAI Launches Its New Family of Models With GPT-5.6," TechCrunch, July 9, 2026. https://techcrunch.com/2026/07/09/openai-launches-its-new-family-of-models-with-gpt-5-6/
10. OpenAI, "API Pricing," OpenAI, As viewed Sept. 4, 2026. https://openai.com/api/pricing/
11. Anthropic, "Claude Fable 5.1 and Mythos 5.1," Anthropic, Sept. 1, 2026. https://www.anthropic.com/claude-fable-and-mythos-5-1
12. Anthropic, "Pricing," Claude Developer Platform, As viewed Sept. 4, 2026. https://platform.claude.com/docs/en/about-claude/pricing
13. MarkTechPost, "Best Agent Sandboxes in 2026: Cold Start, Per-Second Pricing, and Network Policy," MarkTechPost, Aug. 27, 2026. https://www.marktechpost.com/2026/08/27/best-agent-sandboxes-2026-cold-start-pricing-network-policy/
14. KPMG, "KPMG Q2 2026 AI Quarterly Pulse Survey," KPMG, June 24, 2026. https://kpmg.com/us/en/media/news/q2-ai-pulse-2026.html
