Token Scarcity
Token Scarcity
The rationing regime that follows the token-subsidy era. As usage shifts from assisted (a human in the loop, bounded turns) to agentic (continuous loops, fan-out sub-agents), AI demand outpaces a physically constrained token supply — GPUs, power, and data-center capacity can't scale as fast as agent-driven consumption. The subsidy era ends, prices rise, and every AI company is "now in the token-efficiency business."
The thesis
- Demand outruns supply. Agentic use multiplies token consumption per task; supply is gated by physical infrastructure. The gap closes by rationing, not by more subsidy.
- The subsidy era ends. Labs have been selling tokens below cost to win adoption. As demand hardens, that stops — prices rise toward true cost.
- Firms ration. The enterprise response is a managed-input posture: mixed-basket model routing (cheapest sufficient model per task), token budgets, and per-seat caps.
Distinct from Token Maxing
Token Maxing is the prior behavior — the abundance/subsidy-era pathology of burning tokens because they were cheap and consumption was the scoreboard. Token scarcity is the regime that follows: once tokens are priced at cost and supply is constrained, the maximalist behavior is no longer affordable and rationing discipline becomes mandatory. Maxing is the disease of cheap tokens; scarcity is the economics that forces the cure.
2026-06-20 — "Token Reckoning" Economist coverage
Companies Are Scrambling to Curtail Soaring AI Costs (Economist) is the Business-section confirmation that the regime is now operating in the wild:
- Ramp spend data: AI spending up 13× year-on-year. Token-heavy applications (reasoning models, agents that build agents) are the growth driver.
- Uber spent its annual AI budget in four months. One unnamed firm spent $500m on tokens in a single month.
- The distribution is bimodal — top-1% spenders at ~$7,450/month per employee, median client at $11. The Advantage Gap expressed financially.
- Three corporate response patterns now visible: (1) Meta + Amazon killed their token-usage leaderboards; (2) Routing down-tier — Sonnet ~1/20 of Opus, Kimi ~1/20 of Sonnet (three orders of magnitude across the routing decision); (3) Per-seat / per-task caps (Uber's $1,500/month per coding tool is the canonical case).
- Outcome-based pricing emerging: Intercom charges customers only for queries actually resolved by its IT-support agent — the SaaS pricing model that survives the agentic era.
- The lab-subsidy era ends with the IPOs. Sam Altman called mounting customer costs "a huge issue." OpenAI's strategy for winning customers from Anthropic reportedly involves drastic price cuts — but once both labs IPO later in 2026, prices have to rise toward true cost.
- The geography axis: AI bills are "low compared with hiring a developer in San Francisco, but high compared with employing one in Delhi" — token-cost is now a variable in the on/offshore equation. See Indian IT and AI.
2026-06-27 — Rationing is now operational at the labs (Economist)
Americas Data-Centre Backlash Puts the AI Boom at Risk (Economist) gives the top-of-funnel evidence that inference capacity is genuinely constrained:
- "Anthropic has throttled model usage."
- "OpenAI has scrapped its compute-intensive video tool."
- "Microsoft has repriced its coding assistant so steeply that some programmers are returning to the lost art of writing software themselves."
The physical numbers behind it: ~12 GW of US AI compute today → ~10 GW dedicated to inference across the majors → demand grows faster than the ~30 GW new-build queue through 2028, with training alone potentially absorbing 5–16 GW per frontier model.
2026-06-27 — Total-cost-not-per-token is the buried lede on Chinese "cheap AI"
China Is Having Another AI Moment (Economist) adds the second-order refinement that stops the "we'll just switch to DeepSeek/Zhipu" reflex:
- DeepSeek v4: $0.87 per 1M output tokens; Anthropic Fable 5: $50 per 1M output tokens. ~57× per-token gap.
- Du Zheng (Georgia Tech) et al., June 2026: DeepSeek used 23× more tokens than an OpenAI rival to achieve basically the same result on the same tasks.
- Correct comparison metric = total cost of tokens used, not price per token.
- On a software-engineering benchmark, GLM 5.2 ended up costing more than systems from Anthropic and OpenAI.
The Token Scarcity regime therefore extends: even the "escape hatch to open-source Chinese models" has token-efficiency limits. Whichever model you route to, tokens are the variable cost that matters, not seat licenses or per-token headline prices.
2026-07-06 — Whittemore's June retrospective: the pivot month, periodized
The Big Ways AI Just Changed (AI Daily Brief) periodizes the regime change: May = the shift became visible; June = it became operational. New datapoints beyond the Economist coverage above:
- Walmart moved internal AI tools from unlimited usage to token budgets at the start of June — the first named non-tech mega-enterprise on the rationing list (joins Uber's $1,500/month cap, restated here).
- The measurement layer adapted: Artificial Analysis reweighted its intelligence index toward agentic usage — the benchmark infrastructure now assumes token-hungry agentic workloads as the norm.
- Scarcity is being amplified from the physical layer: memory-company stock outperformance as the memory shortage came into focus; compute becoming a tradable market of its own (SpaceX's expanding neocloud deals with Anthropic/Google/Reflection AI; Meta reportedly following). See AI Capex Supercycle.
- Whittemore's caveat worth keeping: only a vanguard sliver of companies is anywhere near token-efficiency problems — most consume "a vanishingly small portion of the intelligence they will ultimately consume." For the average adopter, the live problem is Botsitting, not rationing.
2026-08-02 — the reframe underneath the regime: AI is a labor line, not a software line
6 Questions Shaping Enterprise AI (AI Daily Brief) supplies the conceptual move that makes the whole regime legible to a CFO, reported from inside KPMG's enterprise symposium:
- The lab revenue chart and the enterprise cost chart are the same chart. "The enterprise experiences the inverse side of that revenue chart as a cost chart." Anthropic's run-rate milestones are, from the buyer's seat, a spend curve.
- The category error, named. "AI in the enterprise is not just another category of software spend, but represented something fundamentally different, something more akin perhaps to labor." Software spend is fixed and negotiated annually; labor is variable, consumed continuously, and managed. Most enterprise budgeting machinery is built for the former.
- Seats → tokens is what deflated the bubble narrative. Whittemore credits this reframe directly: "the recognition that we were not talking about seats, but instead talking about tokens did a whole lot to collapse the AI bubble narratives on Wall Street from Q4 of last year."
- Uber's budget burn, independently sourced. The same datapoint this page already carries from Economist reporting arrives here via a completely different path — Whittemore naming Uber as "the most notable" case of enterprises torching annual budgets in months. Second-source corroboration of the pattern.
- His defence of the budget-blowers is worth keeping: "How are we going to expect organizations to effectively budget for the agentic token era of AI when no one knew that that was right around the corner when those budgets were being made?" The failure was forecasting, not discipline.
The new operational requirement this surfaces: observability. Beyond the caps and routing this page already documents, the enterprises furthest along discovered that rationing presupposes measurement — "systems for monitoring and measuring AI usage." Without cost-to-output visibility per group, function and project, allocation decisions have no basis. Whittemore's colour: "You have not seen the word token used more at an event since the height of the crypto era." And no one at the event treated Model Routing products as a silver bullet.
Cross-references
- Token Maxing — the maximalist behavior this regime supersedes
- Managing Enterprise IT Development in the Era of Token Scarcity — the enterprise-IT operating playbook for this regime
- Agentic Loop — the usage shift (assisted → agentic) that drives demand past supply
- Indian IT and AI — the on/offshore axis Token Scarcity compresses
- SaaSpocalypse — outcome-based pricing as the survivable SaaS landing spot
Sources
- The New Dumbest Chart in AI (AI Daily Brief) — "every AI company is now in the token-efficiency business"; the assisted → agentic demand-vs-supply framing
- Companies Are Scrambling to Curtail Soaring AI Costs (Economist) — 2026-06-20; the operating data (Ramp 13×, Uber's $1,500/month per-coding-tool cap, $7,450 top-1%, Sonnet 1/20 routing factor, Intercom outcome-pricing)
- Americas Data-Centre Backlash Puts the AI Boom at Risk (Economist) — 2026-06-27; the top-of-funnel evidence (Anthropic throttling, OpenAI killing video tool, Microsoft repricing Copilot)
- China Is Having Another AI Moment (Economist) — 2026-06-27; the 23× token overuse rebuttal to "we'll just switch to Chinese models"
- 6 Questions Shaping Enterprise AI (AI Daily Brief) — 2026-08-02; the labor-not-software reframe, seats→tokens as the bubble-narrative solvent, and observability as the unmet prerequisite for rationing