Stop Counting Tokens: Why Uber's AI Budget Ran Out in Four Months

Enterprise discussions of AI economics have converged on the token as the unit of account. This paper argues that the token is a supplier metric and that the enterprise unit of account should be the cost per successful outcome (CPSO): all cash spent on an AI workflow divided by the outcomes that meet a stated quality bar. We present the CPSO Model, a closed-form monthly model covering logistic adoption, token consumption that grows quadratically with agent steps, volume-dependent prompt caching, tiered model routing with tier-specific success rates, a lognormal distribution of per-user spend with an exact expression for the effect of spending caps, the operating costs that sit outside a token invoice, the total cost of AI error (TCAE), and a value ledger measured against displaced labor or revenue. We derive six propositions, including a break-even success-rate gap for routing work to cheaper models and a break-even productivity gain for internal adoption. We calibrate the model to public reporting on Uber's 2025 to 2026 rollout of agentic coding tools, which exhausted a full-year AI budget in four months. The calibrated model reproduces the reported adoption path, average and heavy-user spend, and the month the budget ran out. It finds that the budget failed on the adoption curve rather than on token prices; that a $1,500 monthly cap binds for about 2% of engineers while removing about 13% of demand; that tokens are roughly a third of the cash bill; that the program covers its cash cost once each engineer converts about 2.2 hours a month of freed time into output; and that a session moved from a frontier to a mid-tier model pays only if it loses less than 1.6 percentage points of success probability. Applied to the illustrative example in a recent industry white paper, a 78 to 85% reduction in token cost corresponds to a 30% reduction in the complete bill. We state falsification conditions for each claim. The archive includes the LaTeX source, vector figures, and a reference implementation of the model (JavaScript), along with scripts that regenerate every table and figure.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23062189
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Stop Counting Tokens: Why Uber's AI Budget Ran Out in Four Months

Singh Anil Prakash
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

Stop Counting Tokens: Why Uber's AI Budget Ran Out in Four Months

Singh Anil Prakash
preprint en

Abstract

Enterprise discussions of AI economics have converged on the token as the unit of account. This paper argues that the token is a supplier metric and that the enterprise unit of account should be the cost per successful outcome (CPSO): all cash spent on an AI workflow divided by the outcomes that meet a stated quality bar. We present the CPSO Model, a closed-form monthly model covering logistic adoption, token consumption that grows quadratically with agent steps, volume-dependent prompt caching, tiered model routing with tier-specific success rates, a lognormal distribution of per-user spend with an exact expression for the effect of spending caps, the operating costs that sit outside a token invoice, the total cost of AI error (TCAE), and a value ledger measured against displaced labor or revenue. We derive six propositions, including a break-even success-rate gap for routing work to cheaper models and a break-even productivity gain for internal adoption. We calibrate the model to public reporting on Uber's 2025 to 2026 rollout of agentic coding tools, which exhausted a full-year AI budget in four months. The calibrated model reproduces the reported adoption path, average and heavy-user spend, and the month the budget ran out. It finds that the budget failed on the adoption curve rather than on token prices; that a $1,500 monthly cap binds for about 2% of engineers while removing about 13% of demand; that tokens are roughly a third of the cash bill; that the program covers its cash cost once each engineer converts about 2.2 hours a month of freed time into output; and that a session moved from a frontier to a mid-tier model pays only if it loses less than 1.6 percentage points of success probability. Applied to the illustrative example in a recent industry white paper, a 78 to 85% reduction in token cost corresponds to a 30% reduction in the complete bill. We state falsification conditions for each claim. The archive includes the LaTeX source, vector figures, and a reference implementation of the model (JavaScript), along with scripts that regenerate every table and figure.

Zenodo (CERN European Organization for Nuclear Research)
Decent work and economic growth
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Stop Counting Tokens: Why Uber's AI Budget Ran Out in Four Months — Singh Anil Prakash · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS