Microsoft is OpenAI's primary investor and the exclusive reseller of its models via Azure, yet internally it found that engineers were using expensive frontier models for trivial tasks, driving costs that even the company's own finance team couldn't tolerate . Executive VP Jay Parikh sent an email stating "tokenmaxxing is not what we are optimizing for" and announced that divisions would receive AI token budget targets starting July 2026 .
"Tokenmaxxing" became a gamified behavior. Engineers competed on internal leaderboards to maximize token use — the AI equivalent of treating compute like fantasy football points . The practice inflated consumption without measurable business value, as companies discovered that "throwing AI at everything" was raising costs without a proportional spike in productivity .
Satya Nadella himself had to intervene, telling employees to stop using the most powerful AI models for every task and to match the model to the problem . The company switched its internal default model to the cheaper GPT-5.6 and revoked Claude Code licenses for most employees to stop the bleeding .
The industry is moving from token volume as a vanity metric to outcome-based measurement of AI return.
Retiring token leaderboards. Amazon and Meta both retired internal AI token usage leaderboards after employees gamed them . Meta CTO Andrew Bosworth warned staff: "All motion is not progress and token usage alone is not a measure of impact of any kind" .
Replacing token counts with deployment metrics. Amazon replaced its "KiroRank" system with "normalised deployments" — counting AI-assisted code that actually reaches production .
Per-engineer caps and model tiering. Meta is considering per-engineer AI token caps . Companies are moving to "model-level entitlements" that lock the right (cheaper) model to the right role, and setting spend-threshold alerts .
Budget governance over usage maximization. Atlassian capped spending after its monthly AI bill ballooned from roughly $5 million to over $15 million . Citi temporarily revoked access to the newest Claude and ChatGPT models for developers after they became "the primary driver of elevated enterprise consumption" of pooled AI credits . Anthropic shipped enterprise spend-control tools in response to 78% of businesses reporting surprise AI bills .
The core lesson: tokens are not seats — unlike traditional SaaS, generative AI scales costs with every prompt, retry, and agent loop . The industry is now treating AI not as an unlimited productivity elixir but as a governed resource requiring the same financial discipline as any other critical input, measuring success by shipped code, resolved tickets, or revenue impact rather than raw token burn.