Three specific pressures forced Microsoft to impose internal AI discipline:
Unchecked internal costs. Microsoft's own engineering teams had been using frontier AI models freely via GitHub Copilot, driving token consumption and inference costs far beyond what the company considered sustainable. Internal Copilot guidelines acknowledged that many engineers were spending "in the range of hundreds of dollars a month to a few thousand dollars in tokens" . Parikh's email made clear that as of July 2026, every division would operate under an AI token budget target, and individual employees could track their own usage via an internal dashboard
.
Misaligned incentives. Engineers had begun treating high token consumption as a badge of productivity, a practice known across the industry as tokenmaxxing. But Microsoft found that more tokens did not equate to more business value. Parikh reframed the company's goal explicitly: "We are not optimizing for fewer tokens. We are optimizing for impact per token" .
The credibility problem. As a company that sells AI subscriptions — Azure OpenAI Service, GitHub Copilot, Microsoft 365 Copilot — Microsoft could not credibly tell customers to control AI costs while its own teams burned through tokens without restraint . Parikh's email noted that the company would "manage token spend with the same discipline we apply to every other critical resource"
.
Microsoft's move is one data point in a much wider correction playing out across the tech industry. "Tokenmaxxing" — the practice of aggressively pushing AI token consumption as a proxy for innovation — became widespread among tech companies from roughly late 2025 through mid-2026, but it is now rapidly collapsing because the economics did not work .
Uber's budget blowout. Uber Technologies exhausted its entire 2026 AI budget in just a few months. Its CTO publicly declared, "We're coming to the end of the so-called tokenmaxxing era," adding that raw token consumption was not translating into measurable returns . The company is now moving toward cost-per-interaction optimization through prompt caching, more efficient default models, and better visibility into consumption
.
Gartner's warning. Gartner has predicted that AI coding costs will soon surpass average developer salaries under consumption-based pricing models, forcing CFOs across industries to demand cost discipline .
The startup math. Startups have demonstrated that switching models can cut inference costs by roughly 90%. Lindy.ai, for example, moved 100% of its API traffic from Claude to DeepSeek overnight, citing roughly 90% lower inference costs for comparable output quality on its core task set . Such dramatic savings make the old "more tokens = better" mentality financially untenable
.
Forbes and Fortune verdict. Multiple business outlets have declared the tokenmaxxing era structurally over. Enterprises are abandoning raw usage as a metric in favor of efficiency-first deployments, model routing (using cheaper models for simpler tasks), and outcome-based billing that removes the incentive to maximize tokens . Forbes characterized tokenmaxxing as "a transitional phase in AI adoption" that helped normalize usage but is "not a sustainable operating model"
.
The industry pattern is now clear: companies that initially encouraged unfettered AI usage are imposing budgets, switching to cheaper models, and demanding that every token spent tie to a measurable business outcome. ServiceNow's chief customer officer has warned that tokenmaxxing is an AI hype cycle . Fortune's verdict was blunt: "Tokenmaxxing is dead. It didn't produce the AI ROI companies wanted"
.
Microsoft's internal token budgets and GPT-5.6 default are a textbook example of this industry-wide reckoning — coming not from an AI laggard, but from one of the sector's largest investors. If the company selling AI subscriptions needs a budget, the message for every other enterprise is clear.