Goldman’s core warning is that LLM token prices have fallen to roughly $0.97 per million after a 29% August drop; if AI usage does not grow faster than price deflation and efficiency gains, expensive AI infrastructure... The downside is not inevitable: Goldman Sachs Research projects agentic AI could lift monthly to...
Research answer

Create a landscape editorial hero image for this Studio Global article: How is Goldman Sachs’ warning that rapidly falling LLM token prices—down to about $0.97–$0.98 per million tokens after a 29% August decline. Article summary: Goldman’s warning implies a more selective, riskier outlook for AI-infrastructure investors: compute demand can grow sharply while the revenue earned per unit of compute falls even faster. If that happens, utilization, p. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
AI inference is becoming dramatically cheaper. That can broaden access to models and stimulate new workloads, but it also complicates the investment case for the chips, cloud capacity and data centers being built to serve them.
Goldman Sachs’ warning is essentially about a mismatch: infrastructure investment is made years ahead of demand, while the price charged per unit of model usage can fall quickly. Silicon Data’s LLM Token Expenditure Index recently reached $0.97 per million tokens after declining 29% in August, according to reporting on the Goldman note. A separate frontier-model pricing index put prices 84% below its March 2023 baseline. 2
4
5
If usage growth does not offset those declines, the result could be too much costly compute capacity competing for lower-revenue inference workloads. That would matter across the AI infrastructure stack.
A falling token price is not, by itself, bearish. Lower prices can make AI practical for more applications, encourage experimentation and expand total consumption. The challenge for providers is whether growth in paid token volume—and improvements in their own operating efficiency—can exceed the decline in revenue per token.
Goldman’s concern, as reported, is that competitive model pricing may make per-token monetization less durable and ultimately leave the industry with excess infrastructure if computing consumption does not rise fast enough. 1
2
That distinction is important:
The potential mismatch is occurring alongside a very large projected build-out. Goldman Sachs’ baseline estimate totals about $7.6 trillion of AI capital investment across compute, data centers and power from 2026 through 2031. Its model implies annual spending rises from roughly $765 billion in 2026 to $1.6 trillion in 2031; Goldman explicitly notes that these are assumption-sensitive estimates, not certainties. 9
This is why token-price trends matter beyond model providers. A large share of the expected spending supports long-lived assets—accelerators, servers, networking, facilities and power infrastructure. Those investments require sustained workloads and reasonable returns over time.
For semiconductor suppliers, a price-driven slowdown would not necessarily show up first as an immediate collapse in demand. Customers may continue building capacity while they pursue new AI products or attempt to secure supply.
The risk emerges if deployed capacity begins to outrun profitable demand. In that scenario, customers could slow new accelerator and server purchases, become more price-sensitive, or seek greater performance per dollar. Businesses with differentiated hardware, strong software ecosystems or efficient inference offerings may be better insulated than suppliers dependent on broadly similar capacity.
The practical question is whether orders are translating into sustained utilization and revenue-generating workloads—not simply whether equipment is being delivered.
Data-center projects involve major, long-lived commitments to land, power and facilities. If model inference becomes cheaper or more efficient, customers may need less infrastructure for a given level of application output. That could pressure occupancy, lease economics and returns on projects that assumed persistent capacity scarcity.
This does not mean data-center demand disappears. Power availability and the technical requirements of AI deployments can still constrain supply. But a capacity shortage in one location or power market can coexist with weaker economics elsewhere. Investors should separate physical scarcity from evidence that customers are consuming capacity profitably and renewing or expanding commitments.
Cloud platforms and model hosts can benefit from lower inference prices because affordability may bring more developers and enterprise workloads onto their platforms. Yet lower prices also reduce revenue per unit of usage.
To protect gross profit, providers need some combination of:
Competitive pricing in proprietary models is a central part of the risk described in the Goldman-related reporting. 1
2 The key is not the nominal token price alone, but the relationship between customer price, provider cost and total paid usage.
The bullish counterweight is agentic AI: systems that perform multi-step tasks and may consume substantially more tokens than simple chatbot interactions. Goldman Sachs Research forecasts that consumer and enterprise adoption could increase monthly token consumption 24-fold between 2026 and 2030, reaching about 120 quadrillion tokens per month. 7
That forecast makes clear why falling token prices do not automatically signal falling compute demand. If agents move from demonstrations to frequent, valuable tasks, total usage could grow enough to support the infrastructure build-out.
But it is a forecast, not a guarantee. The infrastructure case relies on agents creating recurring workloads that users or businesses will pay for—not merely on a large number of low-cost experimental queries.
Investors assessing the compute-oversupply thesis can focus on a few connected indicators:
Goldman’s warning is not that lower LLM token prices are inherently bad. Cheaper inference can be the mechanism that unlocks wider AI adoption. The risk is that price deflation and model efficiency reduce revenue and required compute faster than agent-driven usage grows.
For AI-infrastructure investors, that turns the story from a simple capacity build-out into a utilization and monetization test. The most resilient businesses may be those able to convert cheaper tokens into durable, high-value workloads while maintaining an advantage in cost, performance or service depth. 1
7
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Goldman’s core warning is that LLM token prices have fallen to roughly $0.97 per million after a 29% August drop; if AI usage does not grow faster than price deflation and efficiency gains, expensive AI infrastructure...
Goldman’s core warning is that LLM token prices have fallen to roughly $0.97 per million after a 29% August drop; if AI usage does not grow faster than price deflation and efficiency gains, expensive AI infrastructure... The downside is not inevitable: Goldman Sachs Research projects agentic AI could lift monthly token consumption 24 fold to about 120 quadrillion by 2030, but that forecast must translate into durable, paid workloads.
The most useful signals are AI revenue growth relative to token price declines, accelerator orders versus actual deployments, and data center utilization and contracted capacity.