Gartner's August 2026 report details the following spending forecasts for AI-optimized cloud infrastructure:
For context, in the early years of the generative AI boom, training dominated infrastructure spending. That dynamic has now completely inverted .
The growth is not limited to inference. Total AI-optimized IaaS spending is forecast to reach $42.3 billion in 2026, a 96.4% increase from $21.5 billion in 2025. In 2027, the market is expected to grow further to $66.1 billion, a 56.5% increase .
Overall worldwide IaaS spending (including non-AI workloads) is forecast to reach $287.3 billion in 2026 (up 29.3% from $222.2 billion in 2025) and $359.9 billion in 2027 .
According to Gartner, two forces are driving inference spending past training:
Gartner analyst Hardeep Singh noted: "This shift is accelerating cloud consumption patterns and creating sustained demand for AI-optimized infrastructure" .
This pattern is confirmed by other analysts. Deloitte's Tech Trends 2026 estimates that inference already accounts for two-thirds of all AI compute in 2026 . The Futurum Group's independent forecast shows managed inference reached $23.1 billion in 2025 (58.6% of infrastructure spend), with inference overtaking training back in 2025 .
Gartner's spending figures sit within a much larger picture:
Despite rapid cost declines — Deloitte notes inference cloud API pricing has fallen nearly 80% year-over-year — overall spending is exploding upward because usage is growing even faster than unit costs are dropping .
Gartner also issued a cautionary finding: at least 50% of GenAI initiatives will exceed their planned budgets by 2028 due to poor architectural choices and a lack of operational expertise .
Inference is the primary culprit. Unlike training — a large upfront expense — inference costs recur every time a user or application calls a model. Gartner expects inference to account for at least 70% of a model's lifetime costs .
Gartner recommends the following strategic actions:
The "Inference Flip" is not merely a spending statistic; it represents a fundamental restructuring of AI infrastructure priorities. The era of building ever-larger models as the primary cost driver is giving way to the era of running those models at scale in production. For cloud providers, this means sustained demand for GPU and ASIC capacity. For enterprises, it means that infrastructure planning must now center on the operational costs of AI, not just the upfront investment in model training.
As Gartner's report makes clear, the shift from development to deployment has arrived, and the budget line items will never look the same again.