Microsoft is deliberately splitting its next-generation AI infrastructure between AMD Helios and Nvidia Vera Rubin to reduce supply chain risk and avoid single-vendor lock-in . This mirrors the company's model-agnostic approach within Azure and Copilot . CEO Satya Nadella confirmed that "we will be among the first cloud providers to deploy next generation rack scale AI infrastructure based on AMD Helios and Nvidia Vera Rubin," giving Microsoft negotiating leverage and operational flexibility if one vendor faces allocation constraints . The strategy is widely described in the industry as "hedging its bets" on AI infrastructure procurement .
Neither AMD nor Microsoft have disclosed exact pricing, but analyst estimates from the Futurum Group provide a clear comparison:
This makes Helios roughly 40% more expensive on a rack-for-rack basis . The premium reflects Helios's higher memory capacity, open architecture, and AMD's integrated CPU-GPU-networking portfolio .
AMD argues that the higher upfront cost of Helios is justified by superior performance efficiency. According to AMD's own lab data, Helios delivers :
AMD CEO Lisa Su presented these figures at the Advancing AI 2026 event, stating that Helios offers "15% or more performance than the competition on the largest models, 50% more HBM capacity, and up to 30% more tokens per dollar" .
However, analysts caution that these are pre-production estimates. The Futurum Group noted that "AMD's 15% performance edge and 30% tokens-per-dollar claims remain vendor math against an unshipped competitor and will be tested in production within two quarters" .
AMD Helios is a purpose-built, liquid-cooled rack-scale system that represents the company's most comprehensive AI hardware offering to date :
| Component | Specification |
|---|---|
| GPUs | 72x AMD Instinct MI455X |
| GPU Architecture | CDNA 5 (2nm/3nm chiplet hybrid at TSMC) |
| Transistors per GPU | 320 billion |
| CPU | 18x 6th Gen AMD EPYC "Venice" (one per compute tray, up to 256 cores per CPU) |
| Networking | AMD Pensando "Vulcano" 800 AI NICs (800 Gbps Ethernet, up to 2.4 Tbps per GPU scale-out) |
| Scale-up Interconnect | UALink (Ultra Accelerator Link tunneled over Ethernet) |
| Scale-out Interconnect | Ultra Ethernet Consortium (UEC) aligned |
| Total HBM4 Memory | 31 TB per rack |
| Memory per GPU | 432 GB (12 stacks of HBM4 at 36 GB each) |
| Memory Bandwidth per GPU | 19.6 TB/s |
| Aggregate Memory Bandwidth | 1.7 PB/s per rack |
| FP4 Performance | 2.9 exaFLOPS per rack |
| FP8 Performance | 1.4 exaFLOPS per rack |
| Scale-up Bandwidth | 260 TB/s |
| Scale-out Bandwidth | 43 TB/s |
| Form Factor | 18 liquid-cooled compute trays (4 GPUs each) + 6 switch trays, OCP Open Rack Wide compliant |
| Per-GPU Performance | 40 PFLOPS FP4, 20 PFLOPS FP8 |
Each compute tray contains four MI455X GPUs paired with a single EPYC Venice CPU and Pensando networking, all connected via the UALink protocol to allow the 72 GPUs to function as a single compute unit .
AMD has secured the following confirmed early Helios customers :
One notable absence: Anthropic has not been publicly named as an early Helios customer in any major announcement as of July 2026 .
Microsoft's decision to deploy both AMD Helios and Nvidia Vera Rubin signals that the AI hardware market has entered a new phase. Rather than a single-vendor dependency, hyperscalers are investing in platform diversity to secure supply, improve negotiating leverage, and access differentiated performance characteristics .
AMD Helios, with its massive 31 TB of HBM4 memory and open Ethernet-based interconnect, is particularly well-suited for large-scale inference workloads, where memory capacity and bandwidth directly drive throughput and cost-per-token . With shipments expected in the second half of 2026, the real test will come when production benchmarks are published and the competing claims can be verified at scale .