MLPerf Inference v6.1 shows that AI inference leadership is increasingly a systems question: AMD posted the largest aggregate 512 GPU MI355X result, while NVIDIA previewed Vera Rubin at up to 3.7× GB300 NVL72 throughp... The round drew a record 30 submitting organizations and 486 datacenter and edge results, while a...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What were the key results and broader implications of MLCommons’ MLPerf Inference v6.1 benchmarks, including AMD’s record-setting MI355X per. Article summary: MLPerf Inference v6.1 showed a more competitive and more deployment-oriented inference market: AMD set aggregate-scale records, NVIDIA demonstrated its next platform while retaining strong Blackwell results, Intel broade. Topic tags: general, documentation, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks,
MLPerf Inference v6.1 is less a single-chip leaderboard than a snapshot of how AI serving is changing. The benchmark’s biggest signals were AMD’s extreme-scale MI355X submission, NVIDIA’s first Vera Rubin preview data, Intel’s broader accessible-inference coverage, and new tests for retrieval-augmented generation (RAG) and edge agents. Together, they make a simple point: the useful unit of comparison is no longer just a GPU—it is a complete, available serving system running a relevant workload under the required latency constraints.
AMD’s headline result came from Crusoe’s 512-GPU Instinct MI355X configuration, which posted the largest aggregate token-throughput figures reported in the v6.1 round. 5 That is an important scale demonstration: it tests more than accelerator performance, touching distributed inference, networking and system integration across a very large deployment.
It is not, however, a universal declaration of platform leadership. A 512-GPU aggregate result should not be compared directly with a single rack or smaller cluster without normalizing for GPU count, model, benchmark scenario and service-level requirements. Aggregate throughput can be the right measure for a large inference fleet, while per-accelerator throughput, rack density, power use, availability and operational complexity can matter more for a specific deployment.
AMD’s case is also a software story. The company highlighted higher performance from the same MI355X hardware over the previous MLPerf cycle, attributing the progress to ROCm optimization. The broader implication is straightforward: compilers, kernels, model runtimes and distributed-serving software can raise the useful capacity of an installed fleet without a hardware refresh.
AMD has also committed to a six-week ROCm release cadence. If that cadence produces sustained, deployable improvements, it could make software velocity a meaningful part of buyers’ cost-per-token calculations. But MLPerf gains should be validated against a buyer’s own model mix, quantization choices, framework, traffic pattern and latency objectives; improvements in one benchmark do not automatically transfer one-for-one to production.
NVIDIA’s Vera Rubin NVL72 made its first MLPerf Inference appearance as a preview submission. NVIDIA reported up to 3.7× higher throughput than GB300 NVL72 on Qwen3-VL, and described a 288-GPU GB300 NVL72 submission that achieved 99% scaling efficiency across four racks. 2
The key qualifier is availability. Preview results are useful evidence of architectural direction and software maturity, but they are not automatically a like-for-like purchasing comparison with generally available systems. Cluster sizes in preview submissions also varied, making per-GPU or otherwise normalized comparisons more informative than total system throughput alone. 8
Blackwell and Blackwell Ultra remain central to the near-term deployment picture. For example, CoreWeave reported leading cloud-provider results on several Blackwell and Blackwell Ultra configurations across Qwen3-VL, DeepSeek-R1 and GPT-OSS-120B workloads. 4 The practical takeaway is segmentation rather than a single winner: large-cluster aggregate capacity, rack-scale serving, multimodal inference and interactive latency can favor different configurations.
Intel’s submissions expanded the visibility of CPU and workstation-class GPU inference. Its four-GPU Arc Pro B70 node combined 128 GB of VRAM and submitted results for Llama, GPT-OSS-120B, Whisper and the new end-to-end RAG workload. Intel also reported software-driven gains on the same Arc Pro B70 configuration versus the prior round. 3
That does not position Arc Pro as a direct replacement for hyperscale accelerator clusters. It does make the benchmark more representative of organizations evaluating smaller, lower-cost or more flexible deployments—especially where memory capacity, local inference or existing server infrastructure shape the decision.
MLCommons reported a record 30 submitting organizations and 486 datacenter and edge results in v6.1. The suite added an end-to-end datacenter RAG benchmark and an Edge Agentic Inference benchmark, reflecting the move from one-shot model prompts to multi-stage AI applications. 9
That expansion matters because agentic workloads introduce bottlenecks that a conventional offline token-throughput test may not expose: multi-turn interactions, tool use, orchestration overhead and responsiveness between steps. MLCommons describes the edge benchmark as reflecting complex, multi-turn tasks such as agentic coding. 9
These are still controlled benchmarks, not a complete representation of live production services. Yet they move evaluation closer to the workloads many teams are now trying to deploy.
MLCommons also previewed MLPerf Endpoints as a possible next direction. The important distinction is between periodic benchmark submissions and observable performance from running serving endpoints. A mature endpoint-oriented program could put more emphasis on behavior under load, latency consistency, availability and operational repeatability—not only a peak result from a fixed benchmark run.
For now, buyers should view that as an emerging measurement direction rather than a finalized replacement for the existing submission process. The v6.1 results themselves underscore why it could be valuable: deployment quality depends on far more than a maximum tokens-per-second figure.
MLPerf v6.1 offers strong evidence that inference competition is being shaped by three connected factors: accelerator capability, software improvement and systems integration. AMD’s large MI355X submission demonstrates the importance of scale; NVIDIA’s Vera Rubin preview signals the next performance step while Blackwell serves current deployments; Intel’s Arc Pro B70 results widen the range of hardware represented; and MLCommons’ new tests acknowledge that RAG and agents create different performance demands.
When evaluating a platform, start with the workload and operating target: the model, context length, concurrency, latency SLO, power envelope, deployment size and cost per useful output. Then compare systems at equivalent scale and availability. That approach is more useful than treating any one MLPerf headline as a universal answer.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
MLPerf Inference v6.1 shows that AI inference leadership is increasingly a systems question: AMD posted the largest aggregate 512 GPU MI355X result, while NVIDIA previewed Vera Rubin at up to 3.7× GB300 NVL72 throughp...
MLPerf Inference v6.1 shows that AI inference leadership is increasingly a systems question: AMD posted the largest aggregate 512 GPU MI355X result, while NVIDIA previewed Vera Rubin at up to 3.7× GB300 NVL72 throughp... The round drew a record 30 submitting organizations and 486 datacenter and edge results, while adding end to end RAG and edge agentic inference tests.
For infrastructure buyers, software progress on existing hardware and reproducible deployment matter alongside accelerator specifications.