Nvidia's Blackwell Ultra GB300 NVL72 system topped the first AgentPerf benchmark, running up to 61,340 concurrent AI coding agents and achieving 20x more agents per megawatt than the previous generation H200 Hopper pl... AgentPerf is the industry's first open benchmark for agentic AI, measuring how many multi step c...

Create a landscape editorial hero image for this Studio Global article: What did Nvidia achieve in the first published results of Artificial Analysis's AgentPerf benchmark, what does this new benchmark measure, a. Article summary: Here are the key findings from the first published results of Artificial Analysis's **AA-AgentPerf** benchmark, announced on June 12, 2026.. Topic tags: general, documentation, general web, user generated. Reference image context from search candidates: Reference image 1: visual subject "We measure real-world performance of AI accelerator systems during language model inference. ## AA-AgentPerf: The Hardware Benchmark for the Agent Era. AA-AgentPerf has been shaped" source context "AI Hardware Benchmarking & Performance Analysis" Reference image 2: visual subject "For years, co-founder and chief executive officer Jensen Huang and other higher-ups at Nvidia have
AI benchmarks are rapidly evolving beyond measuring simple question-and-answer speed. The new frontier is agentic AI—autonomous systems that use tools, write code, and chain together multiple reasoning steps to complete complex tasks. The industry now has its first dedicated benchmark for this demanding workload, and Nvidia's newest hardware has posted dominant initial scores.
On June 12, 2026, Artificial Analysis published the first round of results for its AA-AgentPerf benchmark, designed specifically for agentic AI inference. Nvidia's Blackwell Ultra-based GB300 NVL72 rack-scale system not only achieved the highest overall performance but did so with a dramatic leap in efficiency over the prior generation, running up to 20 times more agents per megawatt than an Nvidia HGX H200 system .
Traditional LLM benchmarks often focus on synthetic queries and single-turn completions. AgentPerf is fundamentally different. It is the industry's first open, multi-vendor hardware benchmark built to stress-test real-world agentic AI workloads .
Instead of generating a single response, AgentPerf replays authentic coding agent trajectories sourced from public repositories across more than 12 programming languages. These trajectories chain together up to 20 sequential LLM calls, interspersed with tool-use simulations that incorporate realistic CPU delays, all while managing growing context windows . The result is a far more demanding test that mimics how a modern AI coding assistant actually behaves when fixing a bug or building a feature.
The core metric is the number of concurrent agents a system can support while meeting a strict service-level objective (SLO) for output token speed and time-to-first-token (TTFT) . The initial results were run on DeepSeek V4 Pro, a large mixture-of-experts (MoE) model chosen as a representative example of the frontier models powering advanced agents
.
On this rigorous new test, the GB300 NVL72 system took a decisive lead. The raw throughput figures from Artificial Analysis's social media post paint a stark picture at the least demanding SLO tier (20 tokens per second, 10 seconds TTFT) :
The 20x efficiency gain over the prior-generation Hopper platform held true at multiple SLO tiers, including the more stringent 60 tokens/s requirement, demonstrating that the performance advantage is not limited to a single testing point .
These numbers do not exist in a vacuum. Nvidia is using the AgentPerf results to cement a narrative around full-stack optimization that goes far beyond raw hardware specs. The company attributes the 20x efficiency gain to a combination of tightly integrated technologies: the NVLink scale-up fabric that links 72 GPUs into a single coherent system, custom CUDA kernels that overlap communication and computation specifically for MoE architectures, and TensorRT LLM optimizations like WideEP, DeepEP, DeepGEMM, and fused MoE kernels that maintain high utilization as the number of concurrent agent sessions scales .
The AgentPerf win also completes a clean sweep of the major AI infrastructure benchmarks for Blackwell Ultra. In MLPerf Inference v5.1, the same GB300 NVL72 system set records on DeepSeek-R1, delivering 1.4x the throughput of the previous Blackwell-based GB200 . In MLPerf Training v5.1, Blackwell Ultra achieved the fastest time-to-train on all seven benchmarks, including pretraining Llama 3.1 405B in just 10 minutes using 5,120 GPUs
.
Crucially, Nvidia is not just publishing records. The company is already pointing to production deployments as evidence that Blackwell Ultra is ready for real agentic workloads. Together AI is using the platform to power agentic coding for Cursor, and DeepInfra is running the AI workforce for Pam.ai on Blackwell . The blog post promoting the results also explicitly fast-forwards to the next architecture, noting that the Vera Rubin platform is now in production, aiming to provide even more capacity for the coming wave of agentic AI
.
For an industry that is increasingly betting its future on autonomous AI agents that can reason, code, and act, Nvidia is making a clear statement: the infrastructure is ready, and it is starting from a position of overwhelming performance leadership.
Studio Global AI
Use this topic as a starting point for a fresh source-backed answer, then compare citations before you share it.
Nvidia's Blackwell Ultra GB300 NVL72 system topped the first AgentPerf benchmark, running up to 61,340 concurrent AI coding agents and achieving 20x more agents per megawatt than the previous generation H200 Hopper pl...
Nvidia's Blackwell Ultra GB300 NVL72 system topped the first AgentPerf benchmark, running up to 61,340 concurrent AI coding agents and achieving 20x more agents per megawatt than the previous generation H200 Hopper pl... AgentPerf is the industry's first open benchmark for agentic AI, measuring how many multi step coding agents can run simultaneously on real workloads, not simple chat completions.
The win is part of Nvidia's full stack push to define the infrastructure for the coming wave of autonomous AI, with partners like Together AI already using Blackwell Ultra to power agentic coding in production.