At Hot Chips on August 24, Nvidia said its Groq 3 LPX rack scale accelerator had entered full production as the low latency, long context inference component of the Vera Rubin platform. It is intended for rapid token generation in agentic AI, complementing Vera Rubin NVL72 systems rather than replacing GPU based tra...
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Nvidia announce at Hot Chips 2026 about the full production launch of its Groq 3 LPX dedicated AI inference accelerator—acquired th. Article summary: At Hot Chips on August 24, Nvidia said its Groq 3 LPX rack scale accelerator had entered full production as the low latency, long context inference component of the Vera Rubin platform.. Topic tags: general web, openai, agents, ai, workflow. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clickbait thumbnai
At Hot Chips on August 24, Nvidia said its Groq 3 LPX rack-scale accelerator had entered full production as the low-latency, long-context inference component of the Vera Rubin platform. It is intended for rapid token generation in agentic AI, complementing Vera Rubin NVL72 systems rather than replacing GPU-based training and broader inference work. 36
Nvidia reported 3,400 output tokens per second on Google’s open-source Gemma 4 31B with a 100,000-token context in an Artificial Analysis benchmark; Nvidia’s more precise cited result is 3,431 tokens per second. 16
Nvidia claimed this made LPX four times faster in responsiveness than the nearest alternative platform for these long-context, latency-sensitive workloads. 6
The product’s core purpose is the decode/generation portion of inference—producing the next token quickly and predictably—while Vera Rubin systems handle the broader training, context processing, and AI-factory workload. 36
Nvidia positioned LPX as a seventh major Vera Rubin component, alongside the Vera CPU and the platform’s other chips; the Vera CPU has 88 custom Olympus cores. 2
Nebius is announced as the first AI-cloud adopter and plans to add Groq 3 LPX to its Token Factory production-inference platform. 6
Groq itself also said it would be among the first adopters, deploying LPX alongside Vera Rubin NVL72 in its purpose-built AI-inference cloud. 4
Nvidia’s framing is that future AI factories should be workload-optimized: GPUs and Rubin systems for flexible, large-scale AI compute, with LPX dedicated to the interactive token-generation bottleneck that determines an agent’s perceived responsiveness. 36
The premise that Nvidia “purchased Groq” is not fully accurate. Available reporting describes a $20 billion deal to license Groq technology and hire key personnel, rather than an outright acquisition of the company; Groq subsequently remained active and raised new funding. 5
The supplied evidence does not adequately substantiate the specific hardware claims of 256 chips per rack, 500 MB of on-die SRAM per chip, Samsung fabrication versus TSMC GPU production, nor the claimed competitive responses from AMD/Cerebras or OpenAI’s 750-token-per-second mode. Insufficient evidence.
Likewise, the evidence supports that SpaceXAI plans to deploy Vera CPUs for agentic-AI applications, but does not establish the detailed claims about tool execution, simulation across orbital satellites, or the exact scope and timing of that deployment. 7
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
At Hot Chips on August 24, Nvidia said its Groq 3 LPX rack scale accelerator had entered full production as the low latency, long context inference component of the Vera Rubin platform.
At Hot Chips on August 24, Nvidia said its Groq 3 LPX rack scale accelerator had entered full production as the low latency, long context inference component of the Vera Rubin platform. It is intended for rapid token generation in agentic AI, complementing Vera Rubin NVL72 systems rather than replacing GPU based training and broader inference work.
[3][6] What Nvidia announced Nvidia reported 3,400 output tokens per second on Google’s open source Gemma 4 31B with a 100,000 token context in an Artificial Analysis benchmark; Nvidia’s more precise cited result is 3,431 tokens per second.