Nemotron 3 Ultra sits atop a three-model family that shares the same hybrid Mamba-Transformer MoE architecture .
| Model | Total Parameters | Active per Token | Highlights |
|---|---|---|---|
| Nemotron 3 Nano | ~4 billion | Smallest | Optimized for local and edge inference on RTX PCs and DGX Spark |
| Nemotron 3 Super | 120 billion | 12 billion | Released March 11, 2026; scored 85.6% on PinchBench, the top open model for agentic tasks |
| Nemotron 3 Ultra | 550 billion | 55 billion | Largest model; highest U.S. open-weights intelligence score at launch |
All three support up to a 1-million-token context window. The Super and Ultra variants include NVFP4 training, LatentMoE, and multi-token prediction layers .
First introduced at GTC 2026 in March and reiterated at Computex, NemoClaw and OpenShell form Nvidia's enterprise agent stack .
The Nvidia Vera CPU is a custom Arm-based processor purpose-built for orchestrating AI workloads in data centers .
The Vera CPU serves as the host processor inside Vera Rubin NVL72 rack-scale systems, where it works alongside next-generation GPUs to power AI factories .
RTX Spark is Nvidia’s first dedicated system-on-a-chip for consumer Windows PCs, co-developed with MediaTek and built on TSMC’s 3nm process . Jensen Huang called it “everything we’ve learned over 33 years distilled into one chip” .
Nvidia positions RTX Spark as “the world’s first Windows PCs purpose-built for personal agents” .
Nvidia confirmed that the Vera Rubin platform has entered full production, with systems expected to reach partners in the second half of 2026 .
At the Computex keynote, Nvidia described Vera Rubin as the foundation of the next generation of AI factories .
Taken together, Computex 2026 was the clearest demonstration yet of Nvidia’s vertical integration strategy:
Jensen Huang’s keynote message was singular: Nvidia now supplies every layer of the AI stack, from silicon and open models to agent runtimes and factory-scale deployment infrastructure.