LangChain's Deep Agents harness, optimized for Nemotron 3 Ultra, delivers inference costs up to 10x lower per run compared to certain leading closed-source models . On Artificial Analysis, Nemotron 3 Ultra costs $0.58 per 1M tokens compared to Claude Sonnet 4.6 at $2.31 — a 75% cost reduction
. Chinese financial media reported a claim that inference costs are ~90% lower than closed models
.
Nvidia states that on LangChain's Deep Agents benchmarks, Nemotron 3 Ultra matched the highest-scoring model on business tasks (parity in quality) . Key benchmark scores include:
Nemotron 3 Ultra achieves up to 5.9x higher inference throughput compared to GLM-5.1-754B-A40B on the 8K token input / 64K token output setting . On a pre-release DeepInfra endpoint, it served over 300 tokens per second
.
LangChain provided Day 0 support for Nemotron 3 Ultra in its Deep Agents harness and is a member of the Nemotron Coalition . LangChain CEO Harrison Chase said enterprises can achieve strong performance at a fraction of the cost of closed models
. Nvidia stated that LangChain's Deep Agents harness completed "more tasks at higher throughput using Nemotron 3 Ultra while delivering inference costs up to 10x lower than some leading closed models"
.
Global partners EY and others such as Abridge, Amdocs, and Box are already engaged with the NemoClaw + Nemotron 3 Ultra stack for building enterprise AI agents . Cadence, Siemens, Synopsys, and Dassault Systèmes are also deploying autonomous AI engineers using Nvidia's NemoClaw blueprints
.
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 ollama.com/library/nemotron-3-ultra Nvidia (NVDA) shares gained 4% on Wednesday, July 9, as the new benchmark results and NemoClaw announcement suggested Nemotron 3 Ultra could challenge closed models on both cost and enterprise performance . Bank of America reiterated its Buy rating on Nvidia stock following the announcement, citing the model's cost-performance advantage and enterprise momentum
.