2026年7月8日,英伟达与LangChain联合发布NemoClaw for LangChain Deep Agents蓝图,基于开源Nemotron 3 Ultra模型——550B总参数/55B活跃参数的MoE混合Mamba Attention架构。 Nemotron 3 Ultra在Artificial Analysis Intelligence Index上获得47.7 48分,为美国开放权重模型最高分;推理成本仅$0.58/百万token,相比Claude Sonnet 4.6的$2.31降低75%。
研究答案

Create a landscape editorial hero image for this Studio Global article: Search & fact-check with cited sources for What did Nvidia announce on Tuesday about its open-source Nemotron 3 Ultra AI model, including it. Article summary: Here is a detailed, sourced fact-check on all aspects of your query.. Topic tags: general, documentation, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clickbait thumbnails, icons, and tiny thumbnail layouts. Make it useful as an illustrative visual, not as factual evidence.
2026年7月8日(周二),英伟达与LangChain联合发布NemoClaw for LangChain Deep Agents蓝图,这是一套基于英伟达开源Nemotron 3 Ultra模型构建企业AI代理的参考架构。该发布重点展示了相比闭源模型的巨大推理成本优势和企业级性能表现。
NemoClaw for LangChain Deep Agents蓝图为企业提供了一套构建开放代理系统的参考架构,结合了Nemotron 3 Ultra和英伟达OpenShell运行时策略控制。该架构专为需要可定制、可投入生产的AI代理且希望降低推理成本的大型企业客户设计。 7月8日发布的NemoClaw v0.0.76版本已将Nemotron 3 Ultra设为Deep Agents Code的英伟达端点默认模型。
针对Nemotron 3 Ultra优化的LangChain Deep Agents套件,每次运行的推理成本相比某些领先闭源模型降低高达10倍。 在Artificial Analysis平台上,Nemotron 3 Ultra每百万token成本为**$0.58**,而Claude Sonnet 4.6为$2.31——成本降低75%。
据中国财经媒体报道,其推理成本较闭源模型降低约90%。
英伟达表示,在LangChain Deep Agents基准测试中,Nemotron 3 Ultra在商业任务上达到了与最高分模型持平的性能(质量对等)。 关键基准测试得分包括:
| 规格 | 详情 |
|---|---|
| 总参数量 | 5500亿(550B) |
| 活跃参数量 | 每token 550亿(55B) |
| 架构类型 | 混合专家(MoE)混合Mamba-Attention(Mamba-Transformer) |
| 上下文长度 | 100万token(1M) |
| 预训练 | 20万亿文本token |
| 后训练 | SFT + RL + 多层同策略蒸馏(MOPD) |
| 许可证 | OpenMDW 1.1 / Linux Foundation 宽松许可证 |
在8K token输入/64K token输出设置下,Nemotron 3 Ultra的推理吞吐量相比GLM-5.1-754B-A40B提升高达5.9倍。 在预发布的DeepInfra端点上,其输出速度超过300 tokens/秒。
LangChain为Deep Agents套件中的Nemotron 3 Ultra提供了Day 0支持,并已成为Nemotron Coalition成员。 LangChain首席执行官Harrison Chase表示,企业可以以闭源模型一小部分的成本获得强大性能。
英伟达则表示,LangChain的Deep Agents套件"使用Nemotron 3 Ultra完成更多任务、实现更高吞吐量,同时推理成本比某些领先闭源模型低10倍"。
全球合作伙伴EY以及Abridge、Amdocs和Box等企业已开始使用NemoClaw + Nemotron 3 Ultra堆栈构建企业AI代理。 Cadence、Siemens、Synopsys和Dassault Systèmes也在使用英伟达的NemoClaw蓝图部署自主AI工程师。
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16ollama.com/library/nemotron-3-ultra 获取2026年7月9日(周三),英伟达(NVDA)股价上涨4%,因新的基准测试结果和NemoClaw发布表明Nemotron 3 Ultra可能在成本和性能两个维度挑战闭源模型。 美国银行在发布后重申了对英伟达股票的买入评级,理由是模型在性价比上的优势和企业级发展势头。
Studio Global AI
此页面包含一个有来源支持的答案,您可以在 Studio Global 内继续。
2026年7月8日,英伟达与LangChain联合发布NemoClaw for LangChain Deep Agents蓝图,基于开源Nemotron 3 Ultra模型——550B总参数/55B活跃参数的MoE混合Mamba Attention架构。
2026年7月8日,英伟达与LangChain联合发布NemoClaw for LangChain Deep Agents蓝图,基于开源Nemotron 3 Ultra模型——550B总参数/55B活跃参数的MoE混合Mamba Attention架构。 Nemotron 3 Ultra在Artificial Analysis Intelligence Index上获得47.7 48分,为美国开放权重模型最高分;推理成本仅$0.58/百万token,相比Claude Sonnet 4.6的$2.31降低75%。
EY、Abridge、Amdocs和Box等企业已采用NemoClaw堆栈;开发者可通过Hugging Face、英伟达NIM、Ollama或本地部署获取模型。