Nvidia and d Matrix are planning a hybrid inference system: GPUs handle compute heavy prefill while d Matrix Raptor XPUs handle token by token decode, connected through NVLink Fusion. NVLink Fusion, MGX reference designs and Nvidia networking are meant to reduce the integration risk of deploying a specialized infere...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: How is AI-chip startup d-Matrix’s September 2026 partnership with Nvidia—integrating Nvidia’s NVLink Fusion into its next-generation Raptor. Article summary: The partnership makes d-Matrix’s specialized inference silicon a first-class component of Nvidia-style AI factories rather than a separate appliance. The intended architecture is heterogeneous: Nvidia GPUs handle compute. Topic tags: general, news, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers,
Generative AI services are increasingly judged by how quickly they begin and continue producing an answer. Nvidia and d-Matrix’s September 2026 collaboration is aimed at that problem: connect d-Matrix’s forthcoming Raptor inference XPUs directly into Nvidia-designed AI infrastructure, so a rack can assign different parts of an LLM request to the processor best suited to them. 1
3
4
The key point is complementarity. This is not an announcement that Raptor replaces Nvidia GPUs. It is a plan for a heterogeneous system in which GPUs perform compute-intensive prefill work and d-Matrix accelerators focus on the repetitive, latency-sensitive decode phase that generates output tokens one at a time. 4
5
An LLM request has distinct performance demands. Prefill processes the prompt and context; decode repeatedly generates the next token. d-Matrix positions its inference architecture, including its current Corsair product, for decode work alongside GPUs, which it describes as strong at compute-intensive workloads. 26
That division can be useful for interactive products where time to the next token matters—such as coding assistants, chat applications and voice experiences. The proposed Raptor platform is aimed at what d-Matrix calls premium, ultra-low-latency token services. 4
8
But separating workload phases creates a systems challenge: the processors must exchange data quickly enough that communication does not erase the benefit of specialization. The collaboration’s infrastructure choices are intended to address that constraint.
Raptor is planned to use NVLink Fusion to join Nvidia’s NVLink scale-up domain, while Spectrum-X provides the scale-out network between systems. Nvidia says this gives d-Matrix a route to connect custom silicon to its wider AI infrastructure platform. 3
The planned d-Matrix rack is based on Nvidia’s MGX reference architecture and is expected to incorporate Nvidia Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs and Spectrum-X Ethernet. 4
8
For d-Matrix, this matters as much as the interconnect itself. Instead of independently qualifying every server, network, power, cooling and supply-chain element around a new processor, the company can build toward a validated rack ecosystem. Nvidia explicitly frames MGX and its broader platform as an accelerated, lower-risk path from custom silicon to large-scale deployment. 3
In practical terms, the partnership is intended to make Raptor deployable as part of an Nvidia-style AI factory rather than as an isolated inference appliance.
d-Matrix has outlined an ambitious rack-scale configuration for Raptor: up to 144 XPUs per rack, 2.3 TB of 3D-stacked DRAM and 7.2 PB/s of aggregate memory bandwidth. The company says the design is intended to support models larger than four trillion parameters at 4-bit precision. These are planned specifications, not independently validated shipping results. 4
The architecture reflects the company’s view that LLM decode is strongly shaped by memory movement and available bandwidth. Raptor’s proposed 3D package combines a DRAM memory chip with an SRAM compute chip, extending d-Matrix’s memory-centric approach. 4
d-Matrix says Raptor is expected to tape out by the end of 2026, while initial availability in an MGX rack is expected in the fourth quarter of 2027. That leaves substantial execution work ahead: tape-out, manufacturing, packaging, system validation and rack-scale deployment all need to succeed before customers can assess real-world performance. 4
8
Nvidia’s role in the collaboration is strategically notable. The company is allowing a specialized external XPU to connect into its rack architecture and networking fabric rather than requiring every inference stage to run on an Nvidia GPU.
Nvidia characterizes its AI platform as “vertically integrated and horizontally open”: it can retain the value of its scale-up fabric, scale-out networking, rack designs and ecosystem while enabling other processors to participate. 3
That expands the set of workloads the platform can address. Customers looking for a decode-focused accelerator may be more likely to remain inside Nvidia-compatible infrastructure if that accelerator can use NVLink Fusion, MGX and Spectrum-X. The result is better understood as controlled complementarity than a retreat from GPUs: Nvidia GPUs remain central to broad-purpose training, inference and the compute-heavy prefill stage, while a specialized device is positioned for decode. 1
4
5
Raptor follows d-Matrix’s Corsair inference accelerator. The company says Corsair is purpose-built for decode in disaggregated inference and works with GPUs rather than replacing them. 26
d-Matrix has publicized performance, cost and energy claims for Corsair-based hybrid systems, but those figures should be treated as company claims unless independently reproducible benchmark details are available. The evidence supplied for this announcement does not establish a universal performance gain across models, batch sizes or deployment configurations. 25
28
The company raised a $275 million Series C at a reported $2 billion valuation, bringing its reported total funding to $450 million; the round included M12, Microsoft’s venture fund, among other investors. 21
30 Reuters also reported that Microsoft backed d-Matrix in a 2023 financing round and that d-Matrix had shipped its first AI chip in November 2024.
1
Those milestones suggest d-Matrix is moving from a single-product inference pitch toward a broader rack-scale roadmap. The Nvidia integration could make that transition more credible to customers that already buy or operate Nvidia-based infrastructure.
The financial terms of the Nvidia agreement were not disclosed. 1 As a result, it is not possible to determine from the public announcement whether the arrangement includes licensing fees, joint engineering commitments, volume purchases, revenue sharing or an Nvidia investment.
The broader conclusion is clearer than the commercial detail: both companies are planning for an inference market where one processor type does not have to execute every step of an LLM request. d-Matrix gains a path into Nvidia-compatible AI factories; Nvidia gains a specialized decode option that can be integrated with its networking and rack platform. Whether that strategy delivers lower latency or better economics at scale will depend on the eventual Raptor hardware, software orchestration and customer benchmarks after the planned 2027 availability. 3
4
8
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Nvidia and d Matrix are planning a hybrid inference system: GPUs handle compute heavy prefill while d Matrix Raptor XPUs handle token by token decode, connected through NVLink Fusion.
Nvidia and d Matrix are planning a hybrid inference system: GPUs handle compute heavy prefill while d Matrix Raptor XPUs handle token by token decode, connected through NVLink Fusion. NVLink Fusion, MGX reference designs and Nvidia networking are meant to reduce the integration risk of deploying a specialized inference processor at rack scale—not to replace Nvidia GPUs.
The deal shows Nvidia’s interest in making its AI factory infrastructure interoperable with specialized third party silicon while keeping NVLink, networking and rack scale systems central to the deployment.