d Matrix plans to use Nvidia NVLink Fusion to place its next generation Raptor inference XPUs inside Nvidia MGX based AI infrastructure, pairing GPU prefill with specialized token decoding. The arrangement gives d Matrix access to Nvidia’s NVLink and Spectrum X networking, plus rack components including Vera CPUs, B...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did AI chip startup d-Matrix announce about integrating Nvidia’s NVLink Fusion into its next-generation Raptor inference processors, ho. Article summary: d-Matrix announced a multiyear roadmap to integrate Nvidia’s NVLink Fusion into its next-generation Raptor inference XPUs, making its specialized decode hardware deployable as part of Nvidia-based AI factories rather tha. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
d-Matrix is taking a route into large-scale AI inference that does not require customers to build a wholly separate accelerator cluster. The startup said it will integrate Nvidia’s NVLink Fusion into its next-generation Raptor inference XPUs, connecting the chips to Nvidia’s MGX rack architecture and broader AI infrastructure platform. The companies describe the effort as a multiyear product roadmap, with initial Raptor XPUs integrated into MGX racks expected in the fourth quarter of 2027. 1
8
NVLink Fusion is Nvidia’s high-bandwidth, low-latency interconnect technology and IP for bringing custom CPUs and XPUs into Nvidia AI infrastructure. For d-Matrix, that means Raptor can connect to NVLink scale-up networking, Spectrum-X scale-out networking and MGX rack designs rather than being deployed only in a bespoke system. 1
14
The planned rack design includes Nvidia Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs and Spectrum-X Ethernet. d-Matrix says this gives its systems a validated path through the surrounding rack, networking, power, cooling and supply-chain work that custom accelerator vendors otherwise must assemble themselves. 1
8
In practical terms, the announcement is about system integration, not merely attaching another chip to a server. A data-center operator could use a common Nvidia-oriented infrastructure approach across GPUs, CPUs and d-Matrix XPUs, instead of engineering dedicated racks for every processor type. 1
The proposed architecture uses disaggregated inference: Nvidia GPUs handle the compute-intensive prefill stage, while d-Matrix’s specialized hardware is aimed at the token-by-token decode stage. Prefill processes an input prompt; decode produces the continuing stream of output tokens.
That division is meant to let each processor address the part of inference for which it is intended. d-Matrix has previously said Parasail would deploy its current Corsair accelerators alongside Nvidia Hopper and Blackwell systems using this same broad prefill-and-decode approach. 22
The important caveat is that the performance benefits remain workload-dependent. Parasail’s claim of up to 10x faster, more cost-efficient services is a company and partner claim for selected workloads, not an independently validated, universal comparison between d-Matrix hardware and GPUs. 22
d-Matrix has said its first Raptor rack will follow Nvidia’s MGX reference architecture. Reporting on the roadmap describes an NVL144-style design intended to connect up to 144 Raptor accelerators in a single NVLink fabric; those specifications remain future plans rather than shipping-product guarantees. 2
The concrete availability milestone supplied by d-Matrix is more limited: initial Raptor XPUs integrated into an Nvidia MGX rack are expected in Q4 2027. 8
Because Raptor is not yet a shipping MGX product, customers should treat performance, capacity and deployment details as roadmap targets until final systems and independently testable benchmarks are available.
At first glance, enabling a third-party inference processor might appear to dilute Nvidia’s role. The strategic logic is the reverse: Nvidia can make its AI platform the preferred foundation even when the compute chip is not an Nvidia GPU.
NVLink Fusion keeps high-value platform layers—including the scale-up fabric, networking, rack architecture and associated infrastructure—inside Nvidia’s ecosystem. For d-Matrix and its customers, that may reduce integration risk and speed deployment. For Nvidia, it makes a fully separate, non-Nvidia cluster less necessary. 1
14
That is the meaning behind Nvidia’s description of its platform as “vertically integrated and horizontally open.”
It does not mean every accelerator is interchangeable or performance-equivalent. “Fungible” in this context is best understood as an operational goal: operators can choose the compute resource that suits a workload while using a common infrastructure foundation.
d-Matrix’s plan is complementary to Nvidia GPUs in the proposed prefill/decode split, but it also occupies the same broad specialized-inference territory as Nvidia’s planned Groq 3 LPU-based systems. Nvidia has presented Groq 3 LPX as a purpose-built decode solution that can pair with Vera Rubin NVL72 infrastructure. 24
That creates a dual relationship. Raptor can complement Nvidia GPUs and consume Nvidia platform infrastructure, while also potentially competing with Nvidia’s own specialized decode offering. The available information does not support a meaningful head-to-head conclusion on latency, throughput, power use, price or availability.
d-Matrix said its November 2025 Series C raised $275 million at a $2 billion valuation, bringing its total funding to $450 million. 27 Its current inference platform is Corsair; Raptor is the next-generation XPU named in the NVLink Fusion roadmap.
Corsair is already the basis for d-Matrix’s partnership with Parasail, which said it would deploy the accelerators alongside Nvidia infrastructure for selected inference workloads. 22 Raptor’s significance is that the company is now planning a more direct route into Nvidia’s rack-scale environment.
Neither company disclosed the deal’s financial terms. As a result, the public record does not establish licensing royalties, minimum-purchase commitments, engineering fees, revenue sharing or any preferential commercial conditions.
The announcement is therefore best read as a significant infrastructure roadmap: d-Matrix gains a path to deploy specialized decode hardware in Nvidia-based AI factories, while Nvidia broadens the range of compute chips that can rely on its interconnect and rack ecosystem. Whether that translates into superior real-world inference economics will depend on final hardware, software integration, pricing and independently reproducible workload results. 1
8
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
d Matrix plans to use Nvidia NVLink Fusion to place its next generation Raptor inference XPUs inside Nvidia MGX based AI infrastructure, pairing GPU prefill with specialized token decoding.
d Matrix plans to use Nvidia NVLink Fusion to place its next generation Raptor inference XPUs inside Nvidia MGX based AI infrastructure, pairing GPU prefill with specialized token decoding. The arrangement gives d Matrix access to Nvidia’s NVLink and Spectrum X networking, plus rack components including Vera CPUs, BlueField 4 DPUs and ConnectX 9 SuperNICs.
It is open to third party compute silicon, but strategically keeps the rack, interconnect and networking layers within Nvidia’s platform.