d Matrix plans to connect its next generation Raptor inference XPUs to NVIDIA’s NVLink Fusion fabric and MGX rack architecture, with initial integrated rack availability targeted for Q4 2027. The proposed MGX design includes NVIDIA Vera CPUs, NVLink switches, BlueField 4 DPUs, ConnectX 9 SuperNICs and Spectrum X Eth...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did AI chip startup d-Matrix announce about integrating Nvidia’s NVLink Fusion technology into its next-generation Raptor inference pro. Article summary: d-Matrix announced a multiyear roadmap partnership to design its next-generation Raptor inference XPUs for NVIDIA NVLink Fusion, beginning with integration into NVIDIA’s MGX rack reference architecture. The aim is a rack. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
d-Matrix has announced a multiyear product-roadmap collaboration with NVIDIA to integrate its upcoming Raptor inference XPUs with NVIDIA NVLink Fusion. The first planned implementation is a Raptor-based system built on NVIDIA’s MGX rack reference architecture, aimed at large-scale, ultra-low-latency AI inference. Initial availability of Raptor in an MGX rack is targeted for the fourth quarter of 2027. 1
2
11
NVLink Fusion is NVIDIA’s platform for linking third-party CPUs and accelerators into its NVLink scale-up networking and broader AI infrastructure. Under the partnership, d-Matrix plans to connect Raptor XPUs through NVLink while using Spectrum-X Ethernet for scale-out networking. 1
In practical terms, this gives d-Matrix a path to build its custom inference silicon into an established NVIDIA rack design rather than create and qualify an entirely separate rack, networking and cooling platform. NVIDIA and d-Matrix describe that as a faster, lower-risk route from custom silicon to large-scale deployment. 1
2
The planned Raptor system is intended to use NVIDIA’s latest MGX reference architecture. d-Matrix says that design will include:
MGX is designed as a modular rack architecture, and NVIDIA says the integration would place Raptor in a broadly deployed liquid-cooled AI infrastructure. 1
The announcement supports interoperability with NVIDIA’s rack-scale platform, but the cited company announcements do not establish a finalized, shipping Raptor configuration specifically branded as an “NVL144” system. Any 144-accelerator or memory-bandwidth figures reported ahead of launch should therefore be understood as roadmap projections, not deployed or independently benchmarked specifications.
The partnership centers on heterogeneous inference: using different processors for different parts of serving a generative AI model.
A common split is to use GPUs for prefill, the compute-intensive stage in which a model processes an input prompt, and a specialized accelerator for decode, the latency-sensitive stage that generates tokens one at a time. d-Matrix positions Raptor for that decode-oriented role. 18
19
Raptor follows the company’s production Corsair platform and extends d-Matrix’s memory-centric design. d-Matrix says Raptor uses a 3D DRAM-stacking approach that brings a DRAM memory die and an SRAM compute die together in one package. The company’s premise is that tighter coupling of memory and compute can better serve low-latency inference workloads. 2
That positioning matters: Raptor is not presented as a wholesale GPU replacement. It is intended to become a specialized component in an NVIDIA-based inference cluster, with each processor assigned work suited to its architecture.
For NVIDIA, supporting a third-party XPU through NVLink Fusion expands the range of hardware that can operate within its networking, rack and infrastructure ecosystem. For customers, the immediate appeal is integration: a specialized inference accelerator can be deployed with NVIDIA’s validated MGX design, NVLink connectivity and Spectrum-X networking rather than assembled as a standalone environment. 1
That makes the arrangement a platform strategy as much as a chip partnership. NVIDIA can keep infrastructure buyers on its rack-scale foundation while allowing specialized silicon to address a distinct inference bottleneck. The commercial outcome will depend on real-world availability, software support, cost and workload performance once Raptor ships.
d-Matrix expects Raptor to tape out by the end of 2026 and targets initial MGX rack availability in Q4 2027. These are forward-looking roadmap dates and may change. 9
11
The company’s current Corsair XPU platform is already in production. In a separate deployment announcement, Parasail said it would pair Corsair accelerators with NVIDIA Hopper and Blackwell GPUs, using GPUs for prefill and Corsair for decode. Parasail and d-Matrix claimed up to 10x faster token generation; that is a company-reported performance claim, not an independent benchmark for Raptor. 19
D-Matrix also reported a $275 million Series C financing in November 2025, valuing the company at $2 billion and bringing its total funding to $450 million. 16
The important development is not simply that d-Matrix has adopted an NVIDIA interconnect. It is that Raptor is being designed to enter NVIDIA’s MGX rack ecosystem as a purpose-built inference component. If the roadmap is delivered, operators could combine NVIDIA CPUs, networking and GPUs with a memory-centric XPU intended for the decode phase of AI serving—without building a separate rack-scale platform from scratch. 1
2
For now, the partnership is a roadmap with a 2027 availability target. The eventual value proposition will need to be tested against shipping hardware, real inference workloads and independently measured latency, throughput and total cost of ownership.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
d Matrix plans to connect its next generation Raptor inference XPUs to NVIDIA’s NVLink Fusion fabric and MGX rack architecture, with initial integrated rack availability targeted for Q4 2027.
d Matrix plans to connect its next generation Raptor inference XPUs to NVIDIA’s NVLink Fusion fabric and MGX rack architecture, with initial integrated rack availability targeted for Q4 2027. The proposed MGX design includes NVIDIA Vera CPUs, NVLink switches, BlueField 4 DPUs, ConnectX 9 SuperNICs and Spectrum X Ethernet, giving d Matrix a route into NVIDIA’s liquid cooled rack ecosystem.
Raptor remains a future product, so its performance and rack scale specifications should be treated as company roadmap targets rather than independently verified results.