d Matrix plans to use Nvidia NVLink Fusion to put its Raptor inference XPUs directly into Nvidia MGX rack scale systems rather than deploy them as isolated PCIe accelerators. The design brings together Raptor XPUs, Nvidia Vera CPUs, NVLink switches, BlueField 4 DPUs, ConnectX 9 SuperNICs and Spectrum X Ethernet—lett...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: How does d-Matrix’s adoption of Nvidia’s NVLink Fusion for its next-generation Raptor inference processors enable the startup to plug its me. Article summary: d-Matrix is using NVLink Fusion to make Raptor a native, rack-scale participant in Nvidia’s AI-factory design rather than a separate PCIe inference card. That gives customers a memory-centric accelerator for latency-sens. Topic tags: general, news, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts w
d-Matrix’s NVLink Fusion agreement with Nvidia is fundamentally an integration play: Raptor is intended to become a component inside Nvidia’s rack-scale AI infrastructure, rather than an inference card connected only through a conventional server expansion path. The goal is to combine a specialized, memory-centric inference processor with the system design, interconnect and networking already used in Nvidia AI factories. 3
10
NVLink Fusion is Nvidia’s platform for bringing partner silicon into Nvidia-based AI infrastructure. For d-Matrix, that means connecting Raptor to Nvidia’s NVLink scale-up networking, MGX rack architecture and Spectrum-X scale-out networking. Nvidia describes this as a faster, lower-risk route from custom silicon to large-scale deployment because the partner can build on a validated rack ecosystem rather than assemble every layer of the system independently. 3
In practical terms, Raptor is planned as part of an MGX reference-design rack. That is the difference between selling a standalone accelerator and participating in a rack-scale system with a shared high-bandwidth fabric and a defined set of infrastructure components. 3
10
The planned Raptor-based design includes the following Nvidia platform elements:
d-Matrix is also working with Astera Labs on custom connectivity intended to support high-throughput data flow in the system. 10
This split is strategically important. d-Matrix supplies the inference XPU, while Nvidia supplies much of the surrounding platform: the reference architecture, scale-up connectivity, networking components and broader AI infrastructure ecosystem. 3
10
Raptor is positioned for inference workloads where token-generation latency matters. A memory-centric accelerator can be particularly relevant for workloads that repeatedly access model weights during generation, while Nvidia’s platform provides the communications and deployment foundation needed to operate at rack scale.
The collaboration does not mean Raptor replaces Nvidia GPUs across all AI work. Instead, it creates a path for heterogeneous systems in which specialized inference silicon can coexist with Nvidia infrastructure. Nvidia’s NVLink Fusion strategy is explicitly aimed at semi-custom AI infrastructure that combines partner silicon with Nvidia rack-scale systems and networking. 26
d-Matrix has said it expects to offer systems with up to 144 Raptor accelerators connected by a single all-to-all NVLink fabric. The company’s reported rack-level targets are 2.3 TB of 3D-stacked DRAM and 7.2 PB/s of aggregate memory bandwidth, with the stated aim of supporting models exceeding four trillion parameters at 4-bit precision. 2
4
Those figures should be treated as forward-looking product targets. Public reporting also contains differing per-accelerator memory and bandwidth descriptions, so the final per-device configuration and usable model capacity will need confirmation when the product is closer to launch. Runtime memory, model metadata, quantization format and system software all affect deployable model size.
Raptor is described as a design combining a 4-nanometer compute die with custom DRAM. d-Matrix expects tape-out by the end of 2026, with initial rack-integrated systems expected in the fourth quarter of 2027. 10
14
That makes this a multi-year product-roadmap collaboration, not a currently shipping joint rack product. 12
Opening its rack-scale platform to partner chips lets Nvidia address customers that want specialized silicon without requiring them to abandon Nvidia’s infrastructure stack. Even when the inference processor is supplied by d-Matrix, the proposed system can still use Nvidia’s NVLink fabric, CPUs, switches, DPUs, SuperNICs, Ethernet networking and MGX architecture. 3
10
That conclusion is an architectural implication, not a disclosed statement about deal economics or revenue sharing. The companies did not disclose financial terms. 1
Nvidia introduced NVLink Fusion to enable semi-custom AI infrastructure with partners. Its initial public partner list included MediaTek, Marvell, Alchip Technologies, Astera Labs, Synopsys and Cadence; Nvidia also said Fujitsu and Qualcomm CPUs could be integrated with Nvidia GPUs through the platform. 26
d-Matrix extends that approach to inference-focused XPUs. The wider significance is less about one chip than about a more modular rack model: customers can potentially combine custom compute silicon with Nvidia’s interconnect and network infrastructure rather than choose between an all-Nvidia system and a fully separate platform.
d-Matrix launched its Corsair inference accelerator in November 2024. In November 2025, the company said it had raised a $275 million Series C at a $2 billion valuation, bringing total funding to $450 million; Microsoft’s M12 was among the investors. 28
39
The Raptor partnership is therefore the company’s attempt to move from a discrete inference-accelerator approach toward a rack-scale deployment model. Whether its promised latency, capacity and bandwidth advantages translate into production deployments will depend on the final silicon, software integration and real-world benchmarks when the planned systems arrive in 2027.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
d Matrix plans to use Nvidia NVLink Fusion to put its Raptor inference XPUs directly into Nvidia MGX rack scale systems rather than deploy them as isolated PCIe accelerators.
d Matrix plans to use Nvidia NVLink Fusion to put its Raptor inference XPUs directly into Nvidia MGX rack scale systems rather than deploy them as isolated PCIe accelerators. The design brings together Raptor XPUs, Nvidia Vera CPUs, NVLink switches, BlueField 4 DPUs, ConnectX 9 SuperNICs and Spectrum X Ethernet—letting d Matrix specialize in inference while using Nvidia’s surrounding rack...
d Matrix says a 144 XPU rack is intended to provide 2.3 TB of 3D stacked DRAM and 7.2 PB/s of aggregate memory bandwidth; these are company stated targets.