d Matrix is using NVLink Fusion to make Raptor a native, rack scale participant in Nvidia’s AI factory design rather than a separate PCIe inference card. That gives customers a memory centric accelerator for latency sensitive decoding while preserving Nvidia’s scale up fabric, networking, control plane hardware, and...
Published byImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: How does d Matrix’s adoption of Nvidia’s NVLink Fusion for its next generation Raptor inference processors enable the startup to plug its me. Article summary: d Matrix is using NVLink Fusion to make Raptor a native, rack scale participant in Nvidia’s AI factory design rather than a separate PCIe inference card.. Topic tags: general web, llm, ai, code, benchmarks. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clickbait thumbnails, icons, and tiny thumbnail layou
d-Matrix is using NVLink Fusion to make Raptor a native, rack-scale participant in Nvidia’s AI-factory design rather than a separate PCIe inference card. That gives customers a memory-centric accelerator for latency-sensitive decoding while preserving Nvidia’s scale-up fabric, networking, control-plane hardware, and rack integration. 3
5
How it plugs in: Raptor will implement Nvidia’s NVLink Fusion interface and fit Nvidia’s liquid-cooled NVL144/MGX reference architecture. NVLink supplies the high-bandwidth, low-latency scale-up fabric between devices; Spectrum-X supplies scale-out networking beyond the rack. 3
5
Nvidia platform components: A Raptor-based rack can use Nvidia’s Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X Ethernet networking. Thus, d-Matrix contributes the XPU while Nvidia still supplies much of the compute-system plumbing, I/O, networking, management, and reference-rack design. 5
Inference rationale: GPUs remain useful for model prefill and general compute, while d-Matrix targets token-by-token decode, where repeatedly reading model weights makes memory capacity and bandwidth especially important. The intended result is lower token latency and a complementary GPU–XPU inference pipeline, not necessarily wholesale GPU substitution. 3
10
Capacity and bandwidth claim: d-Matrix’s stated rack target is 144 Raptor XPUs, totaling 2.3 TB of 3D-stacked DRAM and 7.2 PB/s aggregate memory bandwidth—equivalent to 16 GB and 50 TB/s per XPU. In principle, 2.3 TB is sufficient to hold more than four trillion 4-bit parameters before allowing for metadata, runtime state, and other overhead; actual deployable model size will therefore depend on the model format and system software.
Silicon schedule: Raptor is described as a 4-nm compute die combined with custom DRAM. d-Matrix expects tape-out by the end of 2026 and initial rack-integrated systems in Q4 2027. 10
4
Why Nvidia benefits: Strategically, NVLink Fusion reduces the chance that customers leave Nvidia’s ecosystem when they choose specialized third-party accelerators. Nvidia can retain the rack architecture and sell the fabric, switches, CPUs, DPUs, SuperNICs, Ethernet, and software-adjacent infrastructure—even where the primary inference silicon is not an Nvidia GPU. This is an inference from the architecture and commercial model, not a disclosed statement of deal economics. 3
5
Ecosystem significance: d-Matrix joins the broader NVLink Fusion effort, alongside partners publicly associated with it including AWS, Arm, Intel, Fujitsu, MediaTek, and Marvell. The strategy is to make NVLink a platform interface for heterogeneous AI-factory hardware, analogous to expanding Nvidia’s moat from accelerators into the surrounding system fabric. 3
Company context: d-Matrix launched its Corsair inference accelerator in November 2024. Its later funding history includes Microsoft’s M12 participation; reporting says a $275 million Series C valued the company at $2 billion, while some accounts describe its cumulative financing as about $450 million—these figures should not be conflated as one $450 million round. 2
6 ParaSail has reported up to a tenfold inference-speed improvement from combining Corsair with Nvidia GPUs, but that is a company/customer performance claim rather than an independent benchmark.
12
Economics: The companies did not disclose the collaboration’s financial terms. 1
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
d Matrix is using NVLink Fusion to make Raptor a native, rack scale participant in Nvidia’s AI factory design rather than a separate PCIe inference card.
d Matrix is using NVLink Fusion to make Raptor a native, rack scale participant in Nvidia’s AI factory design rather than a separate PCIe inference card. That gives customers a memory centric accelerator for latency sensitive decoding while preserving Nvidia’s scale up fabric, networking, control plane hardware, and rack integration.
[3][5] How it plugs in: Raptor will implement Nvidia’s NVLink Fusion interface and fit Nvidia’s liquid cooled NVL144/MGX reference architecture.