d-Matrix to Integrate Raptor AI Inference Chips Into Nvidia MGX Racks With NVLink Fusion

Dedicated Servers and GPU

AI inference chipmaker d-Matrix will integrate its next-generation Raptor XPUs into Nvidia’s MGX rack-scale infrastructure using NVLink Fusion, giving the startup a route to deploy its specialized accelerators within an established Nvidia data center architecture.

The Santa Clara, California-based company announced the collaboration on September 10. Its first Raptor systems built around Nvidia’s MGX architecture are expected to become available in the fourth quarter of 2027, while d-Matrix expects the Raptor processor to tape out before the end of 2026.

Deployment timeline

Initial Raptor XPUs integrated into Nvidia MGX racks are targeted for Q4 2027.

Raptor moves into Nvidia’s rack architecture

Under the multi-year product roadmap, d-Matrix plans to integrate Raptor with Nvidia Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs and Spectrum-X Ethernet networking. Astera Labs is also working with d-Matrix on connectivity solutions intended to provide high-throughput data movement across the system.

Raptor is d-Matrix’s next-generation accelerator for AI inference, following the company’s Corsair platform. The new processor extends d-Matrix’s memory-centric architecture with a 3D stacking approach that combines a DRAM memory chip and an SRAM compute chip in a single package.

Rather than creating an independent rack-scale platform around the processor, d-Matrix is designing Raptor for Nvidia’s existing MGX ecosystem. NVLink Fusion provides the high-bandwidth, low-latency scale-up interconnect needed to connect the XPUs, while MGX supplies a common rack architecture and an established ecosystem for networking, power, cooling and system deployment.

Why the infrastructure integration matters

The agreement is significant beyond the Raptor processor itself. Bringing a custom AI accelerator into production at data center scale requires more than silicon: operators need validated networking, rack designs, power delivery, cooling and a supply chain capable of supporting large deployments.

Using Nvidia’s infrastructure allows d-Matrix to concentrate more of its engineering effort on its inference processor while relying on an architecture already designed for large AI systems. Nvidia said the common MGX design can allow data centers to support GPUs, CPUs and third-party XPUs without requiring an entirely separate rack architecture for each processor type.

For infrastructure operators, that creates the possibility of deploying specialized inference hardware alongside Nvidia GPU systems while retaining common elements of the surrounding rack and networking environment.

Nvidia said Raptor-based racks can also operate alongside GPU systems such as Vera Rubin NVL72 for disaggregated inference. That model can assign different stages or types of AI inference workloads to hardware optimized for them rather than requiring every part of the workload to run on the same accelerator architecture.

Inference becomes a target for specialized silicon

d-Matrix is positioning Raptor specifically for latency-sensitive inference workloads, including conversational AI, coding assistants and voice agents. These applications place a premium on the speed and efficiency with which deployed models generate tokens for users.

The partnership also shows how Nvidia is extending its data center footprint beyond systems based exclusively on its own accelerators. NVLink Fusion allows third-party CPU and XPU developers to connect custom processors to Nvidia’s rack-scale infrastructure while retaining their own compute architectures.

Reuters reported that the financial terms of the d-Matrix collaboration were not disclosed. The company has previously attracted backing from Microsoft and has been developing dedicated inference hardware as demand shifts from building large AI models toward operating them at scale.

What happens next

The immediate milestone is silicon rather than deployment. d-Matrix says Raptor is expected to tape out before the end of 2026. That will be followed by the work required to validate the processor and its integration into the MGX rack-scale environment.

Initial availability of Raptor XPUs integrated into Nvidia MGX racks is currently planned for Q4 2027. Until then, performance, deployment scale and customer adoption of the final production systems remain prospective rather than established results.

For data center operators, the more consequential development is the architecture behind the announcement: specialized inference silicon is increasingly being designed to fit into existing rack-scale AI infrastructure rather than requiring a completely separate hardware ecosystem. Whether Raptor gains significant deployment will depend on execution over the next year and on how its inference performance translates from the design stage into production systems.

The announced Q4 2027 availability and pre-year-end 2026 Raptor tape-out are forward-looking deployment targets from d-Matrix and should not be read as completed deployments.