AI inference chipmaker d-Matrix will integrate its next-generation Raptor XPUs into NVIDIA’s rack-scale infrastructure using NVLink Fusion, giving the startup a route to deploy its specialized accelerators within systems built around NVIDIA’s networking, CPUs and data-center architecture.
The companies disclosed the collaboration on September 10. d-Matrix said the multi-year product roadmap will start with Raptor connecting directly into the NVIDIA MGX rack architecture. Initial availability of Raptor XPUs integrated into the MGX rack is expected in the fourth quarter of 2027.
The move is significant for AI infrastructure operators because it is designed to support heterogeneous racks rather than forcing specialized inference silicon into an entirely separate server architecture. d-Matrix is targeting AI labs, hyperscalers and neocloud operators running latency-sensitive inference services such as coding assistants, real-time chatbots and voice agents.
Raptor moves into NVIDIA’s rack architecture
Under the collaboration, Raptor will be designed to connect through NVIDIA NVLink Fusion, which extends NVIDIA’s scale-up interconnect architecture to custom processors. The planned system will also incorporate NVIDIA Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs and Spectrum-X Ethernet networking.
NVIDIA separately confirmed the integration, saying NVLink Fusion will connect Raptor to its scale-up networking and broader AI infrastructure platform. For d-Matrix, that provides access to an established rack design, networking stack, power and cooling infrastructure and supply-chain ecosystem rather than requiring the company to build every element of a rack-scale deployment independently.
d-Matrix is also working with Astera Labs on custom connectivity intended to provide high-throughput data movement across the system. The planned rack will use modular, cable-free trays based on NVIDIA’s MGX ecosystem.
| Component | Planned role |
|---|---|
| d-Matrix Raptor | Purpose-built XPU for AI inference workloads |
| NVIDIA NVLink Fusion | High-bandwidth, low-latency scale-up connectivity |
| NVIDIA MGX | Rack-scale architecture and deployment ecosystem |
| NVIDIA Vera | CPU platform in the planned rack architecture |
| BlueField-4, ConnectX-9 and Spectrum-X | Data processing and scale-out networking infrastructure |
A specialized architecture for inference
Raptor is the planned successor to d-Matrix’s Corsair inference platform. The company says the processor extends its memory-centric architecture through a 3D stacking approach that combines a DRAM memory chip with an SRAM compute chip in a single package.
The architecture is aimed specifically at inference, where already-trained models generate responses and where memory movement, latency and token-generation speed can become major infrastructure constraints. d-Matrix said Raptor is being evaluated by AI hyperscalers and frontier AI labs, although it did not identify those prospective customers.
The company expects Raptor to tape out before the end of 2026. That means the announced rack is still a future product rather than currently deployable infrastructure, and d-Matrix has not disclosed pricing or financial terms for its NVIDIA collaboration.
Mixed compute inside the AI factory
A central part of the plan is the ability to divide inference work between different processor types. d-Matrix said operators could use NVIDIA GPU systems for compute-intensive portions of an inference workload while assigning latency-sensitive token generation to Raptor.
For AI coding workloads, for example, the company describes a configuration in which GPUs handle the compute-heavy prefill phase while its inference XPUs process the decode phase. This is a proposed deployment model rather than a disclosed customer deployment.
NVIDIA’s adoption of partners through NVLink Fusion also broadens the infrastructure available to companies building processors that compete with GPUs for particular workloads. Instead of requiring a wholly separate rack environment, specialized accelerators can be designed around NVIDIA’s interconnect, networking and physical infrastructure.
What happens next
The first technical milestone is Raptor’s expected tape-out before the end of 2026. The companies will then have to move through silicon production, system integration and validation before the planned Q4 2027 availability window.
For data-center and AI infrastructure operators, the practical question will be whether specialized inference processors can deliver enough performance, latency and operating-cost advantages to justify heterogeneous deployments while retaining the operational benefits of a common rack architecture.
The Q4 2027 availability date and pre-2027 tape-out target are company projections. No production deployment of the Raptor-based NVIDIA MGX system has yet been announced.







