AI inference chipmaker d-Matrix has agreed to integrate its next-generation Raptor accelerators with NVIDIA NVLink Fusion and the NVIDIA MGX rack architecture, giving the startup a route to deploy its specialized processors inside NVIDIA-based data center infrastructure.
The companies announced the multi-year collaboration on September 10. Initial availability of Raptor XPUs integrated into MGX racks is expected in the fourth quarter of 2027, while d-Matrix plans to complete the chip’s final design, or tape-out, before the end of 2026.
The agreement targets AI laboratories, hyperscalers and neocloud operators running latency-sensitive inference services such as coding assistants, interactive chatbots and voice agents. Financial terms were not disclosed.
Deployment status
Initial availability: Q4 2027
Raptor has not entered commercial deployment. The chip is expected to tape out before the end of 2026, making the availability date a forward-looking product target.
Raptor moves into NVIDIA’s rack ecosystem
Under the d-Matrix product announcement, Raptor will connect to NVIDIA’s scale-up infrastructure through NVLink Fusion and use the modular MGX rack design. The planned systems will incorporate NVIDIA Vera CPUs, NVLink switches, BlueField-4 data-processing units, ConnectX-9 SuperNICs and Spectrum-X Ethernet networking.
d-Matrix is also working with connectivity supplier Astera Labs on custom components intended to maintain high-throughput data movement across the system. The rack design will use modular, cable-free trays from the MGX ecosystem, reducing the amount of infrastructure d-Matrix must design and validate independently.
NVIDIA describes NVLink Fusion as a platform for connecting third-party processors to its scale-up networking, rack architecture and wider AI infrastructure stack. NVIDIA claims its sixth-generation NVLink can provide 3TB per second of all-to-all bandwidth per XPU, with lower latency and higher packet rates than standard Ethernet connections.
Those figures describe the NVLink Fusion platform rather than measured performance of a commercial Raptor system. d-Matrix has not yet published independent benchmarks for the completed rack configuration.
A memory-focused processor for inference
Raptor is the planned successor to d-Matrix’s Corsair inference platform. Its architecture places a DRAM memory chip and an SRAM-based compute chip together in a vertically stacked package, an approach designed to reduce the data-movement bottleneck encountered when AI models generate tokens.
The proposed rack architecture can also divide inference workloads between processor types. NVIDIA Vera Rubin systems would handle the compute-intensive prefill stage, when a model processes an incoming prompt, while Raptor XPUs would handle the latency-sensitive decode stage that generates the response.
This arrangement makes the collaboration more than a conventional accelerator supply agreement. d-Matrix is adopting NVIDIA’s interconnect, rack, networking and supply-chain framework while retaining its own processor architecture for a specific part of the AI workload.
The Register reported that d-Matrix expects configurations with as many as 144 Raptor accelerators connected through a single all-to-all NVLink fabric by the end of 2027. That configuration remains planned rather than generally available.
Why the agreement matters
For data center operators, a common rack design could make it easier to deploy specialized inference processors alongside NVIDIA GPU systems without constructing a separate power, cooling and networking architecture for each accelerator family.
The agreement also advances NVIDIA’s strategy of keeping its infrastructure platform central even when customers select processors from another supplier. NVLink Fusion gives custom-chip developers access to NVIDIA’s rack-scale ecosystem, while NVIDIA continues to provide CPUs, switches, network adapters, DPUs and software around those accelerators.
For d-Matrix, the immediate consequence is access to a more established deployment path. The larger test will come after tape-out: producing working silicon, validating the complete liquid-cooled rack and delivering commercial systems on the fourth-quarter 2027 schedule.







