Summary: d-Matrix plans to integrate its upcoming Raptor accelerator into NVIDIA MGX racks through NVLink Fusion. Vera Rubin GPUs would process compute-heavy prompts, while Raptor handles sequential token generation. This broadens silicon choice but ties alternative accelerators more closely to NVIDIA's surrounding infrastructure.
Split work inside one rack
The proposed architecture separates two phases of AI inference. During prefill, a model processes the full input context largely in parallel; d-Matrix assigns that work to NVIDIA Vera Rubin. During decode, tokens are generated sequentially, making memory access and response latency more important. Raptor is intended to handle this phase.
According to d-Matrix, Raptor combines a DRAM memory chip and an SRAM compute chip in a stacked package. The company expects the design to tape out before the end of 2026 and targets initial Raptor XPUs in MGX racks for the fourth quarter of 2027. Astera Labs is developing custom connectivity, while the rack design also includes Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 adapters and Spectrum-X Ethernet.
NVLink Fusion opens NVIDIA's interconnect technology and rack reference architecture to third-party CPUs and accelerators. NVIDIA advertises up to 3.6 TB/s per XPU for NVLink 6 in a 72-accelerator domain. The d-Matrix announcement does not specify which NVLink generation or bandwidth its Raptor configuration will implement.
What remains unproven
d-Matrix promises very low latency, better energy efficiency and attractive token economics. No independent measurements, model sizes, response latencies, power figures or total-cost comparisons are available. Splitting prefill and decode is not a free gain either: orchestration and data transfers between unlike accelerators can reduce the benefit. Raptor has not yet taped out, and announced availability is more than a year away.
Pandorex Analysis
For d-Matrix, access to MGX racks and NVIDIA's supply chain is likely more valuable than one interface alone. Data centres could deploy alternative inference silicon without redesigning power, cooling, networking and management from scratch. NVIDIA broadens choice at the chip layer while retaining control over the wider rack and networking model.
Pandorex assessment: The collaboration could make Raptor easier to deploy, but it does not yet prove an economic advantage over GPU-only systems. Independent tests with real models and the cost of moving data between prefill and decode will be decisive.
