Nvidia And d-Matrix Announce Roadmap For Fast Inference
Nvidia has done something unexpected: the AI leader is partnering with its (previous?) competitor d-Matrix to build an alternative disaggregated inference solution. Why not wait for the Groq LPU3, which essentially solves the same problem using fast SRAN for disaggregated inference processing? I have a few ideas I will delve into in this article. ( Nvidia is a client of Cambrian-AI Research, and d-Matrix has been a client in the past .)
When I first met d-Matrix founder and CEO Sid Sheth in 2022, I asked how his company will compete with Nvidia. Sid said he didn’t intend to; he wanted to focus where Nvidia wasn’t and be complementary to the then-emerging AI leader. And now he is doing just that, pairing his company’s SRAM-based AI technology for decode with Nvidia GPUs for encoding, or tokenizing the input query for analysis against the trained neural network.
But Nvidia announced its acquisition of startup Groq’s technology and leadership for some $20B last December, and explained its disintegration strategy at GTC in April 2026. GPUs would do the compute-intensive prefill and attention processing, while the next gen Groq LPU3 would do the memory-intensive decoding to produce output tokens, or answers. As I reported at the time, it sounded great to me .
But recently, the NeoCloud player ParaSail announced that it was standing up Nvidia GPUs with startup d-Matrix’ 1st generation Corsair PCIe cards, mimicking the approach sooner and probably at a lower cost, albeit without the fast NVLink used to connect the processors in the Nvidia LPX rack. Now that approach has been essentially endorsed by Nvidia, which joined d-Matrix in announcing a multi-year roadmap for integrating the d-Matrix 2nd generation platform, code-named Raptor which will include NVLink Fusion.
What Does This Mean for Groq?
Nvidia’s pairing of GPU infrastructure with d-Matrix for decode is best understood as a customer-retention and time-to-value strategy—not a retreat from owning inference or supporting Groq. It gives customers a deployable answer to low-latency token generation now, and can help customers extract more value from the Nvidia GPUs already acquired. Data centers can deploy the d-Matrix Corsair accelerator today alongside existing or new GPUs, and deploy Groq or the next generation Raptor over NVLink (planned for Q4 2027).
Parasail operates Nvidia Hopper and Blackwell fleets already. While new data-center capacity can take years to approve and construct, adding Corsair to those GPUs lets the provider extract more interactive-inference value from deployed Nvidia hardware rather than waiting for a wholesale platform refresh.
That is a highly Nvidia-like move. Instead of requiring customers to wait for the perfect all-Nvidia decode solution, Nvidia enables a mixed architecture that can improve user-visible latency and token economics today. Parasail claims they duo can run up to 10x faster, more cost-efficient inference services, though those are vendor-stated targets rather than independently published production benchmarks.
The partnership protects Nvidia’s most valuable asset: customer platform control. A GPU-only posture risks forcing customers with latency-critical coding agents, consumer chat, voice, and real-time enterprise applications to build a separate inference stack around a rival ASIC. By accepting d-Matrix as a complementary decode tier, Nvidia keeps the large-model serving, prefill, model ecosystem, networking, systems integration, and much of the capital spend anchored to Nvidia infrastructure.
Customers increasingly want the best economics per phase of the request, not ideological purity around a single accelerator. Nvidia can therefore position itself as the orchestrator of the broader AI factory: its GPUs handle the heavy work, while external specialty silicon fills a niche bottleneck when justified by the service-level objective.
The d-Matrix relationship also gives Nvidia time and optionality before its own low-latency inference product is ready at scale. With LPX, Nvidia can offer a more vertically integrated decode tier over NVLink with tighter software, networking, support, and commercial integration. But Nvidia does not need to force that transition prematurely. It can learn from real d-Matrix deployments: what fraction of demand needs low latency, which models and context lengths benefit, what routing software is required, and how much customers will pay for lower time-to-first-token and faster inter-token latency.
But customers could also opt for an Nvidia plus d-Matrix future. d-Matrix has already announced a new memory technology which would stack the next-gen AI processor on top of a memory chip, extending the memory and model size that can be run on a single accelerator. Nvidia has not yet announced a similar capability for Groq 3.
Nvidia is selling outcomes before architecture. For a customer, the relevant question is not whether every token comes from an Nvidia chip; it is whether the service meets latency, throughput, availability, energy, and cost targets while preserving operational simplicity. The d-Matrix pairing answers that question now and reduces the incentive for customers to defect from Nvidia-centered deployments. And these customers, once successful with d-Matrix + Nvidia become a ripe target for the LPU3.
The longer-term risk for Nvidia is that decode becomes a large enough portion of inference spending that specialist silicon gains durable account control. Nvidia’s response is pragmatic: partner while the category is forming, remain indispensable in the high-value portions of the stack, and be prepared to substitute its own LPU 3/LPX solution when the customer and workload justify tighter vertical integration. Or it could decide to offer a choice, with the Groq-based solution offering a higher value and the d-Matrix decoder possibly offering lower prices. (Pricing has not been disclosed.)
This is all conjecture, but you can begin to see that Nvidia is playing the long game; either way, they win: it’s still Nvidia Inside.
Disclosures : This article expresses the author's opinions and
should not be taken as advice to purchase from or invest in the companies mentioned. Cambrian-AI Research is fortunate to have many, if not most, semiconductor firms as our clients, including Cadence Design, Cerebras, IBM, Intel, NVIDIA, Qualcomm Technologies, Synopsys, and scores of investment clients. For more information, please visit our website at https://cambrian-AI.com .