Article 789Z3 d-Matrix drinks the Nvidia Kool-Aid with NVLink Fusion and MGX rack designs

d-Matrix drinks the Nvidia Kool-Aid with NVLink Fusion and MGX rack designs

by
from www.theregister.com - Articles on (#789Z3)
Story ImageAI infrastructure startup d-Matrix on Thursday joined the growing list of chipmakers licensing Nvidia's NVLink Fusion interconnect tech and rack-scale reference designs to make its high-performance inference platform more accessible to customers. Under the deal, d-Matrix will integrate support for NVLink Fusion, a high-speed chip-to-chip interconnect that Nvidia began licensing last year, into future chip designs, including its upcoming Raptor accelerators. As we've previously reported, by embracing the tech, d-Matrix sidesteps many of the challenges associated with scaling its chip architecture across large compute clusters. By using NVLink over alternative interconnects and designing its compute blades around the GPU giant's MGX reference designs, its customers can deploy its chips using the same racks and NVSwitch fabrics as Nvidia. By the end of next year, d-Matrix expects to offer systems with up to 144 Raptor accelerators connected by a single all-to-all NVLink fabric. While there's a lot we don't know about Raptor just yet, at the Hot Chips conference last month the company revealed each Raptor card" would feature 32 GB of ultra-fast 3D-stacked DRAM on board capable of delivering 100 TB/s of memory bandwidth - roughly 4.5 times the memory bandwidth of Nvidia's Rubin GPU. By the looks of things, the XPUs that will power d-Matrix's NVL144 racks are about half the size of the Raptor cards shown off at Hot Chips and include around 16 GB of 3D-DRAM and about 50 TB/s of memory bandwidth - still quite respectable by any measure. With 144 of these per rack, d-Matrix is looking at about 2.3 TB of memory capacity - enough for models exceeding four trillion parameters in size at 4-bit precision - and about 7.2 petabytes a second of peak aggregate memory bandwidth. This is achieved by bonding compute logic atop a stack of DRAM. The result is an in-memory compute platform that offers modest capacity while maintaining memory bandwidth closer to that of SRAM than is achievable using HBM. Memory bandwidth, as you may recall, is the biggest bottleneck for AI inference. The faster your memory, the faster the system can spew out tokens. This is exactly why Nvidia dropped $20 billion last year to license Groq's IP and hire away its engineering talent. The chip's SRAM-heavy dataflow architecture was capable of hitting 150 TB/s per chip, but the tradeoff is that SRAM isn't very space-efficient and the chipmaker could only pack 500 MB of it onto a single die. d-Matrix's Raptor promises to deliver a decent fraction of that bandwidth with 64x higher capacity, which means the company can get away with using far fewer chips per model. Where a trillion-parameter model might need more than 2,000 Groq 3 LPUs at 8-bit precision, a single d-Matrix system would only need about 64 (32 at 4-bit precision). Just like Groq's LPUs, d-Matrix chips can be deployed standalone, or as part of a heterogeneous compute config using GPUs for the compute-intensive prompt processing (prefill) phase of the inference pipeline and its Raptor accelerators for the memory bandwidth bound token generation (decode) phase. And for enterprises interested in the ultra-low latency inference capabilities of Groq, d-Matrix's chips may offer a cheaper point of entry since fewer XPUs would be required. Nvidia's walled garden We looked at Nvidia's emerging IP licensing strategy in more detail last week, but in a nutshell, Nvidia stands to gain a lot more than licensing revenues from NVLink Fusion adopters. In addition to licensing Nvidia's interconnect tech, d-Matrix plans to pair its accelerators with the GPU giant's Vera CPUs, NVSwitch appliances, BlueField and ConnectX NICs, and SpectrumX Ethernet products. In other words, Nvidia stands to make a lot of money even if it's not selling GPUs. It seems a fair number of chip designers are willing to make that kind of deal to avoid having to design their scale up networks or rack systems. Last week, MediaTek joined Marvell, Qualcomm, Arm, Fujitsu, and Amazon Web Services in adopting Nvidia's NVLink Fusion interconnects. Nvidia is so invested in getting folks into its walled garden that it's spending billions on incentives to get them in the door. As part of the MediaTek deal last week, Nvidia invested $3.5 billion in the SoC designer, while it spent $2 billion to get Marvell in the door. No word on whether the latest deal included any such terms. (R)
External Content
Source RSS or Atom Feed
Feed Location http://www.theregister.co.uk/headlines.atom
Feed Title www.theregister.com - Articles
Feed Link https://www.theregister.com/
Reply 0 comments