Article 77T77 AMD inches closer to its goal of making AI suck less ... energy

AMD inches closer to its goal of making AI suck less ... energy

by
from www.theregister.com - Articles on (#77T77)
Story ImageAI's thirst for power remains an ongoing concern, but fear not: AMD says it's making steady progress towards its goal of boosting rack efficiency 20x by the end of the decade. In a blog post published this week, the House of Zen estimates that, as of 2026, its systems are already 4x more efficient than they were in 2024. Certainly, a lot has changed since then. As you may recall, AMD began volume production of the MI300X, its first true datacenter GPU designed for AI, that year. The 750-watt part boasted up to 2.6 petaFLOPS of dense FP8 performance, which at the time made it competitive on paper with Nvidia's Hopper generation of AI accelerators. Since then, AMD has pulled every lever and pushed every button at its disposal to squeeze more FLOPS per watt from its GPU systems, push its memory and scale-up fabrics harder, and optimize its software stack in order to catch up with its larger, more successful rival. This included adding support for 4-bit floating point data types, new memory technologies, increasing interconnect speeds, and transitioning from conventional GPU servers to fully-integrated rack-scale systems. "The counterintuitive thing here... is the bigger the device, the more efficient it is," AMD SVP and Fellow Sam Naffziger told El Reg last year when the chipmaker announced the initiative. Last month, AMD revealed the fruits of its labors with the launch of said rack-scale compute platform, codenamed Helios, which crams 72 MI455X GPUs into a single massive system. Compared to the MI300X, each MI455X boasts between 7.7x and 15.4x higher floating point performance, 2.25x more HBM, 4.4x faster memory, and 4x chip-to-chip interconnect bandwidth. Without question, the chip is faster, but it also requires more than 3x the power. Instead, the biggest performance gains come from just how efficiently AMD can scale AI workloads across the system's six dozen accelerators. As usual, AMD isn't exactly a pioneer here. Nvidia made the leap to rack-scale in late 2024, with the launch of its Grace Blackwell-based NVL72 systems that also pack 72 GPUs into a single rack-sized system. At the time, Nvidia CEO Jensen Huang boasted that compared to an equivalent number of Hopper GPUs, GB200 NVL72 racks delivered a 4x uplift in training and 30x improvement in inference performance. While the benefits of rack-scale architectures are clear, it's worth emphasizing AMD is using a very different methodology to calculate efficiency, by weighting max achieved FLOPS, memory, and interconnect bandwidth differently for training and inference, rather than basing their comparison on real-world application performance. It's also worth nothing that AMD's 4x claim is an estimate. The first Helios units should ship to customers this calendar quarter. We expect the first MLPerf and InferenceX benchmarks to follow not long after. However, assuming AMD can make good on its goals, it says two Helios racks will be able to do the same work that required 570 racks full of kit in 2024. Put more realistically, for the same power customers will be able to deploy 20x more compute, assuming the bubble hasn't already popped by then and taken demand with it. (R)
External Content
Source RSS or Atom Feed
Feed Location http://www.theregister.co.uk/headlines.atom
Feed Title www.theregister.com - Articles
Feed Link https://www.theregister.com/
Reply 0 comments