Article 78QEW AMD's 192 GB Gorgon Halo prices might leave you petrified

AMD's 192 GB Gorgon Halo prices might leave you petrified

by
from www.theregister.com - Articles on (#78QEW)
Story ImageThe first systems powered by AMD's Gorgon Halo system-on-chip (SoC) platform have arrived, boasting up to 192 GB of unified memory on board. That's enough to put DeepSeek V4 Flash on your desk if you can afford to pay a hefty premium for all that RAM. GMKtec's EVO-X5 Pro is among the first to feature the House of Zen's top-specced Ryzen AI Max+ 495 SoC - you can see why we're just going to call it Gorgon Halo from here on out - but with regular pricing starting at $6,799, all that memory doesn't come cheap. But for local AI enthusiasts and perhaps privacy-conscious small businesses, it might be worth it just to run larger, more capable models from the security of their homes and offices. At 4-bit precision, Gorgon Halo is equipped with enough memory to run models to around 345 billion parameters on Linux or 320 billion parameters on Windows. The difference here comes down to memory partitioning. On Windows, you'll need to allocate 160 GB of memory to the GPU, while the Linux kernel and AMD's GPU drivers enable users to take advantage of nearly all of the available memory. This puts high-profile frontier models like the 284 billion parameter DeepSeek V4 Flash or Z.AI's 320 billion parameter GLM-5.3-Flash within reach of Gorgon Halo customers. The chip Announced earlier this year, Gorgon Halo is essentially a factory overclocked version of the Strix Halo APU that powers AMD's AI Halo workstation that we reviewed this summer. Both chips can be had with up to 16 Zen 5 cores, a 40-compute-unit integrated GPU, and an XDNA 2-based NPU to power all the Copilot+ features you didn't ask for but Microsoft is pushing on you anyway. The big difference between the chips is that AMD has bumped up the clocks by about 100 MHz across the board, and Gorgon Halo now supports 192 GB of faster 8,533 MT/s LPDDR5x memory compared to Strix Halo's 128 GB of 8,000 MT/s. In terms of performance, we don't expect the clock bump to make a meaningful difference. With that said, the faster memory could help a little bit for local AI enthusiasts. Large language model (LLM) inference - that is, the process of running pre-trained models - is predominantly memory bound. That means the faster your memory is, the faster the system can generate tokens. But while Gorgon Halo does achieve high memory bandwidth at 273 GB/s, that's only about 6.5 percent faster than Strix Halo at 256 GB/s. At most, you can expect to see a few more tokens a second on highly optimized models. Priced out of the market Gorgon's main attraction is really the prospect of higher memory capacity. But as we mentioned before, the ongoing memory shortage has not been kind to consumer electronics. A year ago, a 128 GB Strix Halo box would have set you back between $2,000 and $3,000 depending on OEM. DDR5 prices have skyrocketed since then and by the time AMD launched its AI Halo workstation this summer, prices had jumped to $4,000. With 50 percent more memory, you can expect Gorgon Halo-based products to be even pricier. The GMKtec unit is already 64 percent more expensive and the brand is usually on the more affordable side of things. More premium brands like Framework and notebooks like HP's ZBook Ultra G3a will likely command a steep premium. Unfortunately for AMD, there's not a lot that can be done. LPDDR5x is in short supply thanks to the AI boom. Each NVL72 rack Nvidia sells is packed with between 36 and 54 TB of the stuff, and that's just CPU memory. The situation is only going to get worse as AMD embraces LPDDR5x for its own AI-optimized Verano Epycs. If you need a lot of memory right now, you're in for a lot of pain, and the situation isn't expected to improve anytime soon. Earlier this year, Samsung warned that memory supplies are likely to remain constrained through 2028, which is bad news for everyone including AMD's top competitor Nvidia. Nvidia has also seen the price of AI workstations based on its GB10 SoC roughly double over the past year, jumping from $3,000-$4,000 a year ago to $6,000-$8,000 today, which doesn't bode well for its Windows SoC debut. Announced at Computex late this spring, Nvidia's RTX Spark is a family of premium Windows systems from major OEMs based around the same GB10 SoC and offering up to 128 GB of memory capacity. Considering these systems will be offered as both mini PCs and notebooks, we wouldn't be surprised to see them command a premium over existing GB10 boxes. Once again, this puts AMD in a position to undercut its competitor by offering more memory for less, but it's awfully hard to cross shop when you can't afford either option. (R)
External Content
Source RSS or Atom Feed
Feed Location http://www.theregister.co.uk/headlines.atom
Feed Title www.theregister.com - Articles
Feed Link https://www.theregister.com/
Reply 0 comments