Article 78XQ0 Software could be the easiest fix for hyperscalers' AI power squeeze, researchers say

Software could be the easiest fix for hyperscalers' AI power squeeze, researchers say

by
Chris Stokel-Walker
from Latest from Tom's Hardware on (#78XQ0)

Data center energy demand is shooting constantly upwards. By 2030, some 945TWh of electricity is expected to be used to meet AI's demand; the International Energy Agency believes - that's as much electricity as Japan uses today. The energy crunch is inescapable, and the industry's answer has largely been to focus on the physical: build more efficient GPUs and secure more power supply from an already tight grid. But is that myopic focus on hardware actually worth it? Instead, could most of the work running on those machines be cut back to save power?

Average power usage effectiveness (PUE) has barely changed for six successive years, according to the Uptime Institute's 2025 survey. PUE can show whether cooling and power systems are wasteful, but doesn't account for whether the software running on the servers is doing useful work. And given servers account for around 60% of electricity demand in a modern data center, while cooling ranges from about 7% in an efficient hyperscale site to more than 30% in a less-efficient enterprise facility, focusing on eking out more efficiency from the software side of the equation seems sensible.

The answer to what seems to be the core question - Can software or algorithms meaningfully help with power and energy?' - is yes, 100%," said Jae-Won Chung, a PhD candidate in computer science and engineering at the University of Michigan and researcher with the ML.Energy initiative, in an interview with Tom's Hardware Premium.

Chung suggests computing systems should be seen as a stack with hardware at the bottom, backed up by systems software, algorithms and applications. The hardware end of that equation is the hardest bit to tackle: chips are slow and costly to replace. But the other three layers are easier to change, and any gains in efficiency made there can compound.

ML.Energy's tests of the Alibaba Qwen 3 235B A22B Thinking model suggest that running inference in FP8 - a slightly lower-precision format - consumed a third less energy than bfloat16 versions on problem-solving tasks. Chung's ownPerseus training optimiseridentifies parts of a large-model training job that have less work to do, then slows them so they finish alongside busier parts. The system cut training energy by up to 30% without reducing throughput or changing the hardware.

This isn't a tweak around the edges being made by tinkerers: some of the biggest GPU firms are doing something similar. Nvidia's Blackwell power profiles fine-tune everything from GPU compute and memory frequencies, power limits, NVLink states and cache settings to fit the workload. Nvidia believes that can save up to 15% of energy while retaining 97% or more of overall performance, allowing power-constrained facilities to run more GPUs, raising throughput by as much as 13%.

The loudest rack in the room'

Chetan Visrolia, business development manager for data center infrastructure at SHI, said in an interview with Tom's Hardware Premium that the most obvious waste he encounters is often legacy equipment supporting old code. It is always the lowest-hanging fruit for quick savings," he said. It can easily be found in any data center, as it is the loudest rack on the floor."

Cloud instances can also be made a smarter size, and AI prompts that are repeatedly used can be stored in cache rather than recomputed on the fly. There are also efficiencies to be made by getting small models to handle routine requests while holding back the larger ones for harder tasks. At the same time, better batching and compilers can raise accelerator utilisation, and crunching down prompts while throwing limits on outputs can also tamp down the number of tokens processed.

Software can also change when and where electricity is drawn. Sophie Hall, a doctoral student at ETH Zurich's Automatic Control Laboratory who has studied workload shifting across Google's global data-center fleet, argues in an interview with Tom's Hardware Premium that electricity consumption isn't necessarily the main problem.

It's more like: when do they use it, where do they use it, and how is it interacting with the grid?" she said. If batch jobs aren't time-sensitive, they could be delayed until local demand falls or routed to a region with spare capacity and lower-carbon electricity. Hall's research uses day-ahead planning and real-time scheduling to respond to grid signals while preserving performance guarantees.

Taking that kind of action was less necessary when hyperscalers built way more capacity than they actually needed and the grid wasn't a bottleneck, Hall said. But the massive demand for AI has changed the equation.

There are limits, though: individual countries' data sovereignty rules can prevent a workload from crossing borders, while moving large datasets across networks is something that inference providers may be wary of in case things go wrong. The data needs to literally, physically, geographically move to a different location," Hall said. That's not that easy."

The scale of the data center demand powered by AI means that software isn't a silver bullet or a substitute for better chips and new electricity infrastructure. But it can act as a control layer that squeezes out every last drop from those increasingly expensive assets, while also helping shift demand to places where there is more capacity.

Jevons ' paradox is also at play, and needs to be considered when thinking about making systems run better. Efficiency doesn't reduce a facility's electricity use if each watt saved gets used up generating even more tokens. Power is the core bottleneck in AI data centers," said Chung. We really want to make the best use of every watt we consume."

External Content
Source RSS or Atom Feed
Feed Location https://www.tomshardware.com/feeds/all
Feed Title Latest from Tom's Hardware
Feed Link https://www.tomshardware.com/feeds.xml
Reply 0 comments