Researchers say software can ease AI data-center power pressure
Energy-aware inference, training and scheduling can cut use without sacrificing much performance, offering operators more headroom before new infrastructure is needed.
Researchers and infrastructure operators are making the case that software changes, not just new chips and more power, can ease AI data-center energy pressure. Work from the University of Michigan’s ML.Energy group found that FP8 inference on Alibaba’s Qwen 3 235B A22B Thinking model used about one-third less energy than bfloat16, and Chung’s Perseus optimizer cut training energy by up to 30% without lowering throughput. Nvidia is also using workload-aware Blackwell power profiles, which it says can save up to 15% of energy while preserving 97% or more of performance and lift throughput by as much as 13%. Separate research from ETH Zurich points to shifting flexible jobs across time and geography to follow grid conditions, though data-sovereignty rules and moving large datasets can limit that approach. The overall message is that software is not a replacement for better hardware or new infrastructure, but it can squeeze more useful work out of existing AI clusters.
Why it matters
For hyperscalers and other operators running large AI clusters, the near-term gains come from making current systems work harder, not only from buying more chips or power. The research points to lower energy use and steadier throughput from software choices in inference, training and scheduling, which can reduce pressure on already constrained data centers. Where workloads can move with grid conditions, software can also help shift demand, though border and data-movement limits narrow that option.
Keep or strike?
Does this story matter, or is it hype? Mark it before you see what everyone else did.
Sources
- Tom's Hardware