Performance per Watt Is the Key Metric for AI Infrastructure Efficiency
NVIDIA
CoreWeave
Perplexity AI
Fireworks AI
Performance per watt is the foundation for AI factories, as power is the main constraint. NVIDIA's Blackwell NVL72 platform delivers up to 25x performance per watt over Hopper for MoE models, and the upcoming Vera Rubin platform builds on this to further improve rack-scale energy efficiency.
NVIDIA argues that power is the inescapable constraint for AI infrastructure, making performance per watt the critical metric for revenue and profitability. With frontier AI models increasingly using mixture-of-experts (MoE) architecture, GPU domain size matters: serving MoE efficiently requires a 72-GPU domain rather than the 8-GPU domain of the Hopper generation. The Blackwell NVL72 platform, with full-stack codesign, delivers up to 25x performance per watt over Hopper for the latest open models. NVIDIA emphasizes that production experience is key: Anthropic, OpenAI, and SpaceXAI use Blackwell NVL72 for inference, while CoreWeave, Perplexity, and Fireworks AI deploy various models on the platform. The next-generation Vera Rubin platform builds on this foundation to further elevate energy efficiency.
- Сокращения
- MoE = Mixture of Experts
- GPU = Graphics Processing Unit
- NVLink = NVIDIA NVLink
- SHARP = SHARP (in-network computing technology, exact expansion not given)
- DSX = NVIDIA Data Center Switch eXtension (likely, not explicit)
Source: NVIDIA blog —
original
