NVIDIA Groq 3 LPX Boosts Vera Rubin Power Efficiency
NVIDIA is integrating the Groq 3 LPX accelerator into its Vera Rubin NVL72 platform to dramatically boost power efficiency and throughput for massive interactive AI models.

NVIDIA will integrate the Groq 3 LPX low-latency accelerator into its Vera Rubin NVL72 platform in the second half of 2026. This integration targets the severe power constraints of modern AI factories, especially during highly interactive workloads. Combining the Groq 3 LPX with the Vera Rubin NVL72 yields up to a 35-fold increase in throughput per megawatt compared to the older GB200 NVL72 architecture when serving highly interactive, long-context models exceeding two trillion parameters.
The efficiency gains stem from the deterministic execution model of the Groq 3 LPX, which enables cycle-exact scheduling of compute and data movement across all 256 LPU chips in a rack. Because the compiler predicts the exact electrical current draw for every clock cycle, it can deploy Preemptive Power (PEP) and Clock Period Synthesis (CPS). PEP adjusts the voltage supply before demand spikes arrive, while CPS adjusts individual clock cycles to smooth out sudden surges. By mitigating these electrical fluctuations, the system slashes voltage droop by over 60 percent. This allows the hardware to run with a tighter safety margin, cutting overall workload power consumption by a low-double-digit percentage.
These chip-level features work alongside rack and factory-level power management. Inside the Vera Rubin NVL72, rack-level capacitors and Intelligent Power Smoothing software absorb bursty spikes, allowing operators to plan infrastructure around sustained demand rather than worst-case peaks. At the factory level, NVIDIA DSX MaxLPS software dynamically reallocates power across different racks to reclaim unused capacity. This optimization allows operators to run up to 40 percent more GPUs and boost token throughput by 35 percent within a fixed power budget.
For AI practitioners, these advancements redefine energy efficiency as the primary metric for data center value. By minimizing the electrical safety margins that typically waste power, the system ensures more of a data center's energy budget goes directly to token generation, making the deployment of massive, interactive models economically viable.
This is our own summary of reporting by NVIDIA Developer Blog


