FUNDING · COMPANIES · #1094
NVIDIA DSX MaxLPS boosts usable GPUs and AI throughput via policy-governed power sharing
In a joint evaluation with Nscale, NVIDIA’s DSX MaxLPS used policy-governed power sharing on GB300 NVL72 systems at Nscale’s Verne campus to run Kimi K2.5 workloads, allowing 192 GPUs under the same 264.4 kW provisioned power budget versus a 140-GPU static baseline and increasing normalized aggregate throughput by 49.2% (tokens/s/W from 4.10 to 6.12). Median and P75 latency remained within 5% of baseline while P99 time-to-first-token rose 17%; NVIDIA outlines a five-stage validation process operators should follow before scaling deployment.
KEY POINTS
- In a joint evaluation with Nscale, NVIDIA’s DSX MaxLPS used policy-governed power sharing on GB300 NVL72 systems at Nscale’s Verne campus to run Kimi K2.5 workloads, allowing 192 GPUs under the same 264.4 kW provisioned power budget versus a 140-GPU static baseline and increasing normalized aggregate throughput by 49.2% (tokens/s/W from 4.10 to 6.12).
- Median and P75 latency remained within 5% of baseline while P99 time-to-first-token rose 17%; NVIDIA outlines a five-stage validation process operators should follow before scaling deployment.
- It demonstrates a practical method to increase AI facility capacity within existing power budgets by reallocating unused power, while highlighting trade-offs in tail latency that operators must validate.
WHY IT MATTERS
It demonstrates a practical method to increase AI facility capacity within existing power budgets by reallocating unused power, while highlighting trade-offs in tail latency that operators must validate.