NVIDIA cuOpt Scales Decision Optimization to 100M Variables
NVIDIA has introduced a multi-GPU linear programming solver to its cuOpt engine, enabling organizations to process massive decision-optimization models with over 100 million variables.

NVIDIA has launched a Multi-GPU Primal-Dual hybrid gradient for Linear Programming (mPDLP) solver within its cuOpt platform. This new tool distributes massive linear programming problems across NVLink-connected GPUs. The development addresses a major bottleneck for enterprise practitioners whose massive planning models previously exceeded single-GPU memory limits or took hours to converge. The solver caps linear programming problems at 2.1 billion nonzeros and cuts peak memory usage per GPU by up to 6x compared to single-GPU PDLP.
To achieve this, the solver uses min-cut partitioning to analyze the sparsity of the constraint matrix, keeping tightly connected dependencies on the same GPU to minimize communication overhead. In benchmarks of over 100 linear programming instances on NVIDIA DGX B200 GPUs, performance gains strongly correlated with problem size. Speedups became noticeable once nonzero entries surpassed 10 million, with mPDLP achieving a 1.2x to 2.5x speedup over the previous D-PDLP method. On the tsp-gaia-10m instance, the solver achieved an 11.4x speedup on the core PDLP steps alone.
For practitioners in supply chain and energy, these scaling improvements translate directly to faster business decisions. Supply chain software company Kinaxis achieved a 3.3x speedup on a consumer packaged goods model containing over 135 million variables using eight NVLink-connected H100 GPUs. Similarly, energy consulting firm PSR saw a speedup of more than 5x on a stochastic energy expansion model with 185 million variables using eight B200 GPUs.
NVIDIA is already working on further optimizations to enhance the solver. Future updates will include weight-aware min-cut and hypergraph partitioning to balance computational loads, overlapping local computation with remote data transfers to hide communication latency, and feasibility polishing to trade absolute optimality for faster iterations. Developers can access these capabilities via the cuOpt mPDLP tutorial and the open-source code on GitHub.
This is our own summary of reporting by NVIDIA Developer Blog



