GPU-Accelerated Supply Chain Optimization: What Our Benchmarks with NVIDIA cuOpt Actually Show

authors

Ashwin Rao
EVP of AI at o9 & Adjunct Professor at Stanford
6 read min
Supply chain optimization problems at global enterprises are genuinely large — tens of millions of variables, tens of millions of constraints, problem structures that can take CPU-based solvers many minutes to solve on production infrastructure. When NVIDIA approached us about integrating cuOpt — their GPU-accelerated optimization solver — into the o9 Digital Brain platform, the first question our engineering team asked was simple: what does this actually do to solve time on real enterprise data?
We now have a clear answer for large-scale Linear Programs. The results are better than we initially reported, and they're worth sharing in detail.
The Test: Real Enterprise Data, Real Scale
Our benchmarks used actual production-scale planning datasets, not synthetic examples. The primary benchmark — a supply chain inventory optimization and production planning workload — represents the scale of complexity o9 customers face in live environments:
| Metric | Value |
| Problem Type | Linear Programming (LP) — Supply Chain Planning / Inventory Optimization |
| Number of Variables | 30,148,213 (~30 million) |
| Number of Constraints | 15,730,620 (~15.7 million) |
| Non-zeros | 94,389,441 (~94 million) |
| Hardware: CPU Baseline | 8-core CPU, 32 GB RAM (AWS c5.4xlarge) |
| Hardware: GPU (cuOpt) | NVIDIA B200 GPU |
| cuOpt Version | 26.06 (main branch, June 2026) |
The Result: ~10x Faster, Near-Parity Solution Quality
On the NVIDIA B200 GPU, cuOpt solved the LP1 supply chain problem in 57.4 seconds with optimal solution status. The same problem took a leading commercial solver at 661.7 seconds on the CPU baseline. That is an approximately 10x reduction in solve time — cutting a nearly 11-minute solve to under one minute.
Equally important: solution quality. The objective value from cuOpt (1,947,311,820) was within 0.008% of best in class commercial solver’s optimal objective (1,947,475,498). Both solvers reported optimal status. This is not a speed-at-the-expense-of-quality result — it is genuine competitive solution quality delivered at GPU speed.
Note: Objective value gap between solvers is 0.008% — within normal tolerance for production planning decisions. Both runs reported optimal solution status.
Why This Matters: Planning Agility at Enterprise Scale
A 10x speedup on a 30-million-variable production planning LP is not a marginal gain. In supply chain planning, solve time directly constrains how many scenarios a planning team can evaluate in a given window.
Consider what this means in practice. A monthly S&OP process that previously relied on overnight batch optimization runs can now iterate multiple times intraday. A production planning team working against tight daily cycles can explore more scenarios, respond to disruptions faster, and run optimization against more current data. Problems that previously required scheduling a long compute window can now be run on-demand.
This shifts the economics of optimization from a scarce, batched resource to something that can be embedded more continuously into planning workflows.
A Note on What Changed Between Runs
It is worth being transparent about the progression of these results, because it illustrates how rapidly this technology is evolving. Our initial benchmarks, run on NVIDIA H100 GPU infrastructure, showed meaningful speedups on smaller datasets but surfaced constraint violation issues on the large LP workload under default settings. Those results required parameter tuning to improve solution quality, and even with tuning, the fastest runs on H100 achieved approximately 8x speedup.
The B200 results in this post represent NVIDIA's latest cuOpt release (v26.06, June 2026) running on their newest GPU architecture. The constraint violation issues observed on H100 under default settings do not appear in the B200 runs — both the default and tighter-tolerance configurations report optimal status with near best-in-class solution quality. The speedup also improved, from approximately 8x on H100 to approximately 10x on B200.
This rate of improvement is one reason we are investing in this collaboration. cuOpt is a rapidly evolving technology, improving meaningfully across releases and hardware generations — and these results show the pace of that progress.
Scope: Where GPU Acceleration Helps Today
These results apply specifically to large-scale LP problems — a common and important class of supply chain optimization that includes large inventory optimization runs, production planning LPs, and network flow problems at enterprise scale. For this problem class, the B200 results represent a compelling, production-relevant case for GPU-accelerated solving.
For MILP workloads and future optimization frontiers including Nonlinear Programming, see the Expanding Horizons section below.
o9's solver-agnostic architecture means customers don't have to pick a single solver. They get the right engine for each problem type — automatically, without re-platforming.
Expanding Horizons: MILP and Nonlinear Programming
Our LP results represent one dimension of a broader collaboration. For Mixed-Integer Linear Programming (MILP) — the problem class that governs production scheduling, facility location, network design, and many of supply chain’s most complex decisions — NVIDIA’s cuOpt MIP solver is under active development. Early benchmarks are promising, and we expect GPU-accelerated MILP to become a meaningful capability as the solver matures. We will publish updated benchmarks as results warrant.
Further out, the o9–NVIDIA collaboration is exploring Nonlinear Programming (NLP) — optimization over nonlinear objective functions and constraints, which are increasingly relevant in AI-driven demand sensing, pricing, and scenario optimization. GPU architectures, with their native support for parallel gradient computation, are well-suited to accelerate the iterative methods that underpin NLP solvers. This is an area where o9’s domain expertise in supply chain math and NVIDIA’s GPU computing infrastructure create a genuinely interesting research and product opportunity. Our Stanford ICME connection is particularly relevant here, as nonlinear and stochastic optimization are active research areas within the department.
What's Coming: Multi-GPU and Accelerated Presolve
NVIDIA has shared two near-term developments for the cuOpt LP solver. First, a multi-GPU version is in development that will support problems exceeding 200 million variables — well beyond the scale of our current benchmark datasets. Second, an accelerated LP presolve is being developed that will further reduce solve times beyond what the B200 results already show.
Both capabilities are on a near-term roadmap and will be immediately accessible to o9 customers through our solver-agnostic platform as they become available.
In my work with graduate students at Stanford's Institute for Computational & Mathematical Engineering (ICME) — which has deep research ties with NVIDIA — I have watched GPU-native optimization evolve from an academic curiosity into a genuine engineering advantage for production workloads. What our B200 benchmarks show is that this transition is happening faster than most expected.
“A 10x speedup with near-parity solution quality on a 30-million-variable enterprise planning problem, running in under one minute — that is a result worth taking seriously. We will continue to push this collaboration forward, expand benchmark coverage across problem types, and bring the results directly to o9 customers.”
Dr. Ashwin Rao
Executive Vice President, AI Strategy and R&D
About o9 Solutions
o9 Solutions is a leading Enterprise Knowledge and AI-powered platform helping companies build Agile, Adaptive & Autonomous Planning & Execution Models for transforming enterprise decision-making in environments of rising volatility and uncertainty. Whether it is improving forecast accuracy, matching demand and supply and driving collaboration across the multi-tier supply chain to improve resilience at optimal costs and inventory, or optimizing new product and commercial initiatives to drive revenue growth and margins, decision-making processes from long-range to tactical to execution horizon can be made faster and smarter and connected on o9’s Digital Brain Platform.
o9 brings together game-changing technology innovations — such as innovative enterprise knowledge graph modeling, big data analytics, advanced algorithms for forecasting, demand/supply balancing, scenario planning, real time learning, collaboration, generative and agentic AI, easy-to-use interfaces and cloud-based delivery, and innovative management methods — as well as organization, process and change management best practices to transform decision-making speed and intelligence.

Neuro-Symbolic Agentic AI for Agile and Adaptive Enterprises
Enterprises don't struggle to gather data. They struggle to turn it into action while it still matters.
By the time the root cause is clear, the window to respond has often closed. This White Paper outlines how neuro-symbolic AI changes that equation, giving leaders a path from signal to grounded, governed decision in the same cycle the problem appears.
About the authors

Ashwin Rao
EVP of AI at o9 & Adjunct Professor at Stanford
Ashwin Rao is EVP-AI at o9 Solutions with the responsibility for o9's AI Strategy & Architecture as well as leading o9's R&D team. Ashwin is also an Adjunct Professor in Applied Mathematics at Stanford University, focusing his research and teaching in the field of Reinforcement Learning (RL), and has written a book on RL with applications in Finance, Supply-Chain and Dynamic Pricing. Previously, Ashwin was the Chief AI Officer at QXO, VP of AI at Target Corporation, Managing Director of Market Modeling at Morgan Stanley, and VP of Quant Trading Strategies at Goldman Sachs. Ashwin has a Ph.D. in Theoretical Computer Science from University of Southern California and a B.Tech in Computer Science from IIT-Bombay. Ashwin resides in Palo Alto, CA.











