OptimAI turns natural-language optimization problems into solver code with a multi-agent LLM pipeline, reaching 88.1% on NLP4LP and 82.3% on Optibench and cutting error rates by 58% and 52% over the prior best
Synopsis
The work introduces OptimAI, an LLM-powered multi-agent framework that takes a natural-language optimization problem through four stages—formulation, planning, solver code generation, and reflective debugging—and adds UCB-based debug scheduling to switch dynamically among candidate plans; under zero-shot prompting it reaches 88.1% accuracy on NLP4LP with GPT-4o+o1-mini and 82.3% on Optibench with DeepSeek-R1, reducing error rates by 58% and 52% over the prior best, while ablations show that removing the planner or code critic drops productivity by 5.8× and 3.1× and that enabling UCB debug scheduling adds a further 3.3× productivity gain.
Figure 1: Overview of the OptimAI Pipeline.
· Page 5Interpretation
OptimAI splits natural-language optimization solving into four stages handled by specialized roles: a formulator that converts the text into a mathematical model with decision variables, objective, constraints, and problem type; a planner that proposes multiple candidate solution strategies and selects solvers before any coding; a coder that generates executable Python solver code; and a code critic that performs reflective debugging using error messages. Compared with prior methods such as OptiMUS and Optibench, the framework simultaneously offers natural-language input, planning before coding, multi-solver support, switching between plans, and distinct-LLM collaboration (Table 1), whereas earlier methods cover only subsets of these capabilities. The paper documents the four-stage pipeline and gives a formal state-update specification in Appendix A, with the full role prompts listed in Appendix B, so the method is reproducible from the prompts and Algorithm 1.
The paper casts the choice of which plan to debug next as a multi-armed bandit problem, using a UCB score to switch dynamically among candidate plans, and adds a decider role that scores plans and a verifier role that checks whether final outputs satisfy all constraints. Prior methods lack the ability to switch between plans (Table 1 marks OptiMUS, Optibench, and CoE as not supporting it), so this work explicitly introduces an exploration-exploitation trade-off into the debugging stage. The ablation in Table 7 on the hard subset of Optibench shows that enabling UCB cuts token usage from 64,552 to 18,072 (about 3.6×) and raises productivity from 0.70 to 2.32 (about 3.3×), while Pass@1 accuracy stays at 69% and executability edges up from 3.4 to 3.5.
Under zero-shot prompting, OptimAI outperforms prior best methods across datasets: 88.1% on NLP4LP with GPT-4o+o1-mini and 82.3% overall on Optibench with DeepSeek-R1, with over 99% of generated code executing without error. The paper reports error-rate reductions of 58% on NLP4LP and 48%, 47%, 68%, and 41% on the four Optibench subsets (Linear w/o Table, Linear w/ Table, Nonlinear w/o Table, Nonlinear w/ Table) relative to the prior best. Results come from Table 3 across four underlying LLMs (GPT-4o, GPT-4o+o1-mini, QwQ, DeepSeek-R1) and two benchmarks; the paper states the DeepSeek-R1 configuration exceeds the prior best on Optibench by 8.1 standard deviations.
The framework allows different stages to be assigned to different LLMs, and experiments observe synergistic gains from heterogeneous model combinations: Llama 3.3 70B or Gemma 2 27B alone reach 59% and 54% accuracy, while assigning Gemma 2 27B as planner and Llama 3.3 70B to the remaining roles raises accuracy to 77%. This moves multi-agent collaboration from homogeneous ensembling toward role-based heterogeneous model assignment and quantifies the resulting gain. Evidence comes from the 3×3 combination matrix in Table 6 on Optibench; the role ablation in Table 8 shows removing the planner requires 4.6× more revisions and drops productivity 5.8×, while removing the code critic increases revisions 3.6× and drops productivity 3.1×.
Perspective
The framework targets users who describe optimization problems in natural language and want executable solver code and a solution automatically; it applies under zero-shot settings to linear, nonlinear, and mixed-integer programming as well as combinatorial tasks such as TSP, JSP, and set covering. The paper states its current performance is comparable to a single skilled programmer and explicitly lists large-scale problems that typically require a team of human experts as beyond the current scope. What practitioners can reuse directly are the four-stage role decomposition, the prompt templates in Appendix B, the UCB debug scheduling in Algorithm 1, and the engineering conclusion that the planner and code critic are not optional; the paper also notes that the optimal number of plans n should increase as the decider becomes more capable, so that hyperparameter should be re-tuned as models improve.
The reported executability metric relies on human evaluation (a score of 4 means a completely correct solution), and inter-rater consistency is not elaborated in the text; Table 4 shows OptimAI's token usage (18,072) is higher than Optibench's (955), which the paper explains as the price of broader problem coverage, so the cost-versus-coverage trade-off depends on the use case. The paper notes the decider is not particularly strong, with about a 39.7% chance of selecting the ultimately successful plan on the first attempt, and expects the optimal plan count to rise as the decider improves, which suggests the current hyperparameter conclusions are tied to specific models. The paper also lists reinforcement-learning fine-tuning of the decider and scaling to large-scale problems requiring expert teams as future directions whose effects remain to be validated.
