OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language Models
Ali AhmadiTeshniziWenzhi GaoMadeleine Udell
Introduces OptiMUS, a modular multi-agent framework that translates natural language problem descriptions into mathematical formulations and debugged solver code, outperforming existing methods by over 30% on challenging mixed-integer linear programming benchmarks.
Mathematical optimization is essential for improving operational efficiency in industries such as manufacturing, logistics, healthcare, and energy. However, translating complex business problems into mathematical models traditionally requires specialized expertise, creating a significant barrier for many organizations. While large language models (LLMs) offer the potential to automate optimization modeling directly from natural language descriptions, standard prompting techniques struggle with long problem narratives, large numerical datasets, and unexecutable or logically flawed code.
The article introduces and evaluates OptiMUS, a modular multi-agent LLM framework designed to formulate and solve linear programming (LP) and mixed-integer linear programming (MILP) problems from plain English text. To evaluate performance under realistic conditions, the researchers also created NLP4LP, a new benchmark dataset of 67 complex, long-description optimization problems drawn from standard operations research textbooks.
OptiMUS operates by preprocessing natural language into a structured representation that decouples large numerical data files from problem descriptions, preventing context overflow. A coordinating manager agent directs a specialized team comprising a formulator, a programmer, and an evaluator. The system builds and maintains a tripartite connection graph linking parameters, variables, and constraints, which enables individual components to be drafted and debugged within concise, targeted prompts. OptiMUS was evaluated across three benchmarks—NL4OPT (simple problems), ComplexOR, and NLP4LP—against standard prompting, Reflexion, and Chain-of-Experts baselines using commercial solvers.
The evaluation produced several key findings regarding accuracy and system scalability:
- OptiMUS outperformed all baseline methods across every benchmark, achieving an accuracy of 78.8% on NL4OPT, 66.7% on ComplexOR, and 72.0% on NLP4LP.
- On challenging datasets with long descriptions, OptiMUS improved accuracy by roughly 30 percentage points over existing methods (e.g., reaching 72.0% on NLP4LP compared to 53.1% for Chain-of-Experts and 35.8% for standard prompting).
- The connection graph kept prompt sizes relatively stable as problem complexity increased (around 3,146 characters on NLP4LP), whereas baseline prompt lengths grew substantially (reaching 3,825 characters).
- System performance depended heavily on underlying reasoning capabilities; substituting GPT-4 with smaller open-source models such as Mixtral-8x7B caused accuracy on NLP4LP to drop sharply from 71.6% to 3.0%.
- Iterative debugging proved critical for complex problems, with the manager frequently calling the programmer and evaluator agents to fix coding errors before attempting mathematical reformulations.
These results demonstrate that modular, agent-based architectures can effectively mitigate context limitations and error rates in LLM-driven technical workflows. By reliably bridging the gap between natural language descriptions and professional solvers without sending full datasets into prompts, the framework reduces the cost and technical barriers to deploying operations research solutions in smaller businesses, public services, and non-profit organizations.
Organizations seeking to implement automated optimization should adopt modular agent architectures with explicit separation between problem logic and numerical data. Teams should prioritize advanced reasoning models for orchestrating tasks and implement robust self-debugging loops. Immediate next steps for further development include fine-tuning smaller, open-weight models on modular prompt templates to lower operational costs, incorporating human-in-the-loop review checkpoints, and expanding reinforcement learning to optimize agent task scheduling.
While OptiMUS achieves strong performance, several limitations remain. When the system fails, 43% to 62% of errors stem from incorrect mathematical modeling or missing constraints rather than coding bugs. Furthermore, LLMs remain susceptible to hallucinations and subtle semantic confusion regarding parameter and variable designations. As a result, OptiMUS is best positioned as an assistive tool to augment human decision-makers rather than an entirely unmonitored solver in high-stakes or safety-critical operational environments.
- Paper: AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, Qingyun Wu et al. (2023). AutoGen establishes the multi-agent conversational framework that underlies specialized modular systems like OptiMUS for automated code generation, formulation, and debugging.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). Program of Thoughts introduces the strategy of disentangling natural language reasoning from computation via external code generation, a concept foundational to how OptiMUS separates formulation from solver execution.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). Least-to-most prompting establishes the technique of decomposing complex problems into sequential subproblems, which OptiMUS adapts via connection graphs to prevent context overflow in mathematical modeling.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). Inner Monologue demonstrates the power of closed-loop execution feedback for dynamic replanning, providing the foundation for OptiMUS's iterative evaluator and programmer debugging loops.
- Paper: CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Guohao Li et al. (2023). CAMEL provides early architectural foundations for collaborative, role-playing multi-agent systems that autonomously decompose and solve technical problems.
- Paper: Mixtral of Experts, Albert Q. Jiang et al. (2024). Mixtral 8x7B serves as a primary open-source baseline evaluated in OptiMUS to analyze how underlying reasoning capabilities affect multi-agent optimization performance.
- Paper: AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML, Patara Trirat et al. (2025). AutoML-Agent extends the multi-agent orchestration of technical workflows seen in OptiMUS to full-pipeline automated machine learning across multimodal tasks.
- Paper: LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery, Pingchuan Ma et al. (2024). This work broadens the paradigm of coupling LLMs with external computational engines by integrating language agents with differentiable physical simulations for bilevel scientific optimization.
- Paper: AI Co-Mathematician: Accelerating Mathematicians with Agentic AI, Daniel Zheng et al. (2026). AI Co-Mathematician generalizes the agentic coordination of mathematical reasoning and code execution to interactive, end-to-end research workflows for human mathematicians.
- Paper: Learning to Orchestrate Agents in Natural Language with the Conductor, Stefan Nielsen et al. (2026). The Conductor advances multi-agent orchestration by using reinforcement learning to learn automated coordination and information routing among heterogeneous language models.
- Paper: Small Language Models are the Future of Agentic AI, Peter Belcak et al. (2025). This work directly addresses OptiMUS's observed trade-offs and future directions by evaluating how compact language models can economically replace massive LLMs in specialized agent subtasks.
- Paper: Toward Efficient Agents: Memory, Tool learning, and Planning, Xiaofang Yang et al. (2026). This survey systematically explores cost-performance and context-efficiency strategies across agent memory, tool use, and planning, building on challenges encountered in complex agent workflows.
