OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language Models

Ali AhmadiTeshniziWenzhi GaoMadeleine Udell

article2024ICML139 citations

Introduces OptiMUS, a modular multi-agent framework that translates natural language problem descriptions into mathematical formulations and debugged solver code, outperforming existing methods by over 30% on challenging mixed-integer linear programming benchmarks.

Listen

Mathematical optimization is essential for improving operational efficiency in industries such as manufacturing, logistics, healthcare, and energy. However, translating complex business problems into mathematical models traditionally requires specialized expertise, creating a significant barrier for many organizations. While large language models (LLMs) offer the potential to automate optimization modeling directly from natural language descriptions, standard prompting techniques struggle with long problem narratives, large numerical datasets, and unexecutable or logically flawed code.

The article introduces and evaluates OptiMUS, a modular multi-agent LLM framework designed to formulate and solve linear programming (LP) and mixed-integer linear programming (MILP) problems from plain English text. To evaluate performance under realistic conditions, the researchers also created NLP4LP, a new benchmark dataset of 67 complex, long-description optimization problems drawn from standard operations research textbooks.

OptiMUS operates by preprocessing natural language into a structured representation that decouples large numerical data files from problem descriptions, preventing context overflow. A coordinating manager agent directs a specialized team comprising a formulator, a programmer, and an evaluator. The system builds and maintains a tripartite connection graph linking parameters, variables, and constraints, which enables individual components to be drafted and debugged within concise, targeted prompts. OptiMUS was evaluated across three benchmarks—NL4OPT (simple problems), ComplexOR, and NLP4LP—against standard prompting, Reflexion, and Chain-of-Experts baselines using commercial solvers.

The evaluation produced several key findings regarding accuracy and system scalability:

  • OptiMUS outperformed all baseline methods across every benchmark, achieving an accuracy of 78.8% on NL4OPT, 66.7% on ComplexOR, and 72.0% on NLP4LP.
  • On challenging datasets with long descriptions, OptiMUS improved accuracy by roughly 30 percentage points over existing methods (e.g., reaching 72.0% on NLP4LP compared to 53.1% for Chain-of-Experts and 35.8% for standard prompting).
  • The connection graph kept prompt sizes relatively stable as problem complexity increased (around 3,146 characters on NLP4LP), whereas baseline prompt lengths grew substantially (reaching 3,825 characters).
  • System performance depended heavily on underlying reasoning capabilities; substituting GPT-4 with smaller open-source models such as Mixtral-8x7B caused accuracy on NLP4LP to drop sharply from 71.6% to 3.0%.
  • Iterative debugging proved critical for complex problems, with the manager frequently calling the programmer and evaluator agents to fix coding errors before attempting mathematical reformulations.

These results demonstrate that modular, agent-based architectures can effectively mitigate context limitations and error rates in LLM-driven technical workflows. By reliably bridging the gap between natural language descriptions and professional solvers without sending full datasets into prompts, the framework reduces the cost and technical barriers to deploying operations research solutions in smaller businesses, public services, and non-profit organizations.

Organizations seeking to implement automated optimization should adopt modular agent architectures with explicit separation between problem logic and numerical data. Teams should prioritize advanced reasoning models for orchestrating tasks and implement robust self-debugging loops. Immediate next steps for further development include fine-tuning smaller, open-weight models on modular prompt templates to lower operational costs, incorporating human-in-the-loop review checkpoints, and expanding reinforcement learning to optimize agent task scheduling.

While OptiMUS achieves strong performance, several limitations remain. When the system fails, 43% to 62% of errors stem from incorrect mathematical modeling or missing constraints rather than coding bugs. Furthermore, LLMs remain susceptible to hallucinations and subtle semantic confusion regarding parameter and variable designations. As a result, OptiMUS is best positioned as an assistive tool to augment human decision-makers rather than an entirely unmonitored solver in high-stakes or safety-critical operational environments.

arXiv: 2402.10172
Cover for OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language Models

Abstract

Optimization problems are pervasive in sectors from manufacturing and distribution to healthcare. However, most such problems are still solved heuristically by hand rather than optimally by state-of-the-art solvers because the expertise required to formulate and solve these problems limits the widespread adoption of optimization tools and techniques. This paper introduces OptiMUS, a Large Language Model (LLM)-based agent designed to formulate and solve (mixed integer) linear programming problems from their natural language descriptions. OptiMUS can develop mathematical models, write and debug solver code, evaluate the generated solutions, and improve its model and code based on these evaluations. OptiMUS utilizes a modular structure to process problems, allowing it to handle problems with long descriptions and complex data without long prompts. Experiments demonstrate that OptiMUS outperforms existing state-of-the-art methods on easy datasets by more than 20% and on hard datasets (including a new dataset, NLP4LP, released with this paper that features long and complex problems) by more than 30%. The implementation and the datasets are available at https://github.com/teshnizi/OptiMUS.

Table of Contents

  • 1. Introduction
  • 2. Background and Related Work
  • 3. Methodology
  • 3.1. Structured Problem
  • 3.2. Agents
  • 3.3. The connection graph
  • 4. Experiments
  • 4.1. Dataset
  • 4.2. Overall Performance
  • 4.3. Ablation Study
  • 4.4. Sensitivity Analysis
  • 4.5. Failure Cases
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. System Design and Software Engineering
  • A.1. Preprocessing
  • A.2. Debugging
  • A.3. Connection Graph
  • B. Optimization Techniques
  • C. Limitations and Weaknesses
  • C.1. Sensitivity of LLMs to and parameter variable names
  • C.2. Hidden Errors
  • D. Applying Existing Methods to Our Dataset
  • E. Prompts
  • E.1. Manager Prompt
  • E.2. Formulation Generation Prompt
  • E.3. Formulation Fixing Prompt
  • E.4. Clause Coding Prompt
  • E.5. Variable Coding Prompt
  • E.6. Debugging Prompt

Knowls

  1. Knowl 1 — OptiMUS modular LLM optimization agent

    model/method

    OptiMUS is a large-language-model agent that converts a natural-language description of a linear program or mixed-integer linear program into a mathematical formulation, executable solver code, and a verified solution. It separates the task into preprocessing, formulation, programming, execution, evaluation, and iterative correction rather than asking one model call to generate the entire optimization program. The experimental implementation uses Python and Gurobi, although the design can target other programming languages and solvers supported by the LLM.

    The system also prompts the formulator to recognize solver-relevant structures, including special ordered sets, indicator variables, general constraints, SAT or constraint-programming structure, and totally unimodular formulations. When an applicable structure is identified, the formulation is adjusted to use the solver's specialized interface where available.

  2. Knowl 2 — NLP4LP benchmark for long and complex optimization descriptions

    data/table

    NLP4LP, introduced with OptiMUS, is a benchmark of 67 natural-language optimization problems: 54 linear programs and 13 mixed-integer linear programs. Its problems cover facility location, network flow, scheduling, portfolio management, and energy optimization, and each instance contains a problem description, a sample parameter-data file, and a known optimal value obtained from a source solution or by manually solving the instance.

    The benchmark was designed to test long descriptions and multidimensional data, which are underrepresented in earlier datasets. The reported comparison is: NL4Opt has 518.0±110.7518.0 \pm 110.7 description characters, 1101 instances, 0 MILPs, and no multidimensional parameters; ComplexOR has 497.1±247.5497.1 \pm 247.5 characters, 37 instances, 12 MILPs, and multidimensional parameters; NLP4LP has 908.9±504.6908.9 \pm 504.6 characters, 67 instances, 13 MILPs, and multidimensional parameters. The publicly available ComplexOR collection contains 37 problems, while 21 problems were gathered from it for the reported experiments.

  3. Knowl 3 — Structured problem representation and preprocessing

    model/method

    OptiMUS first converts the natural-language input into a structured problem state with four components: parameters, clauses, objective, and background context. Each parameter stores a symbol, an inferred shape such as scalar or multidimensional array, and a textual definition. Numerical data from the description is removed from the parameter descriptions and retained as a separate data input, keeping later LLM prompts short.

    A clause is either the objective or a constraint. Initially, each clause contains its natural-language description; later it is augmented with a LaTeX formulation and solver code. The background component is a concise description of the real-world setting and is included in prompts to support contextual reasoning.

    Preprocessing extracts parameters, separates the objective from constraints, and removes redundant, irrelevant, or incorrect constraints. It can also infer implicit modeling requirements, such as nonnegativity or integrality, when they are required by the problem semantics even though they are not stated explicitly. The implementation additionally uses prompts to identify the optimization-problem type and background.

  4. Knowl 4 — Tripartite connection graph for local context retrieval

    model/method

    OptiMUS maintains a tripartite connection graph whose three node types are parameters, clauses, and decision variables. For every clause, the formulator identifies the parameters and variables that occur in its formulation; edges connect the clause to those relevant nodes. If formulation requires an auxiliary variable, the variable is added to the graph and linked to the clause.

    When an agent must formulate, code, debug, or revise a particular clause, OptiMUS retrieves only the connected parameters, variables, formulations, and code. This prevents the LLM from receiving the entire model and allows each objective or constraint to be processed independently. The graph therefore separates large data and long descriptions from the local context needed for a particular LLM call, while also preserving consistency across clauses.

  5. Knowl 5 — Manager-coordinated formulation, coding, and evaluation workflow

    algorithm

    The OptiMUS workflow takes a natural-language optimization problem and its parameter data as input and returns a completed formulation, solver program, and solution. The state P(t)P^{(t)} denotes the structured problem after iteration tt, and the conversation stores messages exchanged by the agents.

    Input: Natural-language optimization problem PP
    P(0)P^{(0)} <- PREPROCESS(PP)
    conversation <- empty list
    repeat
      (agent, task) <- MANAGER(conversation)
      if the manager returns DONE then
        stop
      end if
      (P(t+1)P^{(t+1)}, message) <- agent(P(t)P^{(t)}, task)
      append message to conversation
    until the manager returns DONE
    Output: completed mathematical model, solver code, and evaluated solution

    The manager selects one of three agents and assigns a task. The formulator defines variables, writes or repairs mathematical expressions, introduces auxiliary variables and constraints when needed, and updates the connection graph. The programmer converts formulations into solver code and repairs code marked as erroneous. The evaluator executes the generated program on the supplied data, reports runtime errors, and identifies the associated variable or clause. The manager can repeatedly call agents until the model is judged complete.

    For debugging, OptiMUS incrementally concatenates generated Python code blocks and executes the growing program in a subprocess. When execution fails, the most recently added constraint block is marked erroneous, and its error message plus graph-selected context is sent back to the programmer or formulator.

  6. Knowl 6 — Experimental evaluation protocol and correctness metric

    experimental setup

    OptiMUS was evaluated on NL4OPT, ComplexOR, and NLP4LP against standard prompting, Reflexion, and Chain-of-Experts. The reported main comparison uses GPT-4. Accuracy is the fraction of problem instances for which the generated code executes successfully and returns the correct optimal objective value; an instance is not counted as correct merely because its code compiles or avoids a runtime error. Reference optimal values come from the datasets or from manually solving the instances.

    The evaluation also measures runnable instances, compilation and runtime failures, the effect of disabling debugging, the effect of replacing GPT-4 with GPT-3.5 or Mixtral-8x7B, the number of allowed manager-agent calls, agent-selection frequency, and average prompt length. The authors emphasize accuracy as the primary metric because a short but irrelevant program can otherwise appear successful under compilation- or runtime-error metrics.

  7. Knowl 7 — OptiMUS outperforms prior natural-language optimization methods

    empirical result

    On the reported accuracy metric, OptiMUS is substantially more successful than standard prompting, Reflexion, and Chain-of-Experts across all three datasets. The comparison uses GPT-4 and the percentages are the fractions of correctly solved instances.

    NL4OPT: standard prompting 47.3%, Reflexion 53%, Chain-of-Experts 64.2%, and OptiMUS 78.8%.

    ComplexOR: standard prompting 9.5%, Reflexion 19.1%, Chain-of-Experts 38.1%, and OptiMUS 66.7%.

    NLP4LP: standard prompting 35.8%, Reflexion 46.3%, Chain-of-Experts 53.1%, and OptiMUS 72.0%.

    Thus OptiMUS improves over the strongest listed baseline by 14.6 percentage points on NL4OPT, 28.6 points on ComplexOR, and 18.9 points on NLP4LP. The results support the paper's claim that modular processing and structured context are particularly useful for complex problems.

  8. Knowl 8 — Debugging and model-scale ablations

    empirical result

    An ablation study shows that both iterative debugging and a large language model are important for OptiMUS. In this ablation run, the full GPT-4 system solved 78.7% of NL4OPT, 66.7% of ComplexOR, and 71.6% of NLP4LP instances. Removing debugging reduced accuracy to 72.3%, 57.1%, and 58.2%, respectively. Using GPT-3.5 only for the manager produced 74.9%, 52.4%, and 53.7%, whereas replacing all agents with GPT-3.5 produced 28.6%, 9.5%, and 14.4%. Replacing all agents with Mixtral-8x7B produced 6.6%, 0.0%, and 3.0%.

    The manager substitution has a relatively small effect on NL4OPT because many easy instances need only one formulation-programming-evaluation sequence. Its effect is larger on ComplexOR and NLP4LP, where coordinating repeated agent interactions matters more. The much larger degradation with smaller models is attributed in the paper to their weaker reasoning and their difficulty handling OptiMUS's longer, less conventional modular prompts.

  9. Knowl 9 — Iterative self-improvement preserves prompt scalability

    empirical result

    Allowing the manager to select agents repeatedly improves performance primarily on the difficult datasets. Most NL4OPT instances are solved after one call each to the formulator, programmer, and evaluator, while ComplexOR and NLP4LP often require repeated calls to identify and repair initial mistakes. The sensitivity plot evaluates maximum call budgets from 3 through 10 and shows that additional calls increase accuracy on the difficult datasets, demonstrating the value of iterative self-correction.

    OptiMUS also keeps prompt length comparatively stable as problem difficulty increases. The reported average prompt lengths, in characters, are:

    • NL4OPT: Chain-of-Experts 2003±4562003 \pm 456; OptiMUS 2838±8222838 \pm 822.
    • ComplexOR: Chain-of-Experts 3288±7803288 \pm 780; OptiMUS 3241±11943241 \pm 1194.
    • NLP4LP: Chain-of-Experts 3825±10023825 \pm 1002; OptiMUS 3146±11453146 \pm 1145.

    The programmer and evaluator are selected more often than the formulator on the harder datasets because code errors are frequent and comparatively easy to detect and repair, whereas formulation errors require deeper semantic reasoning.

  10. Knowl 10 — Failure modes and remaining reliability limitations

    limitation

    Among OptiMUS failures, the paper normalizes the three categories to sum to 100% within each dataset. Incorrect modeling accounts for 43.0% of failures on NL4OPT, 62.5% on ComplexOR, and 53.8% on NLP4LP. Missing or wrong constraints account for 36.0%, 12.6%, and 15.4%, respectively. Coding errors account for 21.0%, 24.9%, and 30.8%, respectively. Overall, the percentages of runnable instances are 85.6% for NL4OPT, 76.2% for ComplexOR, and 75.0% for NLP4LP; the percentages solved correctly are 78.8%, 66.7%, and 72.0%.

    The dominant failure on the more complicated datasets is an incorrect mathematical model, while missing constraints are relatively more common on NL4OPT. Runtime evaluation cannot detect hidden semantic errors such as an omitted constraint that still yields executable code or a redundant but incorrect constraint. The paper also notes that LLMs can be biased by variable and parameter names—for example, interpreting price as a decision variable in a production problem because it appeared as a variable in a different training example—but it did not conduct a large-scale experiment on this phenomenon. These limitations mean that OptiMUS is presented as a proof of concept rather than a fully reliable replacement for expert model validation.

Coverage note — Detailed prompt templates and solver-specific code examples were omitted because they implement the summarized pipeline rather than constitute separate load-bearing contributions.

References

  1. 1.Aastrup, J. and Kotzab, H. Forty years of out-of-stock research–and shelves are still empty. The International Review of Retail, Distribution and Consumer Research, 20(1):147–164, 2010.
  2. 2.Akgun, O., Miguel, I., Jefferson, C., Frisch, A., and Hnich, B. Extensible automated constraint modelling. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 25, pp. 4–11, 2011.
  3. 3.Alibaba Cloud. Alibaba cloud mindopt copilot, 2022. URL https://opt.alibabacloud.com/chat.
  4. 4.Antoniou, A. and Lu, W.-S. Practical optimization: algorithms and engineering applications, volume 19. Springer, 2007.
  5. 5.Beale, E. and Forrest, J. J. Global optimization using special ordered sets. Mathematical Programming, 10:52–69, 1976.
  6. 6.Beldiceanu, N. and Simonis, H. A model seeker: Extracting global constraint models from positive examples. In International Conference on Principles and Practice of Constraint Programming, pp. 141–157. Springer, 2012.
  7. 7.Beldiceanu, N. and Simonis, H. Modelseeker: Extracting global constraint models from positive examples. Data Mining and Constraint Programming: Foundations of a Cross-Disciplinary Approach, pp. 77–95, 2016.
  8. 8.Bertsimas, D. and Tsitsiklis, J. N. Introduction to linear optimization, volume 6. Athena scientific Belmont, MA, 1997a.
  9. 9.Bertsimas, D. and Tsitsiklis, J. N. Introduction to linear optimization, volume 6. Athena scientific Belmont, MA, 1997b.
  10. 10.Bessiere, C., Koriche, F., Lazaar, N., and O’Sullivan, B. Constraint acquisition. Artificial Intelligence, 244:315–342, 2017.
  11. 11.Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G. B., Lespiau, J.-B., Damoc, B., Clark, A., et al. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pp. 2206–2240. PMLR, 2022.
  12. 12.Chen, H., Constante-Flores, G. E., and Li, C. Diagnosing infeasible optimization problems using large language models. arXiv preprint arXiv:2308.12923, 2023.
  13. 13.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. Palm: Scaling language modeling with pathways, 2022.
  14. 14.De Raedt, L., Passerini, A., and Teso, S. Learning constraints from examples. In Proceedings of the AAAI conference on artificial intelligence, volume 32, 2018.
  15. 15.Gamrath, G., Berthold, T., Heinz, S., and Winkler, M. Structure-based primal heuristics for mixed integer programming. Springer, 2016.
  16. 16.Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G. Pal: Program-aided language models. In International Conference on Machine Learning, pp. 10764–10799. PMLR, 2023.
  17. 17.Gurobi Optimization. 2023 state of mathematical optimization report, 2023. URL https://www.gurobi.com/resources/report-state-of-mathematical-optimization-2023/.
  18. 18.Holzer, J., Coffrin, C., DeMarco, C., Duthu, R., Elbert, S., Eldridge, B., Elgindy, T., Greene, S., Guo, N., Hale, E., Lesieutre, B., Mak, T., McMillan, C., Mittelmann, H., Oh, H., O’Neill, R., Overbye, T., Palmintier, B., Safdarian, F., Tbaileh, A., Hentenryck, P. V., Veeramany, A., and Wert, J. Grid optimization competition challenge 3 problem formulation. https://gocompetition.energy.gov/sites/default/files/Challenge3_Problem_Formulation_20230126.pdf, 2023. Accessed: Access Date.
  19. 19.Kiziltan, Z., Lippi, M., Torroni, P., et al. Constraint detection in natural language problem descriptions. In IJCAI, volume 2016, pp. 744–750. International Joint Conferences on Artificial Intelligence, 2016.
  20. 20.Li, B., Mellou, K., Zhang, B., Pathuri, J., and Menache, I. Large language models for supply chain optimization. arXiv preprint arXiv:2307.03875, 2023.
  21. 21.Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. Lost in the middle: How language models use long contexts, 2023.
  22. 22.Nace, D. Lecture notes in linear programming modeling, 2020. URL https://www.hds.utc.fr/˜dnace/dokuwiki/_media/fr/lp-modelling_upt_p2021.pdf.
  23. 23.OpenAI. Gpt-4 technical report, 2023.
  24. 24.Paranjape, B., Lundberg, S., Singh, S., Hajishirzi, H., Zettlemoyer, L., and Ribeiro, M. T. Art: Automatic multi-step reasoning and tool-use for large language models. arXiv preprint arXiv:2303.09014, 2023.
  25. 25.Ramamonjison et al, . Augmenting operations research with auto-formulation of optimization models from problem descriptions. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Industry Track, pp. 29–62, Abu Dhabi, UAE, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-industry.4.
  26. 26.Ramamonjison et al, . Nl4opt competition: Formulating optimization problems based on their natural language descriptions, 2023. URL https://arxiv.org/abs/2303.08233.
  27. 27.Saghafian, S., Austin, G., and Traub, S. J. Operations research/management contributions to emergency department patient flow optimization: Review and research prospects. IIE Transactions on Healthcare Systems Engineering, 5(2):101–123, 2015.
  28. 28.Shakoor, R., Hassan, M. Y., Raheem, A., and Wu, Y.-K. Wake effect modeling: A review of wind farm layout optimization using jensen’ s model. Renewable and Sustainable Energy Reviews, 58:1048–1059, 2016.
  29. 29.Shinn, N., Cassano, F., Berman, E., Gopinath, A., k Narasimhan, K., and Yao, S. Reflexion: Language agents with verbal reinforcement learning, 2023.
  30. 30.Singh, A. An overview of the optimization modelling applications. Journal of Hydrology, 466:167–182, 2012.
  31. 31.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., ` Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  32. 32.Wang, Y., Ivison, H., Dasigi, P., Hessel, J., Khot, T., Chandu, K. R., Wadden, D., MacMillan, K., Smith, N. A., Beltagy, I., et al. How far can camels go? exploring the state of instruction tuning on open resources. arXiv preprint arXiv:2306.04751, 2023.
  33. 33.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D. Chain-of-thought prompting elicits reasoning in large language models, 2023.
  34. 34.Williams, H. P. Model building in mathematical programming. John Wiley & Sons, 2013.
  35. 35.Wu, Q., Bansal, G., Zhang, J., Wu, Y., Zhang, S., Zhu, E., Li, B., Jiang, L., Zhang, X., and Wang, C. Autogen: Enabling next-gen llm applications via multi-agent conversation framework. arXiv preprint arXiv:2308.08155, 2023.
  36. 36.Xiao, Z., Zhang, D., Wu, Y., Xu, L., Wang, Y. J., Han, X., Fu, X., Zhong, T., Zeng, J., Song, M., and Chen, G. Chain-of-experts: When LLMs meet complex operations research problems. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=HobyL1B9CZ.
  37. 37.Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X. Large language models as optimizers, 2023.
  38. 38.Yao, E., Liu, T., Lu, T., and Yang, Y. Optimization of electric vehicle scheduling with multiple vehicle types in public transport. Sustainable Cities and Society, 52:101862, 2020.

Citation

MLA
AhmadiTeshnizi, A., et al. “OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2402.10172v1.
APA
AhmadiTeshnizi, A., Gao, W., & Udell, M. (2024). OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language Models. arXiv. http://arxiv.org/abs/2402.10172v1
Chicago
AhmadiTeshnizi, A., W. Gao, and M. Udell. 2024. “OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language Models”. arXiv. http://arxiv.org/abs/2402.10172v1.
Harvard
AhmadiTeshnizi, A., Gao, W. and Udell, M. (2024) “OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.10172v1.
Vancouver
1. AhmadiTeshnizi A, Gao W, Udell M (2024) OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language Models. arXiv

BibTeX

@article{ahmaditeshnizi2024optimus,
  title = {OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language Models},
  author = {AhmadiTeshnizi, Ali and Gao, Wenzhi and Udell, Madeleine},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.10172v1},
  eprint = {2402.10172}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/