Chain of Code: Reasoning with a Language Model-Augmented Code Emulator

Chengshu LiJacky LiangAndy ZengXinyun ChenKarol HausmanDorsa SadighSergey LevineLi Fei-FeiFei XiaBrian Ichter

article2024ICML186 citations

Proposes Chain of Code, a reasoning framework that interweaves Python execution with language model emulation for non-executable pseudocode, boosting reasoning accuracy on BIG-Bench Hard to 84% across both algorithmic and semantic tasks.

Listen

Large language models frequently struggle with complex reasoning problems that blend exact computation with subjective semantic understanding. Natural language reasoning techniques like Chain of Thought excel at linguistic nuances but fail at precise arithmetic or symbolic manipulation, whereas standard code execution methods fail when encountering non-algorithmic semantic subtasks, such as determining sarcasm or categorizing real-world concepts.

The article demonstrates and evaluates Chain of Code, a framework designed to enhance language model reasoning by combining the algorithmic precision of a code interpreter with the semantic commonsense of a language model. The core objective is to allow models to format complex reasoning problems as programs that seamlessly hand off semantic steps to the language model itself when code execution is not possible.

The researchers evaluated this framework across multiple benchmark datasets, primarily the 23 challenging tasks in BIG-Bench Hard and the GSM8K grade-school math suite. They tested several model families, including OpenAI completion models (text-davinci-003 and smaller variants), instruction-tuned models (GPT-3.5 and GPT-4), and PaLM-2. The approach operates in two stages: first, the model generates reasoning steps formatted as structured code or pseudocode; second, a Python interpreter runs the executable lines, and whenever an unexecutable semantic function is caught, a language model emulator (termed an LMulator) simulates the expected output and updates the shared program state.

The findings establish that Chain of Code significantly improves multi-step reasoning performance. On BIG-Bench Hard, the method achieved an overall accuracy of 84% using text-davinci-003, representing a 12% gain over Chain of Thought (72%) and substantially outperforming the average human baseline of 68%. When paired with GPT-4, accuracy reached 91%. The gains were especially pronounced on algorithmic tasks, where Chain of Code achieved 95% accuracy compared to 71% for Chain of Thought. Ablation studies confirmed that both interpreter execution and language model simulation are essential, as relying solely on Python dropped overall accuracy to 48%. Furthermore, the framework scaled effectively to smaller models and demonstrated strong generalization in cross-task prompting and physical robotics tasks.

These results demonstrate that expressing complex problems through program structure reduces calculation errors without sacrificing natural language understanding. For organizations deploying automated decision systems, customer-facing agents, or robotic workflows, this hybrid structure increases reliability, auditability, and execution accuracy across mixed semantic and numerical operations.

Organizations developing reasoning pipelines should consider adopting hybrid code-execution architectures for complex multi-step tasks rather than relying purely on natural language prompting. For production deployments, teams should implement sandboxing and security safeguards, as executing dynamically generated code introduces vulnerabilities if prompts are maliciously constructed.

Key limitations include increased computation time and context length requirements due to multi-turn execution and state tracking. The current implementation tracks state via string parsing into basic Python types, preventing the modification of custom, non-serialized objects. Further work is required to optimize inference latency, evaluate fine-tuned dedicated emulator models, and expand the framework to richer external tool ecosystems.

Cover for Chain of Code: Reasoning with a Language Model-Augmented Code Emulator

Abstract

Code provides a general syntactic structure to build complex programs and perform precise computations when paired with a code interpreter - we hypothesize that language models (LMs) can leverage code-writing to improve Chain of Thought reasoning not only for logic and arithmetic tasks (Chen et al., 2022; Nye et al., 2021; Austin et al., 2021), but also for semantic ones (and in particular, those that are a mix of both). For example, consider prompting an LM to write code that counts the number of times it detects sarcasm in an essay: the LM may struggle to write an implementation for “detect_sarcasm(string)” that can be executed by the interpreter (handling the edge cases would be insurmountable). However, LMs may still produce a valid solution if they not only write code, but also selectively “emulate” the interpreter by generating the expected output of “detect_sarcasm(string)”. In this work, we propose Chain of Code (CoC), a simple yet surprisingly effective extension that improves LM code-driven reasoning. The key idea is to encourage LMs to format semantic sub-tasks in a program as flexible pseudocode that the interpreter can explicitly catch undefined behaviors and hand off to simulate with an LM (as an “LMulator”). Experiments demonstrate that Chain of Code outperforms Chain of Thought and other baselines across a variety of benchmarks; on BIG-Bench Hard, Chain of Code achieves 84%, a gain of 12% over Chain of Thought. In a nutshell, CoC broadens the scope of reasoning questions that LMs can answer by “thinking in code”.

Table of Contents

  • 1. Introduction
  • 2. Chain of Code: Reasoning with a Language Model-Augmented Code Emulator
  • 2.1. Preliminaries
  • 2.2. Chain of Code
  • 2.3. Chain of Code Implementation
  • 2.4. Chain of Code Abilities
  • 3. Experimental Evaluation
  • 3.1. Baselines and Ablations
  • 3.2. Tasks
  • 3.3. Results
  • 4. Related Work
  • 5. Conclusions, Limitations, and Future Work
  • Impact Statement
  • References
  • Appendix
  • A.1. Quantitative results on language reasoning tasks
  • A.2. Quantitative results on the GSM8K Benchmark
  • A.3. Qualitative results on language reasoning tasks
  • A.4. Instruction Tuned Models
  • A.5. Robustness of Chain of Code
  • A.6. Robotics Applications

Knowls

  1. Knowl 1 — Chain of Code combines program structure with language-model execution

    model/method

    Chain of Code (CoC) has two stages. First, given a question, a language model generates a solution as code, pseudocode, or natural-language steps arranged in a code-like program. Second, an executor runs the program, using a conventional interpreter for executable operations and a language model—called an LMulator—to simulate operations that cannot be executed as code. The generated program can therefore use ordinary computation for precise algorithmic work and delegate semantic subtasks, such as deciding whether an item is compostable, to the LMulator. The results from both forms of execution contribute to the same solution.

  2. Knowl 2 — CoC interweaves Python execution and LMulator simulation through shared program state

    model/method

    In the implemented CoC executor, generated code is processed line by line. Python executes a line when possible; if execution fails or the line is not executable, the LMulator receives the question, preceding program lines, and execution history, and generates the resulting program-state update. That update is made available to Python so execution can continue, including through control flow such as loops and conditionals. The final response is taken from the variable named answer when available; for irrecoverable execution errors, the language model produces the answer. This shared-state, line-by-line design permits Python operations and semantic language-model judgments to be interleaved within one program.

  3. Knowl 3 — CoC improves BIG-Bench Hard results under both same-task and cross-task prompting

    empirical result

    The evaluation used the 23 tasks in BIG-Bench Hard (BBH), spanning semantic, numerical, and mixed reasoning. In same-task few-shot prompting, each query had three examples from its task family; in cross-task prompting, the three examples came from different task families. The reported overall results were:

    Model Prompt setting Direct CoT CoC
    text-davinci-003 Same task 55 72 84
    text-davinci-003 Cross task 50 55 61
    PaLM 2-S code variant Same task 49 61 78
    PaLM 2-S code variant Cross task 45 47 47

    Scores are percentages; the human average on the same-task BBH evaluation was 68%. Thus CoC reached 84% with text-davinci-003, 12 percentage points above its Chain of Thought (CoT) result and 29 points above direct answering. In cross-task prompting, text-davinci-003 CoC scored 61%, compared with 55% for CoT and 50% for direct answering; the PaLM 2-S cross-task CoC score tied CoT at 47%. On GSM8K, CoC also exceeded both baselines: text-davinci-003 scored 71% versus 63% CoT and 16% direct in same-task prompting, and 60% versus 55% and 14%, respectively, in cross-task prompting.

  4. Knowl 4 — CoC’s largest aggregate gains occur on algorithmic tasks, while NLP performance matches CoT

    empirical result

    On BBH, the authors’ aggregate breakdown gives the following percentages for direct answering, Chain of Thought (CoT), and CoC (Interweave):

    Task group Direct CoT CoC
    Natural-language tasks 67 74 74
    Algorithmic tasks 41 71 95
    All tasks 55 72 84

    CoC therefore matched CoT on the natural-language task average while substantially exceeding it on the algorithmic average. The execution-type breakdown further shows that CoC scored 100% on tasks with repeated programs executable by Python, compared with 61% for CoT and 38% for direct answering. For different Python-executable programs, the respective scores were 89%, 84%, and 49%; for repeated programs requiring LM execution, 76%, 73%, and 72%; and for different programs requiring LM execution, 72%, 68%, and 53%. These results indicate that the strongest performance was on tasks where the program could be reused across questions and run by Python, while CoC also exceeded the baselines in both LM-execution categories.

  5. Knowl 5 — Ablations indicate that Python, language-model simulation, and intermediate state all contribute

    empirical result

    The BBH ablation study compared line-by-line interweaving with variants that either used only Python or only the LMulator, or attempted the whole program in Python before falling back to LM simulation. Scores are percentages; parentheses give the change from CoC (Interweave) for the same prompt setting.

    CoC variant Same-task prompting Cross-task prompting
    Interweave 84 61
    Try Python, then LM with state 82 (-2) 57 (-4)
    Try Python, then LM final answer 80 (-4) 60 (-1)
    Python only 48 (-36) 35 (-26)
    LMulator with state only 63 (-21) 49 (-12)
    LMulator final answer only 57 (-27) 50 (-11)

    Python-only execution performed poorly across the full benchmark, which includes semantic tasks that are difficult to implement as executable code. LMulator-only variants also fell below interweaving. Maintaining an intermediate state improved LMulator-only performance over producing only a final answer (63% versus 57% same-task; 49% versus 50% cross-task, so the cross-task comparison does not show an improvement). Whole-program Python execution with LM fallback was a simpler alternative with results close to interweaving on these aggregate scores, though it cannot handle cases that require line-level switching between executable operations and semantic judgments.

  6. Knowl 6 — CoC’s advantage over direct answering appears at smaller model sizes as well as larger ones

    empirical result

    Across the evaluated OpenAI model sizes, from text-ada-001 through text-davinci-003, CoC’s performance gains increased as model size increased. Unlike Chain of Thought prompting, which the paper reports improved over direct answering only for the largest model, CoC exceeded direct answering for the smaller models as well. The authors interpret this as evidence that smaller models can benefit from producing structured code as intermediate reasoning, rather than relying only on natural-language reasoning. The paper reports this scaling trend qualitatively rather than giving numerical scores for each model size in the text.

  7. Knowl 7 — CoC also improves results with instruction-tuned chat models

    empirical result

    The authors evaluated instruction-tuned GPT chat models in two settings. In a zero-shot setup, models received a generic instruction to answer directly, reason step by step, or write helpful Python code; generated code was either executed by Python or simulated by an LMulator. In a separate few-shot setup, a system instruction asked the models to follow a completion-model format, and prompts included three same-task or cross-task examples. The reported percentages were:

    Zero-shot model and method Direct CoT CoC Python CoC LM
    gpt-3.5-turbo 51 56 56 45
    gpt-4 70 78 82 75
    Model Few-shot prompt setting Direct CoT CoC
    gpt-3.5-turbo Same task 47 73 79
    gpt-3.5-turbo Cross task 47 60 61
    gpt-4 Same task 69 88 91
    gpt-4 Cross task 67 81 84

    In the few-shot setting, CoC outperformed direct answering and CoT for both models and both prompt settings. In zero-shot prompting, the Python-only or LMulator-only variants did not uniformly outperform CoT; these results should be distinguished from the few-shot interweaving results.

  8. Knowl 8 — CoC results vary little across independently authored prompts on four BBH tasks

    empirical result

    To test prompt robustness, three Python-familiar annotators independently authored CoC prompts for Date Understanding, Logical Deduction, Object Counting, and Penguins in a Table. Same-task evaluation used three examples from the evaluated task; cross-task evaluation used three examples from the other task families. The per-task scores and four-task averages were:

    Prompt setting Annotator Date Understanding Logical Deduction Object Counting Penguins in a Table Average
    Same task A 73 64 92 78 77
    Same task B 68 54 95 88 76
    Same task C 69 43 90 89 73
    Cross task A 41 33 67 76 54
    Cross task B 48 29 78 88 61
    Cross task C 60 30 76 64 57

    The average ranged from 73% to 77% across annotators in same-task prompting, and from 54% to 61% in cross-task prompting. Individual task scores varied more than the same-task averages, supporting the authors’ conclusion that CoC did not require one highly specific prompt-writing style in this evaluation.

  9. Knowl 9 — A one-shot CoC prompt supports real-robot tasks mixing semantic decisions with robot APIs

    experimental setup

    The robotics evaluation used a tabletop scene, a UR5 arm with a vacuum gripper, and a wrist-mounted RGB-D camera. The available perception API, detect_objects(), returned object labels, probabilities, bounding boxes, and segmentation masks; its implementation queried GPT-4V for object descriptions and used Grounding-SAM for localization. The robot also provided pick_place(obj1, obj2) for a scripted pick-and-place action and say(sentence) for speech. The prompt contained one example task: serving a meal according to dietary restrictions. The seven test instructions concerned packing a vegan lunch, assembling a vegetarian sandwich, gathering peanut-butter-sandwich ingredients, preparing tomato-and-egg stir-fry in a pot using Chinese text, sorting paper objects into a color-described container, sorting objects into compost and recycling bins, and helping with a bland steak. CoC used Python for perception and robot actions and the LMulator for semantic judgments such as whether an object was compostable. The authors report successful execution and generalization to new objects, a novel language, and different task domains; they state that line-by-line interweaving was the only evaluated approach capable of these tasks because the robot APIs and semantic judgments had to be alternated.

  10. Knowl 10 — CoC incurs execution costs and has implementation and deployment constraints

    limitation

    The authors identify additional context length and computation time as costs of generating code and then executing it, particularly when Python and LMulator operations are interwoven. They report that code does not help every semantic task; Ruin Names, which asks whether a name edit is humorous, is an example where performance can be worse. Their implementation stores program state as strings and parses it into built-in Python types, so the LMulator cannot modify custom Python objects as implemented. They suggest serialization and deserialization as a possible route to supporting such objects. Separately, the paper warns that executing model-generated Python as though it were benign requires safeguards against maliciously prompted code before real-world deployment.

Coverage note — Related work and background were excluded; no substantial contributed methods, experiments, results, or stated limitations were deliberately omitted.

References

  1. 1.Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  2. 2.Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Gianinazzi, L., Gajda, J., Lehmann, T., Podstawski, M., Niewiadomski, H., Nyczyk, P., et al. Graph of thoughts: Solving elaborate problems with large language models. arXiv preprint arXiv:2308.09687, 2023.
  3. 3.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  4. 4.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  5. 5.Chen, W., Ma, X., Wang, X., and Cohen, W. W. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588, 2022.
  6. 6.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  7. 7.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  8. 8.Dohan, D., Xu, W., Lewkowycz, A., Austin, J., Bieber, D., Lopes, R. G., Wu, Y., Michalewski, H., Saurous, R. A., Sohl-Dickstein, J., et al. Language model cascades. arXiv preprint arXiv:2207.10342, 2022.
  9. 9.Drori, I., Zhang, S., Shuttleworth, R., Tang, L., Lu, A., Ke, E., Liu, K., Chen, L., Tran, S., Cheng, N., et al. A neural network solves, explains, and generates university math problems by program synthesis and few-shot learning at human level. Proceedings of the National Academy of Sciences, 119(32):e2123433119, 2022.
  10. 10.Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G. Pal: Program-aided language models. In International Conference on Machine Learning, pp. 10764–10799. PMLR, 2023.
  11. 11.Gemini Team, G. Gemini: A family of highly capable multimodal models. Technical report, Google, 2023. URL https://storage.googleapis.com/deepmind-media/gemini/gemini_1_report.pdf.
  12. 12.Google, Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  13. 13.Gou, Z., Shao, Z., Gong, Y., Shen, Y., Yang, Y., Duan, N., and Chen, W. Critic: Large language models can self-correct with tool-interactive critiquing. arXiv preprint arXiv:2305.11738, 2023.
  14. 14.Khot, T., Trivedi, H., Finlayson, M., Fu, Y., Richardson, K., Clark, P., and Sabharwal, A. Decomposed prompting: A modular approach for solving complex tasks. arXiv preprint arXiv:2210.02406, 2022.
  15. 15.Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., and Girshick, R. Segment anything. arXiv:2304.02643, 2023.
  16. 16.Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213, 2022.
  17. 17.Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., et al. Solving quantitative reasoning problems with language models, 2022. arXiv preprint arXiv:2206.14858, 2022. URL https://arxiv.org/abs/2206.14858.
  18. 18.Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al. Competition-level code generation with alphacode. Science, 378(6624):1092–1097, 2022.
  19. 19.Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A. Code as policies: Language model programs for embodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pp. 9493–9500. IEEE, 2023.
  20. 20.Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., Zhu, J., et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023.
  21. 21.Mialon, G., Dessì, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Rozière, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., et al. Augmented language models: a survey. arXiv preprint arXiv:2302.07842, 2023.
  22. 22.Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837, 2022.
  23. 23.Mirchandani, S., Xia, F., Florence, P., Ichter, B., Driess, D., Arenas, M. G., Rao, K., Sadigh, D., and Zeng, A. Large language models as general pattern machines. arXiv preprint arXiv:2307.04721, 2023.
  24. 24.Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., Savarese, S., and Xiong, C. Codegen: An open large language model for code with multi-turn program synthesis. arXiv preprint arXiv:2203.13474, 2022.
  25. 25.Ning, X., Lin, Z., Zhou, Z., Yang, H., and Wang, Y. Skeleton-of-thought: Large language models can do parallel decoding. arXiv preprint arXiv:2307.15337, 2023.
  26. 26.Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., et al. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114, 2021.
  27. 27.OpenAI. Gpt-4 technical report, 2023.
  28. 28.Paranjape, B., Lundberg, S., Singh, S., Hajishirzi, H., Zettlemoyer, L., and Ribeiro, M. T. Art: Automatic multi-step reasoning and tool-use for large language models. arXiv preprint arXiv:2303.09014, 2023.
  29. 29.Parisi, A., Zhao, Y., and Fiedel, N. Talm: Tool augmented language models. arXiv preprint arXiv:2205.12255, 2022.
  30. 30.Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., et al. Toolllm: Facilitating large language models to master 16000+ real-world apis. arXiv preprint arXiv:2307.16789, 2023.
  31. 31.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  32. 32.Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761, 2023.
  33. 33.Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A. Progprompt: Generating situated robot task plans using large language models. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pp. 11523–11530. IEEE, 2023.
  34. 34.Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022.
  35. 35.Surís, D., Menon, S., and Vondrick, C. Vipergpt: Visual inference via python execution for reasoning. arXiv preprint arXiv:2303.08128, 2023.
  36. 36.Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
  37. 37.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  38. 38.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023a.
  39. 39.Wang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R. K.-W., and Lim, E.-P. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. arXiv preprint arXiv:2305.04091, 2023b.
  40. 40.Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.
  41. 41.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022a.
  42. 42.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022b.
  43. 43.Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X. Large language models as optimizers. arXiv preprint arXiv:2309.03409, 2023.
  44. 44.Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2022.
  45. 45.Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601, 2023.
  46. 46.Zeng, A., Attarian, M., Ichter, B., Choromanski, K., Wong, A., Welker, S., Tombari, F., Purohit, A., Ryoo, M., Sindhwani, V., et al. Socratic models: Composing zero-shot multimodal reasoning with language. arXiv preprint arXiv:2204.00598, 2022.
  47. 47.Zhou, A., Wang, K., Lu, Z., Shi, W., Luo, S., Qin, Z., Lu, S., Jia, A., Song, L., Zhan, M., et al. Solving challenging math word problems using gpt-4 code interpreter with code-based self-verification. arXiv preprint arXiv:2308.07921, 2023.
  48. 48.Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q., et al. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625, 2022a.
  49. 49.Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A. M., Le, Q. V., Laudon, J., et al. Mixture-of-experts with expert choice routing. Advances in Neural Information Processing Systems, 35:7103–7114, 2022b.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/