Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Charlie SnellJaehoon LeeKelvin XuAviral Kumar

article2024arXiv2,174 citations

Demonstrates that optimizing inference-time computation per prompt enables smaller language models to outperform models fourteen times their size, establishing an adaptive scaling strategy that exceeds standard best-of-N sampling efficiency by more than fourfold.

Listen

The rapid development of large language models has traditionally relied on massive pretraining compute to improve reasoning capabilities. However, pretraining increasingly larger foundation models incurs substantial financial, infrastructure, and deployment costs. A critical question for artificial intelligence deployment is whether spending additional computation at test timeallowing a model to "think longer" during inferencecan serve as a more effective and flexible substitute for scaling up model parameter size.

The article systematically analyzes how to optimize inference-time computation across two core mechanisms: searching against process-based reward models (step-by-step verifiers) and adaptively refining the model's proposal distribution through iterative revisions. The research evaluates how prompt difficulty influences the efficacy of test-time scaling strategies and examines whether additional test-time computation can match or exceed the performance of a substantially larger pretrained model under equivalent computational budgets.

The researchers conducted empirical evaluations on the challenging MATH benchmark using PaLM 2-S* base models fine-tuned specifically to verify solution steps or sequentially revise previous attempts. The study compared parallel sampling baselines against iterative revision chains and tree-search methods like beam search and lookahead search. To allocate compute efficiently, the authors established a compute-optimal framework that predicts question difficulty using a learned verifier and assigns the best-performing search or revision hyper-parameters per difficulty tier across cross-validated test splits.

The analysis yielded several key findings. First, adapting the compute strategy based on problem difficulty improves efficiency by up to fourfold compared to standard best-of-N sampling baselines. Second, the optimal mechanism depends heavily on problem difficulty: easier problems benefit most from purely sequential revisions, whereas harder problems require a balanced mix of parallel exploration and tree search to find valid solution paths. Third, complex tree search methods can over-optimize and exploit verifier flaws on easy problems at high generation budgets, leading to performance degradation. Fourth, in floating-point operations (FLOPs)-matched evaluations, allocating test-time compute to a smaller model outperformed a 14-fold larger pretrained model on easy and intermediate problems, especially in deployment regimes where inference volume is low relative to pretraining data. However, for the most difficult problems, test-time compute gains plateaued, indicating that pretraining scale remains necessary when a model lacks underlying capability.

These findings demonstrate that pretraining compute and test-time compute are not interchangeable one-to-one, but test-time compute offers a compelling path to reduce deployment overhead. Organizations can deploy smaller, highly optimized models on-device or at lower serving costs for routine to moderately complex reasoning tasks, reserving massive foundation models for the most challenging domains. This provides practical leverage to reduce training expenditures and shorten model delivery timelines.

Decision-makers should consider adopting adaptive inference frameworks that allocate computational budgets dynamically according to prompt difficulty rather than using fixed parallel sampling. Future initiatives should focus on lightweight methods for estimating prompt difficulty at runtime to minimize overhead, exploring hybrid search-and-revision methods, and distilling high-quality test-time rollouts back into base models to enable iterative self-improvement.

The study's conclusions are bounded by its focus on mathematical reasoning benchmarks and the reliance on capability-specific fine-tuning for revisions and verification. Furthermore, calculating difficulty metrics introduces extra inference overhead that must be balanced in production. While confidence is high that compute-optimal test-time scaling delivers major efficiency gains on structured reasoning tasks within a model's foundational scope, caution is advised when applying these methods to open-domain tasks or problems entirely beyond the base model's knowledge base.

Cover for Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Abstract

Enabling LLMs to improve their outputs by using more test-time computation is a critical step towards building generally self-improving agents that can operate on open-ended natural language. In this paper, we study the scaling of inference-time computation in LLMs, with a focus on answering the question: if an LLM is allowed to use a fixed but non-trivial amount of inference-time compute, how much can it improve its performance on a challenging prompt? Answering this question has implications not only on the achievable performance of LLMs, but also on the future of LLM pretraining and how one should tradeoff inference-time and pre-training compute. Despite its importance, little research attempted to understand the scaling behaviors of various test-time inference methods. Moreover, current work largely provides negative results for a number of these strategies. In this work, we analyze two primary mechanisms to scale test-time computation: (1) searching against dense, process-based verifier reward models; and (2) updating the model's distribution over a response adaptively, given the prompt at test time. We find that in both cases, the effectiveness of different approaches to scaling test-time compute critically varies depending on the difficulty of the prompt. This observation motivates applying a "compute-optimal" scaling strategy, which acts to most effectively allocate test-time compute adaptively per prompt. Using this compute-optimal strategy, we can improve the efficiency of test-time compute scaling by more than 4x compared to a best-of-N baseline. Additionally, in a FLOPs-matched evaluation, we find that on problems where a smaller base model attains somewhat non-trivial success rates, test-time compute can be used to outperform a 14x larger model.

Table of Contents

  • 1 Introduction
  • 2 A Unified Perspective on Test-Time Computation: Proposer and Verifier
  • 3 How to Scale Test-Time Computation Optimally
  • 3.1 Test-Time Compute-Optimal Scaling Strategy
  • 3.2 Estimating Question Difficulty for Compute-Optimal Scaling
  • 4 Experimental Setup
  • 5 Scaling Test-Time Compute via Verifiers
  • 5.1 Training Verifiers Amenable to Search
  • 5.2 Search Methods Against a PRM
  • 5.3 Analysis Results: Test-Time Scaling for Search with Verifiers
  • 6 Refining the Proposal Distribution
  • 6.1 Setup: Training and Using Revision Models
  • 6.2 Analysis Results: Test-Time Scaling with Revisions
  • 7 Putting it Together: Exchanging Pretraining and Test-Time Compute
  • 8 Discussion and Future Work
  • References
  • A Related Work
  • B Additional Revision Results
  • C Unsupervised Difficulty Bins
  • D PRM Training Details
  • E Comparing PRM Aggregation Strategies
  • F Comparing PRM and ORM
  • G Prompting Details
  • H Revision Model Finetuning Details
  • I Revision Model Selection Criteria
  • J Revision Model Verifier Training
  • K ReSTEM\text{ReST}^{\text{EM}} Revision Model Experiments
  • L Revision Model Example Outputs
  • M PRM Beam Search Example Outputs

Knowls

  1. Knowl 1 — Formulation of Test-Time Compute-Optimal Scaling Strategy

    definition

    Let qq denote an input prompt and y(q)y^*(q) the ground-truth correct response. For a given method of utilizing inference compute (such as sequential answer revisions or tree search against a reward verifier), let θ\theta denote the test-time hyperparameters (such as beam width, lookahead depth, or the ratio of sequential revisions to parallel samples) and NN denote the total test-time computation budget measured in generation samples or tokens.

    Let Target(θ,N,q)\text{Target}(\theta, N, q) denote the distribution over natural language output token sequences induced by the model under hyperparameters θ\theta and budget NN for prompt qq. The test-time compute-optimal scaling strategy selects hyperparameters θq,y(q)(N)\theta^*_{q, y^*(q)}(N) that maximize expected task accuracy:

    θq,y(q)(N)=argmaxθEyTarget(θ,N,q)[1y=y(q)]\theta^*_{q, y^*(q)}(N) = \arg\max_\theta \mathbb{E}_{y \sim \text{Target}(\theta, N, q)} \left[\mathbf{1}_{y = y^*(q)}\right]

    where 1{}\mathbf{1}_{\{\cdot\}} is the indicator function. This formalizes prompt-adaptive allocation of inference compute rather than applying a static sampling or search configuration across all inputs.

  2. Knowl 2 — FLOPs-Matched Exchange Rate Between Model Scale and Test-Time Compute

    theoretical result

    To compare scaling pretraining compute via model parameter size against scaling test-time compute on a smaller model under an identical total floating-point operations (FLOPs) budget, pretraining and inference compute are approximated as:

    X=6NbaseDpretrain,Y=2NbaseDinferenceX = 6 N_{\text{base}} D_{\text{pretrain}}, \quad Y = 2 N_{\text{base}} D_{\text{inference}}

    where NbaseN_{\text{base}} is the number of parameters in the base model, DpretrainD_{\text{pretrain}} is the number of tokens used during pretraining, and DinferenceD_{\text{inference}} is the total number of tokens generated across all inference queries over the model's deployment lifetime. If model parameters are scaled by a factor of MM (yielding Nlarge=MNbaseN_{\text{large}} = M N_{\text{base}}), total compute increases by a factor of MM to M(X+Y)M(X + Y) when using greedy decoding.

    To match this total FLOPs budget using the smaller base model, its test-time inference compute can be increased by a multiplier factor of:

    Multiplier=M+3(DpretrainDinference)(M1)=M+3R(M1)\text{Multiplier} = M + 3 \left(\frac{D_{\text{pretrain}}}{D_{\text{inference}}}\right)(M - 1) = M + \frac{3}{R}(M - 1)

    where R=DinferenceDpretrainR = \frac{D_{\text{inference}}}{D_{\text{pretrain}}} is the ratio of lifetime inference tokens to pretraining tokens. When R1R \ll 1 (e.g., self-improvement pipelines or low-query settings), large test-time compute multipliers can be spent per query while remaining FLOPs-equivalent to pretraining a larger model. When R1R \gg 1 (high-throughput production serving), the allowable inference compute multiplier approaches MM.

  3. Knowl 3 — Prompt Difficulty Estimation for Compute-Optimal Scaling

    model/method

    Because the optimal test-time strategy θq,y(q)(N)\theta^*_{q, y^*(q)}(N) depends on the unknown ground truth y(q)y^*(q), the compute-optimal strategy is approximated as a function of discrete prompt difficulty bins. Questions are categorized into five difficulty quantiles based on base LLM capability:

    1. Oracle Difficulty: Quantiles are established by measuring the base model's pass@1 accuracy across 2048 samples per question against ground-truth correctness.
    2. Model-Predicted (Unsupervised) Difficulty: Quantiles are established without ground-truth labels by averaging the final-answer score produced by a learned verifier (such as a Process Reward Model) across 2048 samples generated by the base model for that prompt.

    For each difficulty bin independently, the hyperparameter configuration θ\theta that maximizes validation accuracy for a given compute budget NN is determined via two-fold cross-validation and deployed on test questions falling in that difficulty bin.

  4. Knowl 4 — Compute-Optimal Test-Time Scaling Efficiency vs Best-of-N Baseline

    empirical result

    Across mathematical reasoning problems in the MATH benchmark using PaLM 2-S*, dynamically selecting test-time compute strategies based on question difficulty improves compute efficiency by up to 4×4\times compared to a standard best-of-N baseline:

    • In process reward model (PRM) search, compute-optimal allocation (e.g., utilizing beam search on harder questions and best-of-N on easier questions) matches or exceeds the accuracy of PRM best-of-N with 64 generations while using only 16 generations.
    • In revision-based generation, compute-optimal allocation (varying the ratio of sequential revisions to parallel samples per difficulty bin) matches the accuracy of parallel best-of-N at 256 generations using only 64 generations.
    • Model-predicted difficulty bins derived entirely from unsupervised PRM score averages yield compute-optimal scaling curves that closely match the performance of oracle difficulty bins derived from ground-truth correctness.
  5. Knowl 5 — Process Reward Model Training via Monte Carlo Rollouts and Last-Step Aggregation

    model/method

    A process-based reward model (PRM) is trained to evaluate step-by-step reasoning without requiring human step annotations:

    1. For each training question, 16 solutions are sampled from the few-shot prompted base LLM (discarding solutions lacking a valid final answer).
    2. At each intermediate step tt, 16 Monte Carlo rollouts are executed to completion using the base LLM at temperature 0 to determine the empirical fraction y[0,1]y \in [0, 1] of rollouts that yield the correct final answer.
    3. The PRM is fine-tuned as a binary classifier using a soft binary cross-entropy loss: LPRM=(ylog(y^)+(1y)log(1y^))\mathcal{L}_{\text{PRM}} = - \left( y \log(\hat{y}) + (1 - y) \log(1 - \hat{y}) \right) where y^\hat{y} is the predicted probability of correctness at that step.

    For scoring a complete candidate solution at inference time, evaluating the PRM prediction solely at the final step outperforms standard step-wise aggregation methods (such as taking the product or minimum of probabilities across all steps). For selecting among multiple generated solutions, best-of-N weighted selection is applied by marginalizing verifier scores across all candidates with identical final answers and choosing the answer with the largest summed score.

  6. Knowl 6 — Iterative Revision Model Training via Edit-Distance Matched Trajectories

    model/method

    To train a language model to iteratively correct reasoning errors without executing multi-turn rollouts during training, offline revision trajectories are constructed:

    1. 64 independent candidate solutions are sampled in parallel from the base model at elevated temperature for each training problem.
    2. For each correct solution, a multi-turn trajectory is formed by prepending k{0,1,2,3,4}k \in \{0, 1, 2, 3, 4\} incorrect sampled solutions in context, where kk is drawn uniformly.
    3. To encourage the model to identify and fix specific errors rather than restarting from scratch, the immediately preceding incorrect attempt is chosen as the incorrect sample with the minimum character-level edit distance to the target correct solution; any remaining preceding slots are filled with randomly selected incorrect samples.
    4. The base model is fine-tuned via supervised fine-tuning (SFT) to output the correct response given the preceding sequence of incorrect attempts. At test time, longer revision sequences are generated using a sliding context window containing the 4 most recent attempts.
  7. Knowl 7 — Tradeoff Between Sequential Revisions and Parallel Sampling by Problem Difficulty

    empirical result

    Evaluating sequential answer refinement against independent parallel sampling reveals distinct trade-offs governed by prompt difficulty:

    • Sequential Superiority: Proposing solutions sequentially conditioned on prior attempts outperforms sampling independent parallel candidates when using either verifier selection or majority voting.
    • Difficulty-Dependent Compute Allocation: On easier problems (difficulty bins 1 and 2), allocating the entire compute budget to purely sequential revisions yields the highest accuracy. On harder problems (difficulty bins 3 and 4), optimal performance requires a mixture of parallel trajectories and sequential revisions within each trajectory (e.g., at a budget of 128 generations, allocating compute across multiple parallel chains of sequential revisions).
    • Revision Degradation: Naive unguided revision converts approximately 38% of correct answers back into incorrect answers on subsequent revision steps, necessitating verifier scoring or majority aggregation across all intermediate states in the chain to maintain performance.
  8. Knowl 8 — FLOPs-Matched Comparison Between Pretraining Scale and Test-Time Compute

    empirical result

    In a FLOPs-matched evaluation comparing a smaller base model (PaLM 2-S*) using compute-optimal test-time scaling against a 14×\sim 14\times larger base model evaluated with greedy decoding:

    • Low Inference Workload (R=Dinference/Dpretrain1R = D_{\text{inference}}/D_{\text{pretrain}} \ll 1): When inference volume is small relative to pretraining data (R=0.16R = 0.16), test-time compute on the smaller model outperforms the 14×14\times larger model across easy and medium questions (achieving +21.6% relative improvement on easy questions with revisions and +19.1% with PRM search).
    • High Inference Workload (R1R \gg 1): When inference volume is large (R=22R = 22), scaling pretraining parameters becomes more compute-efficient than test-time compute on all but the easiest questions.
    • Hard Questions (Difficulty Bins 4 and 5): On the hardest questions where base model pass@1 is close to zero, test-time compute yields minimal improvement, and spending FLOPs on pretraining a larger model is consistently superior (e.g., -37.2% relative performance for test-time compute vs pretraining on hard questions at R1R \gg 1). Test-time compute cannot fully substitute for pretraining scale on out-of-distribution or out-of-capability problems.
  9. Knowl 9 — Search Over-Optimization and Failure Modes in Process Verifiers

    empirical result

    When optimizing solutions against a process-based reward model (PRM) using beam search or lookahead search:

    • Diminishing Returns and Degradation: While beam search (M=4M=4 or M=NM=\sqrt{N}) outperforms best-of-N at low generation budgets on medium and hard problems, larger search budgets lead to performance degradation on easy problems due to verifier exploitation.
    • Pathological Generations: Verifier reward hacking manifests in search discovering degenerative solutions that score highly under the PRM, such as appending repetitive, low-information steps at the end of a derivation or collapsing into 1–2 trivial steps.
    • Lookahead Search Inefficiency: Lookahead search (performing kk-step rollouts at temperature 0 to score intermediate nodes) underperforms standard beam search at equal total sample budgets N×(k+1)N \times (k+1), as the compute spent on simulation rollouts outweighs the value accuracy gained.
  10. Knowl 10 — Degradation of Revision Models Trained with Online Reinforcement Learning (ReSTEM)

    limitation

    Fine-tuning revision models using online iterative self-training via ReSTEM (sampling 64 revision trajectories of maximum length 5 per problem on the training set, filtering at the first correct response, and fine-tuning the base model on correct solutions) severely degrades sequential revision capabilities at test time.

    As the proportion of sequential revisions is increased relative to parallel samples, test accuracy drops substantially. Online trajectory generation during iterative training exacerbates spurious correlations in the multi-turn context, preventing the model from acquiring generalizable error-correction behaviors compared to offline supervised training on edit-distance-filtered pairs.

Coverage note — Specific qualitative solution transcripts shown in the appendices (Figures 17–29) were omitted as they serve only as concrete anecdotal illustrations of the search pathologies and revision patterns formalized in the empirical knowls.

References

  1. 1.Training revision models with synthetic data. Coming soon, 2024.
  2. 2.C. Andrieu, N. De Freitas, A. Doucet, and M. I. Jordan. An introduction to mcmc for machine learning. 2003.
  3. 3.R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. E. Shafey, Y. Huang, K. Meier-Hellstern, G. Mishra, E. Moreira, M. Omernick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G. H. Abrego, J. Ahn, J. Austin, P. Barham, J. Botha, J. Bradbury, S. Brahma, K. Brooks, M. Catasta, Y. Cheng, C. Cherry, C. A. Choquette-Choo, A. Chowdhery, C. Crepy, S. Dave, M. Dehghani, S. Dev, J. Devlin, M. Díaz, N. Du, E. Dyer, V. Feinberg, F. Feng, V. Fienber, M. Freitag, X. Garcia, S. Gehrmann, L. Gonzalez, G. Gur-Ari, S. Hand, H. Hashemi, L. Hou, J. Howland, A. Hu, J. Hui, J. Hurwitz, M. Isard, A. Ittycheriah, M. Jagielski, W. Jia, K. Kenealy, M. Krikun, S. Kudugunta, C. Lan, K. Lee, B. Lee, E. Li, M. Li, W. Li, Y. Li, J. Li, H. Lim, H. Lin, Z. Liu, F. Liu, M. Maggioni, A. Mahendru, J. Maynez, V. Misra, M. Moussalem, Z. Nado, J. Nham, E. Ni, A. Nystrom, A. Parrish, M. Pellat, M. Polacek, A. Polozov, R. Pope, S. Qiao, E. Reif, B. Richter, P. Riley, A. C. Ros, A. Roy, B. Saeta, R. Samuel, R. Shelby, A. Slone, D. Smilkov, D. R. So, D. Sohn, S. Tokumine, D. Valter, V. Vasudevan, K. Vodrahalli, X. Wang, P. Wang, Z. Wang, T. Wang, J. Wieting, Y. Wu, K. Xu, Y. Xu, L. Xue, P. Yin, J. Yu, Q. Zhang, S. Zheng, C. Zheng, W. Zhou, D. Zhou, S. Petrov, and Y. Wu. Palm 2 technical report, 2023.
  4. 4.Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. Lasenby, R. Larson, S. Ringer, S. Johnston, S. Kravec, S. E. Showk, S. Fort, T. Lanham, T. Telleen-Lawton, T. Conerly, T. Henighan, T. Hume, S. R. Bowman, Z. Hatfield-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, and J. Kaplan. Constitutional ai: Harmlessness from ai feedback, 2022.
  5. 5.C. Blakeney, M. Paul, B. W. Larsen, S. Owen, and J. Frankle. Does your data spark joy? performance gains from domain upsampling at the end of training, 2024. URL https://arxiv.org/abs/2406.03476.
  6. 6.G. Chen, M. Liao, C. Li, and K. Fan. Alphamath almost zero: process supervision without process, 2024.
  7. 7.K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman. Training verifiers to solve math word problems, 2021.
  8. 8.Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch. Improving factuality and reasoning in language models through multiagent debate, 2023.
  9. 9.J. S. B. T. Evans. Heuristic and analytic processes in reasoning. British Journal of Psychology, 75(4): 451–468, 1984.
  10. 10.X. Feng, Z. Wan, M. Wen, S. M. McAleer, Y. Wen, W. Zhang, and J. Wang. Alphazero-like tree-search can guide large language model decoding and training, 2024.
  11. 11.L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig. Pal: Program-aided language models, 2023. URL https://arxiv.org/abs/2211.10435.
  12. 12.S. Goyal, Z. Ji, A. S. Rawat, A. K. Menon, S. Kumar, and V. Nagarajan. Think before you speak: Training language models with pause tokens, 2024. URL https://arxiv.org/abs/2310.02226.
  13. 13.D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset, 2021.
  14. 14.J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre. Training compute-optimal large language models, 2022.
  15. 15.J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song, and D. Zhou. Large language models cannot self-correct reasoning yet, 2023.
  16. 16.A. L. Jones. Scaling scaling laws with board games, 2021. URL https://arxiv.org/abs/2104.03113.
  17. 17.D. Kahneman. Maps of bounded rationality: Psychology for behavioral economics. The American Economic Review, 93(5):1449–1475, 2003.
  18. 18.D. Kahneman. Thinking, fast and slow. Farrar, Straus and Giroux, New York, first paperback edition edition, 2013.
  19. 19.L. Kocsis and C. Szepesv'ari. Bandit based monte-carlo planning. In European conference on machine learning, pages 282–293. Springer, 2006.
  20. 20.A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y. Wu, B. Neyshabur, G. Gur-Ari, and V. Misra. Solving quantitative reasoning problems with language models, 2022.
  21. 21.Y. Li, Z. Lin, S. Zhang, Q. Fu, B. Chen, J.-G. Lou, and W. Chen. Making large language models better reasoners with step-aware verifier, 2023.
  22. 22.H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe. Let’s verify step by step, 2023.
  23. 23.A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark. Self-refine: Iterative refinement with self-feedback, 2023.
  24. 24.N. McAleese, R. Pokorny, J. F. Cerón Uribe, E. Nitishinskaya, M. Trębacz, and J. Leike. Llm critics help catch llm bugs. OpenAI, 2024.
  25. 25.OpenAI. Gpt-4 technical report, 2024.
  26. 26.Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, S. Zhao, L. Hong, R. Tian, R. Xie, J. Zhou, M. Gerstein, D. Li, Z. Liu, and M. Sun. Toolllm: Facilitating large language models to master 16000+ real-world apis, 2023. URL https://arxiv.org/abs/2307.16789.
  27. 27.C. Qu, S. Dai, X. Wei, H. Cai, S. Wang, D. Yin, J. Xu, and J.-R. Wen. Tool learning with large language models: A survey, 2024. URL https://arxiv.org/abs/2405.17935.
  28. 28.Y. Qu, T. Zhang, N. Garg, and A. Kumar. Recursive introspection: Teaching foundation models how to self-improve. 2024.
  29. 29.N. Sardana and J. Frankle. Beyond chinchilla-optimal: Accounting for inference in language model scaling laws, 2023.
  30. 30.W. Saunders, C. Yeh, J. Wu, S. Bills, L. Ouyang, J. Ward, and J. Leike. Self-critiquing models for assisting human evaluators, 2022.
  31. 31.A. Setlur, S. Garg, X. Geng, N. Garg, V. Smith, and A. Kumar. Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold. arXiv preprint arXiv:2406.14532, 2024.
  32. 32.Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024.
  33. 33.A. Sharma, S. Keh, E. Mitchell, C. Finn, K. Arora, and T. Kollar. A critical evaluation of ai feedback for aligning large language models, 2024. URL https://arxiv.org/abs/2402.12366.
  34. 34.N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao. Reflexion: Language agents with verbal reinforcement learning, 2023.
  35. 35.A. Singh, J. D. Co-Reyes, R. Agarwal, A. Anand, P. Patil, X. Garcia, P. J. Liu, J. Harrison, J. Lee, K. Xu, A. Parisi, A. Kumar, A. Alemi, A. Rizkowsky, A. Nova, B. Adlam, B. Bohnet, G. Elsayed, H. Sedghi, I. Mordatch, I. Simpson, I. Gur, J. Snoek, J. Pennington, J. Hron, K. Kenealy, K. Swersky, K. Mahajan, L. Culp, L. Xiao, M. L. Bileschi, N. Constant, R. Novak, R. Liu, T. Warkentin, Y. Qian, Y. Bansal, E. Dyer, B. Neyshabur, J. Sohl-Dickstein, and N. Fiedel. Beyond human data: Scaling self-training for problem-solving with language models, 2024.
  36. 36.C. Snell, E. Wallace, D. Klein, and S. Levine. Predicting emergent capabilities by finetuning. Conference on Language Modeling 2024, 2024.
  37. 37.K. Stechly, M. Marquez, and S. Kambhampati. Gpt-4 doesn’t know it’s wrong: An analysis of iterative prompting for reasoning problems, 2023.
  38. 38.R. S. Sutton and A. G. Barto. Reinforcement learning: An introduction. Second edition, 2018.
  39. 39.G. Team. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024.
  40. 40.Y. Tian, B. Peng, L. Song, L. Jin, D. Yu, H. Mi, and D. Yu. Toward self-improvement of llms via imagination, searching, and criticizing, 2024.
  41. 41.H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom. Llama 2: Open foundation and fine-tuned chat models, 2023. URL https://arxiv.org/abs/2307.09288.
  42. 42.J. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins. Solving math word problems with process- and outcome-based feedback, 2022.
  43. 43.K. Valmeekam, M. Marquez, and S. Kambhampati. Can large language models really improve by self-critiquing their own plans?, 2023.
  44. 44.P. Villalobos and D. Atkinson. Trading off compute in training and inference, 2023. URL https://epochai.org/blog/trading-off-compute-in-training-and-inference. Accessed: 2024-07-03.
  45. 45.P. Wang, L. Li, Z. Shao, R. X. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui. Math-shepherd: Verify and reinforce llms step-by-step without human annotations, 2023.
  46. 46.R. Wang, E. Zelikman, G. Poesia, Y. Pu, N. Haber, and N. D. Goodman. Hypothesis search: Inductive reasoning with language models, 2024. URL https://arxiv.org/abs/2309.05660.
  47. 47.J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou. Chain-of-thought prompting elicits reasoning in large language models, 2023.
  48. 48.S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan. Tree of thoughts: Deliberate problem solving with large language models, 2023.
  49. 49.Z. Yuan, H. Yuan, C. Li, G. Dong, K. Lu, C. Tan, C. Zhou, and J. Zhou. Scaling relationship on learning mathematical reasoning with large language models, 2023.
  50. 50.E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman. Star: Bootstrapping reasoning with reasoning, 2022.
  51. 51.E. Zelikman, G. Harik, Y. Shao, V. Jayasiri, N. Haber, and N. D. Goodman. Quiet-star: Language models can teach themselves to think before speaking, 2024. URL https://arxiv.org/abs/2403.09629.

Citation

MLA
Snell, C., et al. “Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters”. arXiv, 2024, http://arxiv.org/abs/2408.03314v1.
APA
Snell, C., Lee, J., Xu, K., & Kumar, A. (2024). Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. arXiv. http://arxiv.org/abs/2408.03314v1
Chicago
Snell, C., J. Lee, K. Xu, and A. Kumar. 2024. “Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters”. arXiv. http://arxiv.org/abs/2408.03314v1.
Harvard
Snell, C. et al. (2024) “Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2408.03314v1.
Vancouver
1. Snell C, Lee J, Xu K, Kumar A (2024) Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. arXiv

BibTeX

@article{snell2024scaling,
  title = {Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters},
  author = {Snell, Charlie and Lee, Jaehoon and Xu, Kelvin and Kumar, Aviral},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2408.03314v1},
  eprint = {2408.03314}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/