Generative Verifiers: Reward Modeling as Next-Token Prediction
Lunjun Zhang$^{1,2}$, Arian Hosseini$^{1,3,}$, Hritik Bansal$^{1,4,}$, Mehran Kazemi$^{1}$, Aviral Kumar$^{1,5}$ and Rishabh Agarwal$^{1}$
$^{*}$ Core Contribution, $^{1}$ Google DeepMind, $^{2}$ University of Toronto, $^{3}$ Mila, $^{4}$ UCLA, $^{5}$ Carnegie Mellon University
Abstract
Verifiers or reward models are often used to enhance the reasoning performance of large language models (LLMs). A common approach is the Best-of-N method, where N candidate solutions generated by the LLM are ranked by a verifier, and the best one is selected. While LLM-based verifiers are typically trained as discriminative classifiers to score solutions, they do not utilize the text generation capabilities of pretrained LLMs. To overcome this limitation, we instead propose training verifiers using the ubiquitous next-token prediction objective, jointly on verification and solution generation. Compared to standard verifiers, such generative verifiers (GenRM) can benefit from several advantages of LLMs: they integrate seamlessly with instruction tuning, enable chain-of-thought reasoning, and can utilize additional test-time compute via majority voting for better verification. We demonstrate that GenRM outperforms discriminative, DPO verifiers, and LLM-as-a-Judge, resulting in large performance gains with Best-of-N, namely $5% \rightarrow 45.3%$ on algorithmic tasks and $73% \rightarrow 93.4%$ on GSM8K. In easy-to-hard generalization settings, we observe improvements of $28% \rightarrow 44.6%$ on MATH, and $37.9% \rightarrow 53.5%$ on MMLU abstract algebra. Furthermore, we find that training GenRM with synthetic verification rationales is sufficient to pick out subtle errors on math problems. Finally, we demonstrate that GenRM scales favorably with model size and test-time compute.
Corresponding author(s): [email protected], {aviralkumar, rishabhagarwal}@google.com
Executive Summary: Large language models often produce incorrect answers on reasoning problems such as math and algorithmic tasks, even when they appear confident. A common fix, called Best-of-N, generates many candidate solutions and uses a verifier to select the best one. Current verifiers are usually trained as simple classifiers that output a score, yet they leave the text-generation strengths of large models unused.
This work set out to test whether training verifiers with the standard next-token prediction objective would improve verification accuracy and unlock additional capabilities already present in language models.
The authors trained generative verifiers, called GenRM, on two families of tasks. For algorithmic string problems they used programmatically generated verification rationales; for grade-school math they used synthetic rationales produced by the same model that created the solutions. They compared GenRM and its chain-of-thought variant against standard discriminative reward models, direct preference optimization verifiers, and prompted LLM judges, measuring performance with Best-of-N selection on held-out test sets and on harder transfer tasks.
GenRM-CoT with majority voting raised the fraction of GSM8K problems solved from 73 % to 93.4 % and improved algorithmic tasks from roughly 5 % to 45 %. The same verifier, trained only on grade-school problems, lifted performance on high-school competition math from 28 % to 44.6 % and produced similar gains on college-level MMLU algebra items. Generative verifiers scaled more favorably with model size and with extra test-time compute than discriminative baselines, and jointly training the model to generate and verify solutions improved both capabilities.
These gains matter because they show that a single model can now perform both generation and verification without custom classifier heads or extra human annotation. The approach therefore lowers the cost of reliable reasoning and makes it easier to add verification to existing language-model pipelines.
The results support wider adoption of generative verification for coding, alignment, and open-ended tasks. Immediate next steps include testing process-level supervision, combining the method with reinforcement learning, and running controlled pilots on new domains to confirm the observed scaling trends hold. The main limitations are reliance on the quality of synthetic rationales and evaluation on a narrow set of reasoning benchmarks; confidence is high for the reported tasks but moderate for untested domains until further data are collected.
1. Introduction
Section Summary: Large language models often produce convincing but flawed reasoning, so a common fix is to generate many candidate solutions and use a verifier to pick the best one. Traditional verifiers simply assign numerical scores and therefore miss the chance to use the model's own text-generation abilities, such as step-by-step reasoning. The authors instead train a generative verifier (GenRM) that answers “Is the answer correct?” by producing a “Yes” or “No” token, optionally after first writing an explicit reasoning chain, and show that this approach yields large gains on math and algorithmic tasks while scaling favorably with extra computation.
![**Figure 1:** **Generative Verifiers outperform standard verification approaches** in terms of Best-of-N on reasoning tasks, with a fixed generator. Here, $\Delta$ represents the improvement in number of problems solved with Best-of-N using GenRM-CoT. GenRM-CoT leverages the generation capabilities of LLMs, enabling a finetuned verifier to utilize chain-of-thought verification to detect subtle reasoning errors. For algorithmic tasks, we report average performance using Gemma-2B on Last Letter Concat ([1]) and BBH Word Sorting ([2]). For math reasoning, we train Gemma2-9B verifiers on GSM8K and evaluate their performance on GSM8K test (middle) and easy-to-hard\* generalization on MATH500 ([3]). For math tasks, LLM-as-a-Judge utilizes Gemini 1.0 Pro, which we used for synthetic verification rationales for training. For each task, the generated solutions in Best-of-N are the same; the only difference is the verifier. Math tasks use model-generated verification rationales for training GenRM-CoT. Data will be released at: [https://sites.google.com/view/generative-reward-models](https://sites.google.com/view/generative-reward-models).](https://ittowtnkqtyixxjxrhou.supabase.co/storage/v1/object/public/public-images/de2q47nm/all_intro_plot_iclr_v2.png)
While large language models (LLMs) demonstrate remarkable capabilities, they often confidently make logical and factual mistakes ([4]). These mistakes pose a significant challenge for reasoning problems, where a single mistake can invalidate the solution. A common strategy to address this issue is Best-of-N ([5, 6]): the LLM generates N candidate solutions for a given problem, and a learned reward model, referred to as a "verifier", ranks these solutions and picks the most suitable one. The effectiveness of this strategy hinges on how accurate the verifier is, making it crucial to identify better approaches for training verifiers.
LLM-based verifiers for reasoning are typically trained as discriminative reward models (RMs) to assign numerical scores to candidate solutions, which is then used to classify them as correct or incorrect ([6, 3, 7]). However, this scoring approach does not utilize the text-generation capabilities that LLMs are fundamentally designed for. As a result, discriminative RMs miss out on the inherent strengths of generative LLMs, such as unified instruction tuning ([8]), chain-of-thought (CoT) reasoning ([1]), and utilizing additional inference-time computation for better performance ([9, 10]). While LLM-as-a-Judge ([11]), which simply prompts off-the-shelf generative LLMs, also offers the above advantages, it typically underperforms trained LLMs-based verifiers on reasoning tasks, which we also observe in Figure 1.

In this work, we propose training verifiers with next-token prediction, which we call GenRM, to leverage the text generation capabilities of LLMs (Figure 2). Concretely, to produce a numerical score for a solution, the verifier now uses a prompt such as `Is the answer correct?', and represents the score as the probability of a single text token (e.g., 'Yes' or 'No'). GenRM naturally supports CoT reasoning ([1]): it can be trained to reason explicitly by generating a verbalized rationale before predicting correctness using 'Yes' or 'No'token (Figure 3), assuming rationales are available during training. We can further boost verification accuracy of CoT verifiers using majority-voting ([9]): sampling multiple CoT rationales and calculating the average score of the 'Yes' token across all rationales, enabling the use of inference-time compute for verification. Moreover, GenRM's next-token prediction training enables unifying solution generation with verification, which has been difficult with DPO verifiers ([12, 13]), potentially improving verification through positive transfer from solution.
Our results show that GenRM outperforms discriminative RMs, LLM-as-a-Judge, and self-consistency on algorithmic string manipulation and math reasoning tasks (Figure 1). Best-of-N performance further improves with GenRM-CoT that uses majority-voting, nearly matching performance with oracle verifier on algorithmic tasks. On GSM8K, when using a Gemma2-9B GenRM-CoT verifier on solutions from Gemini 1.0 Pro, we observe an improvement from $73% \rightarrow 93.4%$ in terms of the number of problems solved, surpassing GPT-4 and Gemini 1.5 Pro. Furthermore, GenRM-CoT trained on grade-school math problems exhibit easy-to-hard generalization, solving 17% more high-school competition problems in MATH500 ([3]) with Best-of-32. Moreover, we find that generative verifiers scale more favorably than discriminative verifiers as we increase model capacity, and outperform LLM-as-a-Judge as we scale inference-time compute with majority voting. Overall, generative verifiers hold significant potential for improving the reasoning capabilities of LLMs.

2. Preliminaries
Section Summary: An autoregressive language model produces text by predicting one token at a time, with each token’s probability conditioned on the input and all prior tokens; these probabilities are obtained by applying a softmax with adjustable temperature to the model’s output scores. Standard supervised fine-tuning trains the model by minimizing the cross-entropy loss between its next-token predictions and the target tokens in a dataset of problems and solutions. To improve reasoning at inference time, the best-of-N method draws multiple candidate answers and selects the highest-scoring one with a verifier; such verifiers are typically trained as binary classifiers that output a correctness probability, while an alternative approach simply prompts an unmodified LLM to judge solutions.
An autoregressive language model generates an output sequence ${\mathbf{y}} = (y_1, y_2, \ldots, y_T)$ given a input context ${\mathbf{x}}$ (e.g., math problem) by predicting tokens one at a time, based on the previously generated tokens. Assuming that the language model is parameterized by $\theta$, the conditional probability distribution of generating a sequence ${\mathbf{y}}$ given context ${\mathbf{x}}$ is
$ p_{\theta}({\mathbf{y}} \mid {\mathbf{x}}) = \prod_{t=1}^T p_{\theta}(y_t \mid {\mathbf{x}}, y_{<t}) $
with the convention $y_{<1} = \emptyset$ and ${\mathbf{y}}{<t} = (y_1, y_2, \ldots, y{t-1})$. For ease of notation, we define $p_\theta(y_t \mid {\mathbf{x}}) := p_\theta(y_t \mid {\mathbf{y}}{<t}, {\mathbf{x}})$. For a vocabulary size $M$, the probability of predicting the $t$-th token $y_t$, $p\theta(y_t \mid {\mathbf{x}})$, is determined using a softmax with temperature $\gamma$ on logit scores $z$ of all the tokens: $p_\theta(y_t \mid {\mathbf{x}}) = \frac{\exp(z_t/\gamma)}{\sum_{i=1}^M \exp(z_i/\gamma)}, \quad \textrm{where}\ z_t = \mathrm{logit}\theta(y_t \mid {\mathbf{x}}, {\mathbf{y}}{<t})$. Higher values of temperature $\gamma$ introduce more randomness, while setting temperature $\tau=0$ makes the output deterministic, which corresponds to greedy decoding.
Next-token prediction is the typical approach for pre-training and fine-tuning LLMs. In particular, supervised fine-tuning (SFT) minimizes the cross-entropy loss between the model's predicted next token and the actual target token in a given sequence. Given a dataset ${\mathcal{D}} = {(x, y)}$ of input context ${\mathbf{x}}$ and target response ${\mathbf{y}}$, the SFT loss is given by:
$ {\mathcal{L}}{\text{SFT}}(\theta, {\mathcal{D}}) = -\mathbb{E}{({\mathbf{x}}, {\mathbf{y}}) \sim {\mathcal{D}}} \left[\sum_{t=1}^{| {\mathbf{y}}|} \log p_{\theta}(y_t \mid {\mathbf{x}}, {\mathbf{y}}_{<t}) \right].\tag{2} $
Best-of-N is a widely-used approach to improve the reasoning performance of LLMs ([6, 3]). Specifically, given a test problem, we sample N candidate solutions from a generator LLM. These candidates are then scored using a learned verifier or reward model, and the highest-scoring solution is selected as the final answer. A better verifier increases the chance of selecting the correct solution, improving test accuracy.
Discriminative Verifiers. The prevalent approach of training verifiers for reasoning domains is to fine-tune an LLM as a classifier on a dataset of correct and incorrect solutions generated from a fixed LLM, using the binary cross-entropy loss. To do so, these verifiers directly assign a numerical score $r_\theta({\mathbf{x}}, {\mathbf{y}}) \in [0, 1]$ to estimate the probability that a solution ${\mathbf{y}}$ is correct for a problem ${\mathbf{x}}$. As such, these verifiers do not utilize the text generation the capabilities of LLMs. Given a reward-modeling (RM) dataset ${\mathcal{D}}{RM} = {\mathcal{D}}{\text{incorrect}} \mathbin{\bigcup} {\mathcal{D}}_{\text{correct}}$, we train discriminative RMs as follows:
$ \begin{aligned} {\mathcal{L}}(\theta, {\mathcal{D}}{RM}) =& -\mathbb{E}{({\mathbf{x}}, {\mathbf{y}}^+) \sim {\mathcal{D}}{\text{correct}}} \left[\log r\theta({\mathbf{x}}, {\mathbf{y}}^+)\right] -\mathbb{E}{({\mathbf{x}}, {\mathbf{y}}^-) \sim {\mathcal{D}}{\text{incorrect}}} \left[\log (1 - r_\theta({\mathbf{x}}, {\mathbf{y}}^-)) \right], \nonumber\ \quad \text{where} \quad &r_\theta({\mathbf{x}}, {\mathbf{y}}) = \text{sigmoid}(z_{cls}), \quad \text{and} \quad z_{cls}=\mathrm{logit}_\theta(cls \mid {\mathbf{y}}, {\mathbf{x}}) \end{aligned} $
where ${\mathbf{y}}^+$ are correct and ${\mathbf{y}}^-$ are incorrect solutions, and $cls$ corresponds to a special vocabulary token. In this work, we always use a balanced data mixture between correct (${\mathcal{D}}{\text{correct}}$) and incorrect (${\mathcal{D}}{\text{incorrect}}$) problem-solution pairs.
LLM-as-a-Judge does not finetune a verifier from a pretrained LLM, but simply prompts the LLM to perform the task of verification or self-critique ([11, 14]). LLM-judge sometimes uses reference-guided grading, where the LLM is given a reference solution to compare to.
3. GenRM: Verification as Next-Token Prediction
Section Summary: GenRM reframes verification as a text generation task by training an LLM to output a "Yes" or "No" token that indicates whether a solution is correct, rather than producing a separate numerical score. This approach preserves the model's original generation capabilities and allows it to be trained jointly on both creating solutions and checking them, simply by mixing the two types of examples during fine-tuning. The method can also incorporate step-by-step reasoning before the final Yes/No decision and improve reliability at test time by averaging predictions across multiple generated reasoning chains.
Discriminative LLM-based ( do not utilize the text generation capabilities of pretrained LLMs. To address this issue, we propose training generative verifiers, which we call GenRM, using standard next-token (. To do so, GenRM represents solution correctness using the LLM's probability distribution over tokens, instead of predicting a separate numerical score. This keeps the generation abilities of GenRM intact as the verification decision is just another token, while also enabling several advantages that come for "free" with LLMs, such as unified training for solution generation and verification, chain-of-thought reasoning, and inference-time computation.
3.1 Direct Verifier
In its simplest form, GenRM predicts whether a solution is correct using a single 'Yes'or 'No' token (Figure 3, top). This can be done by maximizing $\log p_\theta(\text{Yes'} \mid ({\mathbf{x}}, {\mathbf{y}}^+))$ for correct solutions ${\mathbf{y}}^+$ and $\log p_\theta(\text{No'} \mid ({\mathbf{x}}, {\mathbf{y}}^-))$ for incorrect solutions ${\mathbf{y}}^-$. To do so, we minimize the SFT loss ( on the dataset ${\mathcal{D}}_{\mathrm{Direct}}$ containing problem-solution pairs and a Yes or 'No' verification token:
$
\boxed{
{\mathcal{D}}_{\mathrm{Direct}} = {({\mathbf{x}}, {\mathbf{y}}^+, {\mathbf{I}}), \text{Yes'}\} \mathbin{\bigcup} \{({\mathbf{x}}, {\mathbf{y}}^-, {\mathbf{I}}), \text{No'}}}, \quad {\mathbf{I}} = \text{`Is the answer correct (Yes/No)?'}
$
At inference, we use the likelihood of the 'Yes'token as the verifier's score for re-ranking solutions:
$ r_{\text{Direct}}({\mathbf{x}}, {\mathbf{y}}) = p_{\theta}(\text{Yes} \mid {\mathbf{x}}, {\mathbf{y}}, {\mathbf{I}}). $
This score takes into account the verifier's confidence about its correctness prediction, which reduces the chance of being wrong at test-time when using a binary 'Yes' or 'No' prediction.
3.2 Unifying Generation and Verification
GenRM seamlessly integrates reward modeling, which distinguishes between correct and incorrect solutions, with SFT for generating correct solutions. This can be done by simply changing the data mixture in the SFT ( to include both verification and generation tasks. Given a verification dataset ${\mathcal{D}}{\text{verify}}$, which can be ${\mathcal{D}}{\text{Direct}}$ or ${\mathcal{D}}_{\text{CoT}}$ (discussed below) of problems-solution pairs with correctness tokens (optionally with CoT rationales), GenRM minimizes the loss:
$ \boxed{ {\mathcal{L}}{{\text{GenRM}}}(\theta, {\mathcal{D}}{\text{verify}}) = {\mathcal{L}}{\text{SFT}}(\theta, {\mathcal{D}}{\text{verify}}) + \lambda {\mathcal{L}}{\text{SFT}}(\theta, {\mathcal{D}}{\text{correct}})},\tag{3} $
where $\lambda > 0$ is a hyperparameter that controls the mixture ratio between verification (${\mathcal{D}}{\text{verify}}$) and generating correct solutions (${\mathcal{D}}{\text{correct}}$). This unified training can improve verifier and generation performance via positive transfer between these two related tasks: how to generate a correct solution, and whether a solution is correct. By default, we train GenRM verifiers using the unified loss (.
3.3 Chain-of-Thought Verifiers (GenRM-CoT)
Since verification often involves nuanced reasoning, generative verifiers can naturally benefit from CoT ([1]). Specifically, we can generate intermediate reasoning steps or critique (CoT) before making a decision about the solution correctness, which may identify subtle reasoning errors missed by direct verifiers (Figure 3, bottom). To train CoT verifiers, we can minimize the SFT loss ${\mathcal{L}}{{\text{GenRM}}}$ on the dataset ${\mathcal{D}}{\text{CoT}}$ containing problem-solution pairs as inputs, and corresponding verification rationales ${\mathbf{v}}_\textbf{CoT}$ appended with a final question ${\mathbf{I}}$ and 'Yes'or 'No' token as targets:
$
\boxed{
{\mathcal{D}}\text{CoT} = {\left({\mathbf{x}}, {\mathbf{y}}^+, {\mathbf{I}}\textbf{CoT}\right), ({\mathbf{v}}_\textbf{CoT}, {\mathbf{I}}, \text{Yes'})\} \mathbin{\bigcup} \{\left({\mathbf{x}}, {\mathbf{y}}^-, {\mathbf{I}}_\textbf{CoT}\right), ({\mathbf{v}}_\textbf{CoT}, {\mathbf{I}}, \text{No'})}}
$
where ${\mathbf{I}}\textbf{CoT}=$ 'Let's verify step by step.'. Notably, these rationales can either be human or LLM-generated, both of which we explore in this work. During inference, we first generate a CoT rationale ${\mathbf{v}}{\textbf{CoT}}$ from GenRM-CoT and then use the probability of 'Yes' for assigning the correctness score:
$ \begin{aligned} r_{\text{CoT}}({\mathbf{x}}, {\mathbf{y}}) &= p_{\theta}(\text{Yes} \mid {\mathbf{x}}, {\mathbf{y}}, {\mathbf{I}}\textbf{CoT}, {\mathbf{v}}\textbf{CoT}, {\mathbf{I}}), \quad \text{where}\ \ {\mathbf{v}}{\textbf{CoT}} \sim p{\theta}(\cdot \mid {\mathbf{x}}, {\mathbf{y}}, {\mathbf{I}}_\textbf{CoT}), \end{aligned} $
Compared ( that only uses the instruction ${\mathbf{I}}$ to produce a score, the above CoT reward additionally conditions on ${\mathbf{I}}\textbf{CoT}$ and self-generated ${\mathbf{v}}\textbf{CoT}$ before getting a score via instruction ${\mathbf{I}}$.
Inference-time compute for CoT verifier. When sampling verification CoTs, the generative verifier can use different reasoning paths and yield different correctness probabilities for the same problem-solution pair. As such, we would like to marginalize out these reasoning paths to select the most consistent correctness answer ([9]). To do so, we use majority voting where we first generate $K$ verification CoT rationales, and average the CoT-verifier score for these rationales:
$ r_{\text{MajV@K}}({\mathbf{x}}, {\mathbf{y}}) = \dfrac{1}{K} \sum_{i=1}^{K} p_{\theta}\left(\text{Yes} \mid {\mathbf{x}}, {\mathbf{y}}, {\mathbf{I}}\textbf{CoT}, {\mathbf{v}}{\textbf{CoT}}^{(i)}, {\mathbf{I}}\right), \quad \text{where}\ \ {\mathbf{v}}{\textbf{CoT}}^{(i)} \sim p{\theta}(\cdot \mid {\mathbf{x}}, {\mathbf{y}}, {\mathbf{I}}_\textbf{CoT})\tag{4} $
Since individual verification rationales from CoT verifiers can have reasoning errors, majority voting can mitigate the impact of such errors by averaging correctness scores across multiple rationales. Importantly, this means that GenRM-CoT can leverage additional inference-time compute to improve its accuracy, which discriminative verifiers cannot do. Unless otherwise specified, we report GenRM-CoT performance based on majority voting with 32 votes, that is, $K=32$ (.
Synthetic Verification CoT Rationales for Training Verifying LLM solutions with human-generated rationales can become increasingly expensive and challenging as LLMs surpass human reasoning abilities. To address this challenge, we explore using synthetically-generated rationales on GSM8K. One naive approach is to simply use the 'Let's verify step by step' prompt given a problem-solution pair, and keep the generated rationales only when they accurately verify the correctness of a solution ([15, 16]). However, such rationales (after filtering based on final yes/no responses) are still often of poor quality, due to 50% accuracy from random guessing.
To improve the quality of synthetic rationales, we provide a reference solution in addition to the problem and solution to verify (see Table 3), making it easier for an LLM to point out any reasoning error in the provided solution. This idea is similar to reference-guidance grading ([11]). Here, a reference solution could be any model-generated solution that arrives at the correct final answer. After initial data generation, we then filter the synthetic rationales using their verification correctness. Note that we condition on a reference solution only to generate training data, but do not include it during actual finetuning of the verifier, so that there is no train/test mismatch.

4. Experiments
Section Summary: In the experiments, researchers test generative verifiers (GenRM) that rely on next-token prediction and chain-of-thought reasoning against standard discriminative verifiers and other baselines across algorithmic string-manipulation tasks and math problems. They train models on simpler examples and measure how well they generalize to longer or harder problems, using Best-of-N accuracy to see how effectively each verifier ranks multiple candidate solutions. The results show that GenRM with chain-of-thought reasoning outperforms or matches the baselines, often achieving strong performance with far fewer solutions and better detection of subtle errors.
In this section, we evaluate the efficacy of next-token prediction and chain-of-thought reasoning for verification compared to standard verification approaches. To this end, we compare GenRM and standard verifiers on a number of reasoning tasks to answer the following questions: (1) How does ${\text{GenRM}}$ compare to discriminative verifiers and other approaches? (2) Does unified training of GenRM improve generation and verification performance? (3) Can GenRM effectively utilize CoT reasoning to improve its performance? (4) How does GenRM scale with model size and inference-time compute?
Tasks. We focus on the following tasks and put details about data generation in Appendix A:
- Algorithmic reasoning. We use two difficult string manipulation tasks, namely Last Letter Concatenation ([1]) and Word Sorting from Big-Bench ([2]). We train verifiers on word lists of length 2, 3, 4, and evaluate their generalization on length 5, 6. Note that this is a case of length generalization* for the verification task.
- Math reasoning. We train grade-school math verifiers on the GSM8K dataset from [6] that popularized test-time verification. We evaluate these verifiers on the GSM8K test set as well as their easy-to-hard generalization* on much harder MATH dataset ([17]), using the same held-out set of 500 MATH problems as [3]. We also evaluated model performance on the mathematical tasks in MMLU ([18]) dataset.


Baselines. We compare GenRM to the following verification approaches:
- Discriminative RM ([6]) or ORM is the prevalent approach for training verifiers for test-time re-ranking on reasoning tasks (Section 2), and serves as our main baseline.
- LLM-as-a-Judge ([11]) uses an off-the-shelf pretrained LLM for verification. To do so, we use a CoT prompt to produce 32 verification rationales that is used for correctness prediction and pick the majority-vote correctness answer.
- DPO ([12]): Following [13], we use this preference optimization approach for training verifiers on preference pairs with incorrect and correct solutions.
- Self-consistency ([9]): A simple approach to use test-time compute without* verifiers: sample multiple solutions from the LLM generator and pick the most common answer.
Note that self-consistency and test-time verification are complementary approaches, and can be often combined via weighted self-consistency to further boost performance, as shown in Figure 8.
Evaluation protocol. Following [6, 3], we primarily use Best-of-N performance in terms of the percentage of problems solved using a fixed generator (Section 2) with learned verifiers, and report average accuracy on the test set. We also report test RM accuracy, which measures whether the verifier accurately classifies incorrect and correct solutions. While these two metrics are correlated, RM accuracy only evaluates the verifier's point-wise accuracy, while Best-of-N evaluates the verifier's ability to rank solutions for choosing the correct one.
Models & Training. For training verifiers, we use open-weights Gemma models ([19, 20]), specifically Gemma-2B for algorithmic tasks, and Gemma 2B, 7B, and Gemma-2 9B for GSM8K. For solution generation as well as LLM-as-a-Judge, we use Gemma 2B for algorithmic tasks and Gemini 1.0 Pro ([21]) for GSM8K. For verification CoT rationales, we generate oracle rationales for algorithmic tasks programmatically (Table 2); for GSM8K, we generate synthetic rationales using Gemini 1.0 Pro with reference-guided grading (Table 3). See Appendix B for hyperparameter details.
4.1 Generative Verifiers Outperform Standard Verification Approaches
GenRM outperforms LLM-as-a-Judge and DPO verifiers (Figure 1), while performing comparably or slightly better than discriminative verifiers (Figure 16). GenRM-CoT substantially improves the Best-of-N performance over GenRM. In particular, on the algorithmic tasks with oracle verification CoTs, GenRM-CoT nearly matches the oracle verifier performance.
On GSM8K, GenRM-CoT consistently outperforms other methods (Figure 5, middle), even though the synthetic CoT rationales for training may contain errors. Qualitatively, GenRM-CoT is able to detect subtle reasoning errors that are missed by discriminative or direct GenRM verifiers (see Figure 2, Figure 4, and Figure 15).
::: {caption="Table 1: Performance of different methods on math tasks from the MMLU ([18]) dataset. The evaluation uses an easy-to-hard generalization setting, where the verifier is trained only on grade school math. We highlight the absolute improvement of GenRM-CoT over the Disc-RM baseline. Notably, Gen-CoT demonstrates stronger performance across all tasks, with the improvements being more significant on harder tasks."}

:::

Easy-to-Hard Generalization. Without any training on MATH, GenRM-CoT results in a $6.4\times$ better sample efficiency than discriminative verifiers as we increase the number of solutions to verify, and surpasses the strong self-consistency baseline (Figure 5, right). While [22] demonstrate that discriminative verifiers trained on easy MATH problems can generalize to harder MATH problems, GenRM-CoT exhibits a much stronger generalization from grade-school* math problems tohigh-school competition problems in MATH (see Figure 6 for a score breakdown by subject areas and difficulty levels) and college-level math in MMLU (see Table 1).

Leveraging Self-Consistency with Verifiers. Self-consistency and test-time verification can be easily combined to boost Best-of-N performance. To do so, we use weighted self-consistency or majority-voting ([23, 24, 22]) where we weight each solution according to the verifier's score, and select the final answer with the largest weight (see Appendix C for details). Figure 8 shows that weighted SC can indeed improve the vanilla self-consistency (SC); in particular, weighted SC based on GenRM-CoT requires 2.5x fewer solutions than its counterpart based on Discriminative RM to reach the same performance.
4.2 Synergy Between Generation and Verification


Unifying solution generation with verification, as done by GenRM using next-token prediction, consistently improves verification performance across all tasks, as illustrated in Figure 9. This improvement is observed for both direct and CoT-based generative verifiers, suggesting that teaching the verifier to imitate correct solutions generally helps. However, adding too much solution generation data can decrease verification performance of GenRM (Figure 18).
Incorporating CoT verification data into the generator's training mix leads to better solution generation performance for the GenRM-CoT verifier itself, as evidenced in Figure 10 by the improved Best-of-N scores with the oracle verifier (Pass@N). This suggests that teaching a generator to perform CoT verification using next-token prediction can deepen its understanding of the generation process itself. Overall, unifying solution generation and verification is mutually beneficial.
4.3 Scaling Model Size and Inference-time Compute


Scaling Test-Time Compute with GenRM-CoT can be done by sampling multiple CoTs and applying majority voting, as described in (. As shown in Figure 11, GenRM-CoT verifier's performance scales gracefully with number of votes at test time, under all three Gemma model sizes (2B, 7B, 9B), outperforming greedy decoding performance within 2 votes. Notably, across model scales, the finetuned GenRM-CoT verifier outperforms LLM-as-a-Judge, which also utilizes the same CoT approach and number of majority votes, but prompts a more capable Gemini 1.0 Pro model than Gemma models which we finetune as verifiers.
Scaling model size. In Figure 12, we show that generative verifiers, especially GenRM-CoT, perform better than discriminative RMs across model sizes, both in terms of reward modeling accuracy and Best-of-N performance. Intuitively, bigger models are more capable of text generation, allowing GenRM-CoT finetuning to better tap into its chain-of-thought reasoning ability for verification. Furthermore, these results demonstrate that larger models generalize better using the same data, which matches what we expect from scaling model parameter counts under the next-token prediction loss.


4.4 Synthetic Rationales: Quantity and Quality Matter
Our results on math reasoning tasks indicate that CoT verifiers can outperform discriminative and direct verifiers without requiring human-written verification rationales, highlighting the potential of LLM-generated rationales. We find that both the quality and quantity of these synthetic rationales matter. As shown in Figure 13, using reference-guided grading during rationale generation (§ 3.3) significantly improves verification performance. Furthermore, using multiple rationales per solution also improves performance, as shown in Figure 14. We suspect that this is because model-generated rationales may contain errors, such that training on multiple rationales per solution can result in an "ensembling" effect that prevents overfitting to such errors ([25]).
Importantly, unlike prior work, our results on math reasoning tasks do not require a more capable model ([26, 27]) or humans ([28, 29]) for generating verification rationales: we use the same model (Gemini 1.0 Pro) to both generate solutions to verify and synthetic verification rationales for training.

5. Related Work
Section Summary: Traditional reward models and verifiers are typically trained as discriminative classifiers that output numerical scores for solution correctness or preferences, without leveraging language models' text-generation abilities. In contrast, approaches like GenRM frame verification as next-token prediction—such as generating "Yes" or "No" after chain-of-thought reasoning—and can be trained on synthetic data to outperform both untrained LLM judges and stronger off-the-shelf models. This generative setup also enables unifying solution generation and verification within a single model via simple next-token training, avoiding the limitations of methods like DPO that tie rewards implicitly to policy logits or rely on separate discriminative heads.
Reward models (RMs) and verifiers. Conventionally, RMs and verifiers are trained as discriminative models via binary classification: given a prompt and a corresponding solution or a pair of solutions), the model is either trained to predict the correctness of the solution ([6, 3, 7, 23, 30, 31]) or a preference between the two solutions ([32, 33]). Concretely, the RM directly produces a numerical continuous-valued score, which is then plugged into a classification (. As such, discriminative verifiers do not utilize the generation capabilities of LLMs. In contrast to discriminative RMs, GenRM represents the correctness decision using the log probability of specific tokens, for example 'Yes'and 'No'. Posing verification as generating "yet another token" allows it to tap better into the generation capabilities of LLMs, by making it straightforward to employ CoT reasoning and additional inference-time compute for better verification.
LLM-as-a-Judge. Another line of work that poses verification as next-token prediction simply prompts* off-the-shelf LLMs to act as a verifier when provided with a rubric and a template for grading ([11, 14, 34, 35]) or many-shot ICL examples ([36]), butwithout any specific training for the same. Perhaps unsurprisingly, we find in our experiments that using more powerful LLMs (Gemini 1.0 Pro) as a judge is worse than our trained GenRM using weaker Gemma models (Figure 1, Figure 11), highlighting the necessity of training* generative verifiers. Our generative verifiers also exhibit good out-of-distribution generalization, which might be due to better calibrated uncertainty estimates from training ([37]). More generally, even the strong proprietary LLMs, such as GPT-4 ([38]) and Gemini ([39]), fall behind trained RMs on popular leaderboards ([40]), and this gap is much larger for reasoning.
Using CoTs for reward models. Prior works have also used critiques or CoT to extract preference and verification signals using LLM-as-a-Judge ([41, 42, 43]); in contrast to these works, GenRM utilizes model-generated CoTs directly for training the verifier. Upon inference, a GenRM-CoT produces its own CoTs, which it then uses to make decisions on correctness, unlike [27] that simply uses CoTs from a separate highly-capable LLM. In contrast to prior work that utilizes high-quality data from humans to train critique models ([29]) or train discriminative* RMs for generating code critiques ([28]), we show that GenRM can be trained from purely synthetic, model-generated critiques. Concurrent work ([26]) trains an RM to produce response critiques for preference pairs generated using a much more capable LLM, which are then passed as input into a RM head, separate from the base LLM. Unlike GenRM which uses next-token prediction, their RM head is trained discriminatively akin to standard RMs. While this approach allows them to leverage CoT, it does not allow them to unify solution generation and verification as a result of a discriminative RM head, which GenRM seamlessly enables (Section 4.2). Moreover, their synthetic critiques are not filtered for correctness, which would lead to poor verification CoTs on reasoning tasks (§ 3.3).
Unified generation and verification. One of the hallmark properties of GenRM is that the same generative verifier can be co-trained with a generation (: when given a problem, the model is trained to produce a solution, whereas when given a problem and a candidate solution, it is trained to verify this candidate. This is related to DPO ([12]) and its application to learning verifiers in reasoning ([13]), which aims to unify generation (policy) and verification (reward models) by representing the reward implicitly using the logits of a policy and training the policy with a reward-modeling loss. For reasoning, this type of model tying has been shown to exhibit erroneous extrapolation and degradation in learned representations, which prior work has attempted to address with additional techniques ([44, 45, 46, 47]). Of these, while [47] train a reward model with an auxiliary generative SFT loss, note that this loss is applied on a separate head for regularization purposes and is discarded after training; unlike GenRM no text is produced when querying the RM. In addition, compared to DPO, GenRM uses a simpler next-token prediction loss, does not require a reference policy, and obtains significantly better verification performance (Figure 1, Figure 5).
6. Conclusion & Future Work
Section Summary: The paper introduces Generative Verifiers (GenRM), which treat answer checking as a next-word prediction task rather than a simple yes-no judgment. This method outperforms standard verifiers, supports step-by-step reasoning at test time, and merges answer generation with verification inside a single model, improving performance on both. The authors outline several next steps, such as applying the approach to coding and image generation, using reinforcement learning for training, and integrating it into broader reinforcement pipelines for language models.
In this paper, we have introduced Generative Verifiers (GenRM), which recast verification as next-token prediction. GenRM is more performant than discriminative verifiers, and unlocks the use of chain-of-thought reasoning and inference-time compute for better verification. GenRM also unifies generation and verification into a single LLM, and demonstrates that such a unification benefits both generation and verification. Moreover, we show that synthetic model-generated rationales, which can be error-prone, are sufficient to teach GenRM how to use verification CoT to pick out tricky errors on math reasoning tasks (see Figure 2, Figure 4, Figure 15, and Appendix E).
The framework of generative verification offers a solid foundation for future work. Promising directions include extending this framework to broader tasks such as coding, alignment, text-to-image generation ([48]), and open-ended generation ([49]). Furthermore, leveraging process-level supervision ([3]) and training CoT verifiers with reinforcement learning (RL) can result in more accurate generative verifiers. Given GenRM's compatibility with all the existing tools designed to improve LLMs, exploring enhancements through techniques like retrieval-augmented generation ([50]), many-shot learning ([36]), multi-staged prompting ([51]), and tool use ([52]) would be interesting. Finally, incorporating generative verifiers into RL pipelines for LLMs warrants further investigation.
Acknowledgements
Section Summary: The research was carried out while three of the authors were completing internships at Google. The writers thank a group of colleagues for offering feedback on an early draft of the paper and for taking part in useful conversations. They also recognize two other people whose help in preparing the computing setup made the Gemma experiments possible.
This work was done during LZ, AH, and HB's internship at Google. We thank Hugo Larochelle, Minqi Jiang, Aleksandra Faust, Ankit Anand, Guillaume Desjardins, Doina Precup, and Charlie Snell for feedback on an earlier version of this paper and informative discussions. We thank Chirag Nagpal and Katrin Tomanek for support in setting up infrastructure that was crucial for running Gemma experiments.
Author Contributions
Section Summary: LZ led the overall project, carried out nearly all the experiments and analysis, and wrote most of the paper. AH and HB handled specific baseline methods and tasks, while MK provided high-level advice. AK and RA together came up with the idea, guided the work, ran extra experiments, and contributed to the writing, with RA also implementing parts of the evaluation approach and hosting LZ during the research.
LZ led the project, and ran almost all of the experiments and ablation studies, and wrote and edited the paper. AH was responsible for the discriminative RM baselines and DPO baselines. HB was responsible for the word-sorting task, and helped set up evaluations and DPO. MK advised the project, and provided feedback on writing. AK conceived the project with RA, advised LZ, provided feedback on paper, and helped run additional experiments during rebuttal. RA hosted LZ as a student researcher, proposed several ideas and experiments, implemented reference-guided grading, ran experiments on additional datasets during rebuttal, wrote the initial draft and advised the project.
Appendix
Section Summary: The appendix describes methods for generating synthetic training data for verifiers across tasks such as last-letter concatenation, word sorting, and grade-school math problems, where model-generated solution attempts are paired with algorithmically created ground-truth verification chains of thought. It also details the hyperparameter choices and training procedures for generative verifiers, discriminative reward models, and DPO-based models, including learning rates, optimizers, data balancing, and regularization techniques. Additional notes cover data filtering to address imperfections in automated answer checking.
A. Training Data Generation for Verifiers
::: {caption="Table 2: Algorithmic reasoning tasks that we consider. In thes tasks, we can generate ground-truth verification chain-of-thoughts as the training data for a generative verifier. Those synthetic tasks help us understand whether a generative verifier can outperform a discriminative verifier in the ideal scenario* where there is no noise in the verification CoT training data."}

:::
- Last Letter Concatenation ([1]): Given a list of words, the task is to concatenate the last letters of each word (for instance, "Noah Paul Elisha Rebecca"* $\rightarrow$"hlaa"). To generate the training data, for each length ${2, 3, 4}$, we generate $350$ problem queries by randomly sampling from the set of words in original training set; for each problem query, we generate 128 attempts from Gemma-2B ([19]) model. This gives us a total of about 50K training data points after de-duplication. We train verifiers on examples of lengths ${2, 3, 4}$ (here the length refers to how many words are in the input list), and evaluate the verifier performance on length 6. We use the format in Table 2 to algorithmically generate ground-truth verification CoT for training.
- Word Sorting ([2]): Given a list of words, sort them in alphabetical order. We train verifiers on a dataset comprised of ${2, 3, 4}$ words in each example, and evaluate the performance on length $5$. For each length, we generate 4096 lists of words as the problem queries; for each problem, we generate 64 attempts from Gemma-2B. After de-duplication and filtering out invalid responses, we have a total of about 100K training data points. We also algorithmically generate ground-truth verification CoT for training (see Table 2).
- Grade School Math ([6]): We follow the original train/test split and use 1.3K problems for test, 128 problems for validation, and about 7.2K problems for training. We generate 50 solutions per problem, and randomly sample at max 16 correct solutions and 16 incorrect solutions per problem as the training set. We evaluate the verifier performance on 16 solutions per problem in the test set.
::: {caption="Table 3: We use model-generated rationales as CoT training data on GSM with the above prompt with Gemini 1.0 Pro. Specifically, we show the model another solution that arrives at the correct answer, which is privileged information that does not exist at test time. This does not require a more capable model: we use the same model to generate solutions and synthetic rationales in the training data."}

:::
::: {caption="Table 4: Zero-shot prompt for our LLM-as-a-Judge evaluation results based on Gemini 1.0 Pro."}

:::
B. Hyper-parameters for Verifier Training
For Gemma-based verifiers, we pick the best checkpoint based on validation accuracy of verification on held out problems and solutions. We always use data balancing between 50% correct solutions and 50% incorrect solutions in training.
GenRM verifiers
After doing a sweep of learning rates (LR), we find that an LR of $[2e-6, 1e-6, 5e-7]$ works well for our tasks considered (with LR= $2e-6$ generally being the best). We use a weight decay of $1e-2$, and do not apply any dropout. We use the Adam optimizer ([53]) with decoupled weight decay ([54]) and a gradient norm clipping of $1.0$. We use a linear warmup of $1000$ gradient steps, and a cosine decay schedule that decays to $10%$ of the peak learning rate after a decay period. We finetune for 300K steps with a batch size of 64, and use seqio ([55]) library to create data mixtures.
Discriminative RMs
We finetune Gemma-based discriminative RMs by using a special token's logit for classification. We chose the best performing ORM on our validation sets by launching a large sweep over learning rates $[1e-7, 5e-7, 1e-6, 2e-6, 3e-6, 5e-6]$, weight decay $[1e-3, 1e-2, 1e-1]$ and dropouts $[1e-3, 5e-3, 1e-2, 0]$. We also schedule the learning rate with a linear ramp up and a cosine decay. We use a Z-loss $=10^{-4}\cdot \log^{2}Z$ (where $Z$ is the softmax normalizer of all logits) for regularization purposes ([56, 57]). Results obtained with learning rate $1e-7$ and dropout= $0$.
DPO
We first finetune Gemma-based generative models using SFT on correct solutions to obtain a reference policy $\pi_{\text{ref}}$, and then initialize from this reference policy to train generator $\pi_{\text{DPO}}$ with the DPO loss on a dataset of pairs of correct and incorrect solutions. We conduct a hyper-parameter sweep for both the learning rate (LR) and the $\beta$ coefficient in DPO loss: for LR we sweeped $[1e-7, 5e-7, 1e-6, 2e-6]$ and found $1e-6$ to work best; for $\beta$ we considered $[0.01, 0.1, 0.5, 1.0, 2.0]$ and used $0.1$. After DPO is trained, instead of using $r=\log \pi_{\text{DPO}}(\text{solution}\mid \text{question}) - \log \pi_{\text{ref}}(\text{solution}\mid \text{question})$ as the score (as defined in DPO's derivation), we find that directly the sequence log probability of the final DPO policy $\log\pi_{\text{DPO}}(\text{solution}\mid \text{question})$ as the score (without subtracting the log prob from reference policy) results in better performance in verification (see Figure 20); a similar finding was also noted in ([13]).
C. Additional Details
Data filtering for synthetic verification CoT
Since the answer checker (either based on string matching or Sympy library ([58])) is not perfect, there will inevitably be false negatives in the model-generated solutions. Besides, it is possible for a solution to arrive at the right answer with an incorrect reasoning path, so there will also be false positives in solutions. We use the following strategy to mitigate the issue of false negatives and false positives: when selecting the synthetic verification rationales (generated under reference guidance) for training, we only keep the rationales from solutions where more than 50% of verification rationales agree with the correctness returned by the answer checker.
Weighted Self-Consistency
typically sums the verifier scores (across solutions) for each answer, and picks the answer with the highest summed scores. We find that summing the top-K scores (rather than summing all scores) for each answer slightly improves performance. This means that for each answer, we only consider the correctness of its top-K solutions. We use K=6 for GSM and K=4 for MATH.
D. Additional Results
Ablating generation loss weight ($\lambda$) in GenRM. Adding too much generation data negatively impacts verification, while intermediate values yield the best results, as shown in Figure 18. By default, all GenRM experiments use unified training for verification with solution (, with $\lambda=1/3$ for algorithmic tasks and $\lambda=1/4$ for GSM8K.
Data scaling for CoT verifiers. GenRM-CoT shows that the GenRM-CoT performance improves as we increase the number of solutions per problem from 8 to 32, in terms of RM accuracy and Best-of-N Accuracy, as shown in Figure 17.





E. Examples Verification rationales from GenRM-CoT: GSM8K Test and MATH500
::: {caption="Table 5: GenRM CoT Example 1"}

:::
::: {caption="Table 6: GenRM CoT Example 2"}

:::
::: {caption="Table 7: GenRM CoT Example 3"}

:::
::: {caption="Table 8: GenRM CoT Example 4"}

:::
::: {caption="Table 9: GenRM CoT Example 4 (Continued)"}

:::
::: {caption="Table 10: GenRM CoT Example 5"}

:::
::: {caption="Table 11: GenRM CoT Example 6"}

:::
::: {caption="Table 12: GenRM CoT Example 7"}

:::
::: {caption="Table 13: GenRM CoT Example 8"}

:::
::: {caption="Table 14: GenRM CoT Example 9"}

:::
::: {caption="Table 15: GenRM CoT Example 10"}

:::
::: {caption="Table 16: GenRM CoT Example 11"}

:::
::: {caption="Table 17: GenRM CoT Example 12"}

:::
::: {caption="Table 18: MATH (Transfer from GSM): GenRM-CoT Example 1"}

:::
::: {caption="Table 19: MATH (Transfer from GSM): GenRM-CoT Example 2"}

:::
::: {caption="Table 20: MATH (Transfer from GSM): GenRM-CoT Example 3"}

:::
References
Section Summary: This section compiles a numbered list of academic citations, primarily recent papers and preprints on large language models. The references focus on methods for improving reasoning, verification of outputs, and training techniques, along with related benchmarks and model evaluations. They represent the foundational sources drawn upon in the discussed research.
[1] J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022.
[2] M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, et al. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
[3] H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe. Let's verify step by step. arXiv preprint arXiv:2305.20050, 2023.
[4] Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y. Zhang, Y. Chen, et al. Siren's song in the ai ocean: a survey on hallucination in large language models. arXiv preprint arXiv:2309.01219, 2023.
[5] E. Charniak and M. Johnson. Coarse-to-fine n-best parsing and maxent discriminative reranking. In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL’05), pages 173–180, 2005.
[6] K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
[7] P. Wang, L. Li, Z. Shao, R. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui. Math-shepherd: A label-free step-by-step verifier for llms in mathematical reasoning. arXiv preprint arXiv:2312.08935, 2023.
[8] H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
[9] X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.
[10] B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. Ré, and A. Mirhoseini. Large language monkeys: Scaling inference compute with repeated sampling. arXiv preprint arXiv:2407.21787, 2024.
[11] L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36, 2024.
[12] R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2024.
[13] A. Hosseini, X. Yuan, N. Malkin, A. Courville, A. Sordoni, and R. Agarwal. V-star: Training verifiers for self-taught reasoners. arXiv preprint arXiv:2402.06457, 2024.
[14] Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022.
[15] A. Singh, J. D. Co-Reyes, R. Agarwal, A. Anand, P. Patil, P. J. Liu, J. Harrison, J. Lee, K. Xu, A. Parisi, et al. Beyond human data: Scaling self-training for problem-solving with language models. arXiv preprint arXiv:2312.06585, 2023.
[16] E. Zelikman, Y. Wu, J. Mu, and N. Goodman. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35:15476–15488, 2022.
[17] D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021.
[18] D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
[19] Gemma Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024a.
[20] Gemma Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, et al. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024b.
[21] G. T. Google, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
[22] Z. Sun, L. Yu, Y. Shen, W. Liu, Y. Yang, S. Welleck, and C. Gan. Easy-to-hard generalization: Scalable alignment beyond human supervision. arXiv preprint arXiv:2403.09472, 2024.
[23] J. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins. Solving math word problems with process-and outcome-based feedback. arXiv preprint arXiv:2211.14275, 2022.
[24] Y. Liu, A. Singh, C. D. Freeman, J. D. Co-Reyes, and P. J. Liu. Improving large language model fine-tuning for solving math problems. arXiv preprint arXiv:2310.10047, 2023.
[25] E. Zhang, V. Zhu, N. Saphra, A. Kleiman, B. L. Edelman, M. Tambe, S. M. Kakade, and E. Malach. Transcendence: Generative models can outperform the experts that train them. arXiv preprint arXiv:2406.11741, 2024.
[26] Z. Ankner, M. Paul, B. Cui, J. D. Chang, and P. Ammanabrolu. Critique-out-loud reward models. arXiv preprint arXiv:2408.11791, 2024.
[27] Z. Ye, F. Greenlee-Scott, M. Bartolo, P. Blunsom, J. A. Campos, and M. Gallé. Improving reward models with synthetic critiques. arXiv preprint arXiv:2405.20850, 2024.
[28] N. McAleese, R. M. Pokorny, J. F. C. Uribe, E. Nitishinskaya, M. Trebacz, and J. Leike. Llm critics help catch llm bugs. arXiv preprint arXiv:2407.00215, 2024.
[29] W. Saunders, C. Yeh, J. Wu, S. Bills, L. Ouyang, J. Ward, and J. Leike. Self-critiquing models for assisting human evaluators. arXiv preprint arXiv:2206.05802, 2022.
[30] L. Luo, Y. Liu, R. Liu, S. Phatale, H. Lara, Y. Li, L. Shu, Y. Zhu, L. Meng, J. Sun, et al. Improve mathematical reasoning in language models by automated process supervision. arXiv preprint arXiv:2406.06592, 2024.
[31] F. Yu, A. Gao, and B. Wang. Ovm, outcome-supervised value models for planning in mathematical reasoning. In Findings of the Association for Computational Linguistics: NAACL 2024, pages 858–875, 2024.
[32] N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021, 2020.
[33] R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021.
[34] S. Kim, J. Shin, Y. Cho, J. Jang, S. Longpre, H. Lee, S. Yun, S. Shin, S. Kim, J. Thorne, et al. Prometheus: Inducing fine-grained evaluation capability in language models. In The Twelfth International Conference on Learning Representations, 2023.
[35] Z. Ling, Y. Fang, X. Li, Z. Huang, M. Lee, R. Memisevic, and H. Su. Deductive verification of chain-of-thought reasoning. Advances in Neural Information Processing Systems, 36, 2024.
[36] R. Agarwal, A. Singh, L. M. Zhang, B. Bohnet, S. Chan, A. Anand, Z. Abbas, A. Nova, J. D. Co-Reyes, E. Chu, et al. Many-shot in-context learning. arXiv preprint arXiv:2404.11018, 2024.
[37] S. Kapoor, N. Gruver, M. Roberts, K. Collins, A. Pal, U. Bhatt, A. Weller, S. Dooley, M. Goldblum, and A. G. Wilson. Large language models must be taught to know what they don't know. arXiv preprint arXiv:2406.08391, 2024.
[38] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
[39] G. Team, M. Reid, N. Savinov, D. Teplyashin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv e-prints, pages arXiv–2403, 2024.
[40] N. Lambert, V. Pyatkin, J. Morrison, L. Miranda, B. Y. Lin, K. Chandu, N. Dziri, S. Kumar, T. Zick, Y. Choi, et al. Rewardbench: Evaluating reward models for language modeling. arXiv preprint arXiv:2403.13787, 2024.
[41] W. Yuan, R. Y. Pang, K. Cho, S. Sukhbaatar, J. Xu, and J. Weston. Self-rewarding language models. arXiv preprint arXiv:2401.10020, 2024.
[42] T. Wu, W. Yuan, O. Golovneva, J. Xu, Y. Tian, J. Jiao, J. Weston, and S. Sukhbaatar. Meta-rewarding language models: Self-improving alignment with llm-as-a-meta-judge. arXiv preprint arXiv:2407.19594, 2024.
[43] T. Wang, I. Kulikov, O. Golovneva, P. Yu, W. Yuan, J. Dwivedi-Yu, R. Y. Pang, M. Fazel-Zarandi, J. Weston, and X. Li. Self-taught evaluators. arXiv preprint arXiv:2408.02666, 2024.
[44] R. Y. Pang, W. Yuan, K. Cho, H. He, S. Sukhbaatar, and J. Weston. Iterative reasoning preference optimization. arXiv preprint arXiv:2404.19733, 2024.
[45] A. Setlur, S. Garg, X. Geng, N. Garg, V. Smith, and A. Kumar. Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold. arXiv preprint arXiv:2406.14532, 2024.
[46] A. Pal, D. Karkhanis, S. Dooley, M. Roberts, S. Naidu, and C. White. Smaug: Fixing failure modes of preference optimisation with dpo-positive. arXiv preprint arXiv:2402.13228, 2024.
[47] R. Yang, R. Ding, Y. Lin, H. Zhang, and T. Zhang. Regularizing hidden states enables learning generalizable reward model for llms. arXiv preprint arXiv:2406.10216, 2024.
[48] Z. Lin, D. Pathak, B. Li, J. Li, X. Xia, G. Neubig, P. Zhang, and D. Ramanan. Evaluating text-to-visual generation with image-to-text generation. arXiv preprint arXiv:2404.01291, 2024.
[49] M. Besta, L. Paleari, A. Kubicek, P. Nyczyk, R. Gerstenberger, P. Iff, T. Lehmann, H. Niewiadomski, and T. Hoefler. Checkembed: Effective verification of llm solutions to open-ended tasks. arXiv preprint arXiv:2406.02524, 2024.
[50] S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J.-B. Lespiau, B. Damoc, A. Clark, et al. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pages 2206–2240. PMLR, 2022.
[51] S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 36, 2024.
[52] T. Schick, J. Dwivedi-Yu, R. Dess`ı, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36, 2024.
[53] D. P. Kingma. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
[54] I. Loshchilov and F. Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
[55] A. Roberts, H. W. Chung, A. Levskaya, G. Mishra, J. Bradbury, D. Andor, S. Narang, B. Lester, C. Gaffney, A. Mohiuddin, C. Hawthorne, A. Lewkowycz, A. Salcianu, M. van Zee, J. Austin, S. Goodman, L. B. Soares, H. Hu, S. Tsvyashchenko, A. Chowdhery, J. Bastings, J. Bulian, X. Garcia, J. Ni, A. Chen, K. Kenealy, J. H. Clark, S. Lee, D. Garrette, J. Lee-Thorp, C. Raffel, N. Shazeer, M. Ritter, M. Bosma, A. Passos, J. Maitin-Shepard, N. Fiedel, M. Omernick, B. Saeta, R. Sepassi, A. Spiridonov, J. Newlan, and A. Gesmundo. Scaling up models and data with $t5x$ and $seqio$. arXiv preprint arXiv:2203.17189, 2022. URL https://arxiv.org/abs/2203.17189.
[56] A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023.
[57] M. Wortsman, P. J. Liu, L. Xiao, K. Everett, A. Alemi, B. Adlam, J. D. Co-Reyes, I. Gur, A. Kumar, R. Novak, et al. Small-scale proxies for large-scale transformer training instabilities. arXiv preprint arXiv:2309.14322, 2023.
[58] A. Meurer, C. P. Smith, M. Paprocki, O. Čert'ık, S. B. Kirpichev, M. Rocklin, A. Kumar, S. Ivanov, J. K. Moore, S. Singh, et al. Sympy: symbolic computing in python. PeerJ Computer Science, 3:e103, 2017.