Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements

Jiacheng LiuWenya WangDianzhuo WangNoah A. SmithYejin ChoiHannaneh Hajishirzi

article2023EMNLP65 citations

Presents VERA, a standalone commonsense verification model trained on millions of statements that outperforms systems like GPT-4 in estimating the plausibility of declarative claims and detecting errors in language model outputs.

Listen

Modern generative language models demonstrate remarkable capabilities across diverse tasks, yet they frequently generate text containing basic commonsense mistakes and lack built-in mechanisms to quantify uncertainty. Because these failures degrade user trust and pose operational risks in deployment, there is an urgent need for automated systems that can evaluate the commonsense plausibility of machine-generated statements.

The main objective of the article is to introduce and evaluate VERA, a general-purpose plausibility estimation model designed to estimate the commonsense correctness of declarative natural language statements without relying on external text retrieval.

To construct VERA, the researchers trained a 5-billion-parameter neural network using a two-stage process on approximately 7 million statements compiled from 19 question-answering datasets and two commonsense knowledge bases. The training pipeline integrated three distinct learning objectives: standard binary classification, multi-class ranking across statement groups, and supervised contrastive learning to separate similar true and false statements. To enhance robustness, the authors automatically augmented the training data with model-generated incorrect statements and implemented a post-training temperature calibration step to align confidence scores with actual correctness probabilities.

The evaluations yielded several key findings regarding model accuracy and utility. First, when applied to commonsense problem-solving benchmarks in a verification format, VERA achieved an average accuracy of 85.5% on seen datasets and 81.7% to 83.4% on unseen datasets, consistently outperforming leading general-purpose models such as GPT-3.5, ChatGPT, GPT-4, and Flan-T5. Second, when deployed to filter noisy knowledge generated by other language models, VERA enhanced the downstream accuracy boost of knowledge-augmented pipelines by 46% for GPT-3 and by 233% for Rainier. Third, in an evaluation on real-world errors generated by ChatGPT, VERA accurately identified mistakes with 91% precision and 74% recall, yielding an overall F1 score of 82%. Finally, performance scaling analyses revealed steady accuracy gains as model size increased, showing no signs of performance saturation at the 5-billion parameter scale.

These findings demonstrate that dedicated verification models can serve as effective, calibrated safety layers to supervise larger generative systems, reducing hallucination risks and improving automated reasoning quality. Unlike generative question-answering systems that require seeing multiple answer choices simultaneously, standalone verification models provide reliable confidence scores for individual declarative statements across diverse scientific, physical, and social domains.

Organizations developing or deploying generative language applications should consider integrating modular verification systems like VERA into their post-generation pipelines to filter knowledge and flag commonsense errors. When implementing such verifiers, system designers must account for trade-offs, such as the increased computational cost of evaluating multiple candidates individually. Further research and piloting are recommended to extend verification techniques to multi-sentence contexts, complex compositional claims, and domain-specific knowledge.

Readers should interpret these results with appropriate caution due to specific operational limitations. VERA is specifically designed for single-sentence commonsense assertions and is not trained to evaluate encyclopedic facts, long-form compositional text, or moral and ethical judgments. Additionally, while the model demonstrates strong general plausibility scoring, it remains susceptible to syntactic variations such as negations and paraphrases, and it is intended as a research prototype rather than an autonomous decision-making system.

arXiv: 2305.03695
  • Paper: CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge, Alon Talmor et al. (2019). This paper establishes the CommonsenseQA benchmark, one of the foundational commonsense question-answering datasets converted and utilized by Vera for plausibility verification training.
  • Paper: PIQA: Reasoning about Physical Commonsense in Natural Language, Yonatan Bisk et al. (2019). This work introduces the PIQA physical commonsense reasoning benchmark, providing key source datasets and formulations adapted by Vera into declarative plausibility statements.
  • Paper: FEVER: a Large-scale Dataset for Fact Extraction and VERification, James Thorne et al. (2018). This foundational paper formalizes large-scale automated claim verification and evidence entailment, establishing the core verification paradigms that Vera generalizes to open commonsense statements.
  • Paper: Training Verifiers to Solve Math Word Problems, Karl Cobbe et al. (2021). This study demonstrates the power of training dedicated verifier models to score and filter candidate language model outputs, inspiring Vera's retrospective verification architecture.
  • Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). This benchmark details how generative language models frequently imitate common human misconceptions and errors, framing the core challenge of detecting machine-generated implausibilities that Vera is designed to solve.
  • Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). This work analyzes factual consistency evaluation across multiple generation formats, establishing foundational evaluation criteria for standalone verification models.
  • Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). This paper examines probing language models for factual and commonsense relations, motivating Vera’s transition from cloze-based extraction to direct statement plausibility estimation.
Cover for Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements

Abstract

Today’s language models can be remarkably intelligent yet still produce text that contains trivial commonsense errors. Therefore, we seek a retrospective verification approach that can reflect on the commonsense plausibility of the machine text, and introduce VERA, a general-purpose model that learns to estimate the commonsense plausibility of declarative statements. To support diverse commonsense domains, VERA is trained on ~7M commonsense statements that are automatically converted from 19 QA datasets and two commonsense knowledge bases, and using a combination of three training objectives. When applied to solving commonsense problems in the verification format, VERA substantially outperforms existing models that can be repurposed for commonsense verification, even including GPT-3.5/ChatGPT/GPT-4, and it further exhibits generalization capabilities to unseen tasks and provides well-calibrated outputs. We find that VERA excels at filtering machine-generated commonsense knowledge and is useful in detecting erroneous commonsense statements generated by models like ChatGPT in real-world settings.

Table of Contents

  • 1 Introduction
  • 2 Problem Definition and Scope
  • 3 Method
  • 3.1 Data Construction
  • 3.1.1 From Commonsense QA Datasets
  • 3.1.2 From Commonsense KBs
  • 3.2 Model Training
  • 3.2.1 Model Architecture
  • 3.2.2 Batching
  • 3.2.3 Training Objectives
  • 3.2.4 Two-Stage Training
  • 3.3 Inference and Calibration
  • 4 Experimental Setup
  • 4.1 Training Details
  • 4.2 Evaluation and Baselines
  • 5 Evaluation Results
  • 5.1 Solving Multiple-Choice and Boolean Commonsense Problems
  • 5.2 Filtering LM-generated Commonsense Knowledge
  • 5.3 Preliminary Study on Detecting Commonsense Errors made by ChatGPT
  • 5.4 Analysis
  • 6 Related Work
  • 7 Conclusion and Future Work
  • Limitations
  • Acknowledgments
  • References
  • A More Details on Datasets
  • A.1 Dataset-Specific Special Handling
  • A.2 Conversion to Declarative Statements
  • B More Details on Method
  • B.1 Training Objectives
  • B.2 Calibration
  • C More Details on Experimental Setup
  • C.1 Definition of Metrics
  • C.2 Details on Baseline Models
  • D More Evaluation Results
  • E Further Analysis

Knowls

  1. Knowl 1 — VERA’s plausibility-estimation task and scope

    definition

    VERA is designed to estimate the plausibility of a standalone commonsense statement. An input statement xx must be natural-language, declarative rather than interrogative, self-contained, objectively classifiable as correct or incorrect, and judgeable using widely held commonsense knowledge about the world. Encyclopedic facts, such as national capitals, are outside the intended scope.

    The model outputs a real-valued score s(x)∈[0,1]s(x)\in[0,1]. A score of 11 represents complete confidence that xx is correct, and a score of 00 represents complete confidence that xx is incorrect. The default correctness boundary is s(x)=0.5s(x)=0.5, equivalently a zero classification logit.

  2. Knowl 2 — Large-scale construction of labeled commonsense statements

    model/method

    VERA is trained on approximately 7 million labeled statements assembled from 19 commonsense question-answering datasets and two commonsense knowledge bases. The data are organized into statement groups: statements derived from the same multiple-choice question or knowledge-base entry. A multiple-choice group contains one correct statement and one or more incorrect statements; a boolean question produces a one-statement group whose original yes/no label is retained.

    For multiple-choice data, each question and answer option is converted into a declarative statement using question-to-statement conversion, cloze substitution, continuation concatenation, or the option itself when no question is present. For boolean data, the question is converted into a statement using the “yes” option and its original label. The 19 QA datasets contribute approximately 200,000 groups and 400,000 statements. The two knowledge bases contribute approximately 1.6 million groups and 6 million statements: each correct knowledge-base entry is retained, and three incorrect variants are produced by replacing its subject with randomly selected knowledge-base subjects.

    To reduce overfitting to human-authored distractors, nine QA datasets are augmented with language-model falsehoods. For each multiple-choice question, a small language model samples 50 possible answers; the three least-probable answers whose generation probability is below 0.150.15 are added as incorrect statements. The authors selected the threshold through manual inspection and observed that answers above 0.150.15 were more likely to be plausible.

  3. Knowl 3 — VERA scoring architecture

    model/method

    Given a statement xx, VERA uses a pretrained transformer language model to encode the statement and takes the hidden representation of its final end-of-sequence token. This representation is passed through a linear projection and a sigmoid to obtain a plausibility score:

    h(x)=fLM(x),z(x)=wTh(x)+b,s(x)=σ(z(x))=11+e−z(x).h(x)=f_{\mathrm{LM}}(x),\qquad z(x)=w^{\mathsf T}h(x)+b,\qquad s(x)=\sigma(z(x))=\frac{1}{1+e^{-z(x)}}.

    Here h(x)∈Rdh(x)\in\mathbb{R}^{d} is the final end-of-sequence hidden vector, fLMf_{\mathrm{LM}} is the transformer encoder or decoder, w∈Rdw\in\mathbb{R}^{d} and b∈Rb\in\mathbb{R} are learned parameters, z(x)∈Rz(x)\in\mathbb{R} is the classification logit, and s(x)∈[0,1]s(x)\in[0,1] is the estimated plausibility. The final-token representation is used because it can summarize the complete input for both bidirectional encoders such as T5 and left-to-right decoders such as LLaMA.

  4. Knowl 4 — Joint binary, groupwise, and contrastive training objectives

    equation

    VERA trains on batches in which statements from the same statement group are kept together. For a statement xix_i with binary correctness label yi∈{0,1}y_i\in\{0,1\} and score s(xi)s(x_i), the binary loss is

    Lbin(xi,yi)=−yilog⁡s(xi)−(1−yi)log⁡(1−s(xi)).\mathcal{L}_{\mathrm{bin}}(x_i,y_i)=-y_i\log s(x_i)-(1-y_i)\log\bigl(1-s(x_i)\bigr).

    For a multiple-choice group Xj={xj1,…,xjCj}X_j=\{x_{j1},\ldots,x_{jC_j}\} with exactly one correct statement xj∗x_{j^*}, VERA also minimizes the groupwise multiclass loss

    Lmc(Xj)=−log⁡exp⁡z(xj∗)∑c=1Cjexp⁡z(xjc).\mathcal{L}_{\mathrm{mc}}(X_j)=-\log\frac{\exp z(x_{j^*})}{\sum_{c=1}^{C_j}\exp z(x_{jc})}.

    This objective forces the model to distinguish statements that share the same question or knowledge-base source but have different correctness labels. It is not applied to single-statement boolean groups.

    For supervised contrastive learning, let B\mathcal{B} be the set of statements in a batch, let P(i)={k∈B:k≠i, yk=yi}P(i)=\{k\in\mathcal{B}:k\neq i,\ y_k=y_i\} be the positive examples for anchor xix_i, and let N(i)={k∈B:yk≠yi}N(i)=\{k\in\mathcal{B}:y_k\neq y_i\} be its negative examples. With cosine similarity cos⁡(⋅,⋅)\operatorname{cos}(\cdot,\cdot) and temperature τ>0\tau>0, the contrastive loss is

    Lctr(xi)=−log⁡∑k∈P(i)exp⁡ ⁣(cos⁡(h(xi),h(xk))/τ)∑k∈P(i)∪N(i)exp⁡ ⁣(cos⁡(h(xi),h(xk))/τ).\mathcal{L}_{\mathrm{ctr}}(x_i)=-\log\frac{\sum_{k\in P(i)}\exp\!\left(\operatorname{cos}(h(x_i),h(x_k))/\tau\right)}{\sum_{k\in P(i)\cup N(i)}\exp\!\left(\operatorname{cos}(h(x_i),h(x_k))/\tau\right)}.

    The total training objective is a weighted combination L=αLbin+βLmc+γLctr\mathcal{L}=\alpha\mathcal{L}_{\mathrm{bin}}+\beta\mathcal{L}_{\mathrm{mc}}+\gamma\mathcal{L}_{\mathrm{ctr}}, with the paper’s settings α=1.0\alpha=1.0, β=1.0\beta=1.0, γ=0.1\gamma=0.1, and τ=0.05\tau=0.05. The binary loss is balanced within each statement group by averaging the contributions from statements with the same label, preventing multiple-choice groups with many incorrect options from dominating training.

  5. Knowl 5 — Two-stage training procedure and implementation scale

    algorithm

    The model is trained in two consecutive stages because knowledge-base data are much larger but noisier than QA-derived data.

    Input: a pretrained T5 encoder or LLaMA decoder, knowledge-base statement groups, QA-derived statement groups, and the three-loss objective.

    Output: a VERA plausibility estimator.

    Initialize the language model and the scoring layer.
    Stage A:
        Mix the statement groups from Atomic2020 and GenericsKB without reweighting.
        For 50,000 optimization steps:
            Sample a batch containing 64 statement groups.
            Truncate each statement to at most 128 tokens.
            Cap each training group at four statements.
            Compute the binary, multiclass, and supervised contrastive losses.
            Update all parameters with Adam.
    Stage B:
        Continue from the Stage A checkpoint.
        Mix the 19 QA-derived training datasets without reweighting.
        For 50,000 optimization steps:
            Sample and process batches using the same group, token, and loss rules.
            Update all parameters with Adam.
    Return the Stage B model.

    For the T5 encoder, the learning rate is 1×10−51\times10^{-5}; for LLaMA it is 2×10−62\times10^{-6}. The principal model, VERA-T5, starts from the approximately 5-billion-parameter T5-v1.1-XXL encoder. VERA-LLaMA starts from LLaMA-7B. The authors found this two-stage procedure better than training on either source alone or on a single mixture of both sources.

  6. Knowl 6 — Post-hoc temperature calibration of plausibility scores

    model/method

    The uncalibrated VERA model tends to be overconfident. During inference, VERA applies temperature scaling to its logit without changing the model parameters:

    z~(x)=z(x)T,scal(x)=σ(z~(x)),\tilde z(x)=\frac{z(x)}{T},\qquad s_{\mathrm{cal}}(x)=\sigma\bigl(\tilde z(x)\bigr),

    where z(x)z(x) is the original scoring logit, T>0T>0 is a learned scalar temperature, and scal(x)s_{\mathrm{cal}}(x) is the calibrated plausibility score. The temperature is selected on the combined development sets of the seen datasets to minimize expected calibration error, using 10 equal-sized confidence bins.

    Temperature scaling preserves the ordering of scores and the sign of logits. Consequently, it does not change multiple-choice argmax decisions, ROC or precision-recall rankings, or the binary decision boundary at z=0z=0 and s=0.5s=0.5; it only makes the numerical confidence estimates better aligned with observed correctness frequencies.

  7. Knowl 7 — Problem-solving performance on seen and unseen benchmarks

    data/table

    VERA applies to a multiple-choice problem by converting every answer option into a statement and selecting the statement with the highest plausibility score. For a boolean problem, it predicts correct when s≥0.5s\geq0.5. The evaluation distinguishes seen benchmarks, whose training data were used by VERA; unseen type 1 benchmarks with similar task structure; and unseen type 2 benchmarks with more distant structures, including contextualized event completion and entity-based reasoning. The reported aggregate is the unweighted mean development-set accuracy across benchmarks.

    The summary accuracy table on page 6 reports the following percentages:

    Could not parse LaTeX table

    VERA-T5 exceeds the strongest baseline, Flan-T5, by 6 percentage points on seen benchmarks, 4 points on unseen type 1, and 4 points on unseen type 2. Its expected calibration error is no higher than 3% across these benchmark groups after calibration. The paper notes that ChatGPT and GPT-4 results may be underestimated because their APIs did not expose raw token logits, while some unseen benchmarks were included in Flan-T5’s pretraining or instruction-tuning data.

  8. Knowl 8 — Filtering language-model-generated commonsense knowledge

    empirical result

    VERA can filter statements generated by another language model before they are supplied to a downstream commonsense QA system. In the evaluated pipeline, a knowledge generator produces candidate statements, VERA retains only statements with score above 0.50.5, and UnifiedQA-large answers the original question using the retained statements. The generators are few-shot GPT-3 (davinci) and Rainier-large; the evaluation uses the same settings as the Generated Knowledge Prompting framework.

    The pipeline results table on page 7 reports these development-set accuracies:

    Could not parse LaTeX table

    VERA is also the strongest filter on the two seen annotated-generation benchmarks, SKD_anno and I2D2_anno, in both AUROC and average precision; on I2D2_anno it improves AUROC by 2 percentage points over the task-specific I2D2 critic. On the unseen Rainier_anno benchmark, it is comparable to the strongest baselines. Calibration error is no higher than 8% on all three filtering benchmarks.

  9. Knowl 9 — Detecting commonsense errors in ChatGPT outputs

    empirical result

    In a preliminary in-the-wild evaluation, the authors collected 27 online anecdotes about commonsense errors produced by ChatGPT and manually rewrote each erroneous statement into a correct version, yielding 54 statements. VERA’s predictions on the incorrect statements achieved 91% precision, 74% recall, and an F1F_1 score of 82%.

    The example table on page 8 shows that VERA generally assigns low scores to the original erroneous statement and higher scores to its manually corrected counterpart: this pattern occurs in 7 of 9 displayed pairs. For example, the claim that a marble less dense than mercury would sink receives score 0.040.04, whereas the corrected claim that it would float receives 0.960.96. The analysis also exposes failures: VERA gives score 0.860.86 to the claim that a solar eclipse can be followed by a lunar eclipse the next day, and score 0.800.80 to the claim that a diagonal line can be drawn in a triangle, treating both as correct.

  10. Knowl 10 — Ablation and scaling evidence for VERA’s design

    empirical result

    The ablation study removes one training component at a time and evaluates development-set accuracy. Removing the supervised contrastive loss primarily damages performance on unseen datasets, supporting its role in generalization. Removing the knowledge-base pretraining stage reduces performance broadly, indicating that the large-scale commonsense knowledge-base data are useful despite their noise. Language-model-generated falsehoods help most on unseen benchmarks, with a small sacrifice on seen benchmarks. The multiclass loss is particularly useful for multiple-choice tasks, whereas removing the binary loss substantially harms boolean-task performance.

    The scaling experiment trains VERA with progressively larger T5 encoders, from approximately 40M parameters through 100M, 300M, 1B, and 5B. Accuracy increases steadily with model size on both seen and unseen type 1 benchmarks, with no evidence of saturation at 5B parameters. Within the tested models, VERA-T5 also consistently outperforms VERA-LLaMA, which the authors associate with the T5 encoder’s bidirectional connectivity.

  11. Knowl 11 — Operational limitations and verification-format tradeoffs

    limitation

    VERA is intended for objective, single-sentence commonsense statements and is not designed for encyclopedic facts, fictional-world reading comprehension, moral commonsense judgments, very long or highly compositional inputs, or contextualized and defeasible statements. It can still output a score for out-of-scope inputs because it has no explicit scope guard. The model is not fully robust to syntactic changes such as paraphrases and negations, may reflect bias or toxicity present in its training data, and is a research prototype rather than a system for real-world decision-making.

    The verification format also has task-level costs. Compared with generative QA, it can lose accuracy when the competing answer choices cannot be compared jointly, it cannot itself generate an answer, and it must score all CC choices separately rather than solving a CC-way question in one generation. In the authors’ comparison, a QA-format model trained on the same multiple-choice data exceeded VERA by 1.5 percentage points on seen multiple-choice benchmarks. Verification nevertheless supports direct classification of declarative statements and confidence estimation for answers produced by generative systems.

Coverage note — Detailed dataset-specific preprocessing, per-dataset score breakdowns, baseline implementation prompts, and the exact expected-calibration-error formula were omitted because they support the central method and headline results rather than constitute separate load-bearing contributions.

References

  1. 1.Stephane T Aroca-Ouellette, Cory Paik, Alessandro Roncone, and Katharina Kann. 2021. Prost: Physical reasoning about objects through space and time. In Findings.
  2. 2.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Yih, and Yejin Choi. 2019. Abductive commonsense reasoning. ArXiv, abs/1908.05739.
  3. 3.Chandra Bhagavatula, Jena D. Hwang, Doug Downey, Ronan Le Bras, Ximing Lu, Keisuke Sakaguchi, Swabha Swayamdipta, Peter West, and Yejin Choi. 2022. I2d2: Inductive knowledge distillation with neurologic and self-imitation. ArXiv, abs/2212.09246.
  4. 4.Sumithra Bhakthavatsalam, Chloe Anastasiades, and Peter Clark. 2020. Genericskb: A knowledge base of generic statements. ArXiv, abs/2005.00660.
  5. 5.Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019. Piqa: Reasoning about physical commonsense in natural language. ArXiv, abs/1911.11641.
  6. 6.Ali Borji. 2023. A categorical archive of chatgpt failures. ArXiv, abs/2302.03494.
  7. 7.Kaj Bostrom, Zayne Sprague, Swarat Chaudhuri, and Greg Durrett. 2022. Natural language deduction through search over statement compositions. In Conference on Empirical Methods in Natural Language Processing.
  8. 8.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. ArXiv, abs/2005.14165.
  9. 9.Jifan Chen, Eunsol Choi, and Greg Durrett. 2021. Can nli models verify qa systems’ predictions? ArXiv, abs/2104.08731.
  10. 10.Michael Chen, Mike D’Arcy, Alisa Liu, Jared Fernández, and Doug Downey. 2019. Codah: An adversarially authored question-answer dataset for common sense. arXiv preprint arXiv:1904.04365.
  11. 11.Hyung Won Chung, Le Hou, S. Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Wei Yu, Vincent Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed Huai hsin Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. ArXiv, abs/2210.11416.
  12. 12.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. ArXiv, abs/1803.05457.
  13. 13.Ernest Davis. 2023. Benchmarks for automated commonsense reasoning: A survey. ArXiv, abs/2302.04752.
  14. 14.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  15. 15.Andrew S. Gordon, Zornitsa Kozareva, and Melissa Roemmele. 2011. Semeval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In International Workshop on Semantic Evaluation.
  16. 16.Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, and Sourab Mangrulkar. 2022. Accelerate: Training and inference at scale made simple, efficient and adaptable. https://github.com/huggingface/accelerate.
  17. 17.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. In International Conference on Machine Learning.
  18. 18.Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, and Yejin Choi. 2020. Comet-atomic 2020: On symbolic and neural commonsense knowledge graphs. In AAAI Conference on Artificial Intelligence.
  19. 19.Albert Qiaochu Jiang, Sean Welleck, Jin Peng Zhou, Wenda Li, Jiacheng Liu, Mateja Jamnik, Timothée Lacroix, Yuhuai Wu, and Guillaume Lample. 2022. Draft, sketch, and prove: Guiding formal theorem provers with informal proofs. ArXiv, abs/2210.12283.
  20. 20.Jaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras, and Yejin Choi. 2022. Maieutic prompting: Logically consistent reasoning with recursive explanations. In Conference on Empirical Methods in Natural Language Processing.
  21. 21.Saurav Kadavath, Tom Conerly, Amanda Askell, T. J. Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zachary Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Daniel Hernandez, Josh Jacobson, John Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom B. Brown, Jack Clark, Nicholas Joseph, Benjamin Mann, Sam McCandlish, Christopher Olah, and Jared Kaplan. 2022. Language models (mostly) know what they know. ArXiv, abs/2207.05221.
  22. 22.Daniel Khashabi, Yeganeh Kordi, and Hannaneh Hajishirzi. 2022. Unifiedqa-v2: Stronger generalization via broader cross-format training. ArXiv, abs/2202.12359.
  23. 23.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. Unifiedqa: Crossing format boundaries with a single qa system. In Findings.
  24. 24.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised contrastive learning. In Advances in Neural Information Processing Systems, volume 33.
  25. 25.Tushar Khot, Peter Clark, Michal Guerquin, Peter Alexander Jansen, and Ashish Sabharwal. 2019. Qasc: A dataset for question answering via sentence composition. ArXiv, abs/1910.11473.
  26. 26.Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. CoRR, abs/1412.6980.
  27. 27.Hector J. Levesque, Ernest Davis, and L. Morgenstern. 2011. The winograd schema challenge. In International Conference on Principles of Knowledge Representation and Reasoning.
  28. 28.Bill Yuchen Lin, Seyeon Lee, Rahul Khanna, and Xiang Ren. 2020. Birds have four legs?! numersense: Probing numerical commonsense knowledge of pretrained language models. ArXiv, abs/2005.00683.
  29. 29.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In ACL, pages 3214–3252.
  30. 30.Alisa Liu, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2022a. Wanli: Worker and ai collaboration for natural language inference dataset creation. In Conference on Empirical Methods in Natural Language Processing.
  31. 31.Jiacheng Liu, Skyler Hallinan, Ximing Lu, Pengfei He, Sean Welleck, Hannaneh Hajishirzi, and Yejin Choi. 2022b. Rainier: Reinforced knowledge introspector for commonsense question answering. In Conference on Empirical Methods in Natural Language Processing.
  32. 32.Jiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck, Peter West, Ronan Le Bras, Yejin Choi, and Hannaneh Hajishirzi. 2021. Generated knowledge prompting for commonsense reasoning. ArXiv, abs/2110.08387.
  33. 33.Xiao Liu, Da Yin, Yansong Feng, and Dongyan Zhao. 2022c. Things not written in text: Exploring spatial commonsense from visual signals. ArXiv, abs/2203.08075.
  34. 34.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. ArXiv, abs/1907.11692.
  35. 35.Nicholas Lourie, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Unicorn on rainbow: A universal commonsense reasoning model on a new multitask benchmark. In AAAI Conference on Artificial Intelligence.
  36. 36.Gary Marcus and Ernest Davis. 2023. Chatgpt/llm errors (public). https://docs.google.com/spreadsheets/d/1kDSERnROv5FgHbVN8z_bXH9gak2IXRtoqz0nwhrviCw/edit.
  37. 37.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Conference on Empirical Methods in Natural Language Processing.
  38. 38.N. Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James F. Allen. 2016. A corpus and cloze evaluation for deeper understanding of commonsense stories. In North American Chapter of the Association for Computational Linguistics.
  39. 39.Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015. Obtaining well calibrated probabilities using bayesian binning. Proceedings of the ... AAAI Conference on Artificial Intelligence. AAAI Conference on Artificial Intelligence, 2015:2901–2907.
  40. 40.Yasumasa Onoe, Michael J.Q. Zhang, Eunsol Choi, and Greg Durrett. 2021. Creak: A dataset for commonsense reasoning over entity knowledge. ArXiv, abs/2109.01653.
  41. 41.OpenAI. 2022a. Introducing chatgpt.
  42. 42.OpenAI. 2022b. Models - overview - gpt-3.5.
  43. 43.OpenAI. 2023. Gpt-4 technical report. Technical report, OpenAI.
  44. 44.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  45. 45.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. Winogrande: An adversarial winograd schema challenge at scale. Commun. ACM, 64:99–106.
  46. 46.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019. Social iqa: Commonsense reasoning about social interactions. ArXiv, abs/1904.09728.
  47. 47.Shikhar Singh, Nuan Wen, Yu Hou, Pegah Alipoormolabashi, Te-Lin Wu, Xuezhe Ma, and Nanyun Peng. 2021. Com2sense: A commonsense reasoning benchmark with complementary sentences. In Findings.
  48. 48.Zayne Sprague, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett. 2022. Natural language deduction with incomplete information. In Conference on Empirical Methods in Natural Language Processing.
  49. 49.Oyvind Tafjord and Peter Clark. 2021. General-purpose question-answering with macaw. ArXiv, abs/2109.02593.
  50. 50.Oyvind Tafjord, Peter Clark, Matt Gardner, Wen tau Yih, and Ashish Sabharwal. 2018. Quarel: A dataset and models for answering questions about qualitative relationships. ArXiv, abs/1811.08048.
  51. 51.Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. 2022. Entailer: Answering questions with faithful and truthful chains of reasoning. In Conference on Empirical Methods in Natural Language Processing.
  52. 52.Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark. 2019. Quartz: An open-domain dataset of qualitative relationship questions. ArXiv, abs/1909.03553.
  53. 53.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. Commonsenseqa: A question answering challenge targeting commonsense knowledge. ArXiv, abs/1811.00937.
  54. 54.Alon Talmor, Ori Yoran, Ronan Le Bras, Chandrasekhar Bhagavatula, Yoav Goldberg, Yejin Choi, and Jonathan Berant. 2021. Commonsenseqa 2.0: Exposing the limits of ai through gamification. ArXiv, abs/2201.05320.
  55. 55.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. Fever: a large-scale dataset for fact extraction and verification. ArXiv, abs/1803.05355.
  56. 56.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aur’elien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. ArXiv, abs/2302.13971.
  57. 57.Giuseppe Venuto. 2023. chatgpt-failures. https://github.com/giuven95/chatgpt-failures.
  58. 58.David Wadden, Kyle Lo, Lucy Lu Wang, Shanchuan Lin, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or fiction: Verifying scientific claims. ArXiv, abs/2004.14974.
  59. 59.Cunxiang Wang, Shuailong Liang, Yili Jin, Yilong Wang, Xiaodan Zhu, and Yue Zhang. 2020. Semeval-2020 task 4: Commonsense validation and explanation. In International Workshop on Semantic Evaluation.
  60. 60.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Huai hsin Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. ArXiv, abs/2201.11903.
  61. 61.Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017. Crowdsourcing multiple choice science questions. ArXiv, abs/1707.06209.
  62. 62.Peter West, Chandrasekhar Bhagavatula, Jack Hessel, Jena D. Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi. 2021. Symbolic knowledge distillation: from general language models to commonsense models. In North American Chapter of the Association for Computational Linguistics.
  63. 63.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019. Huggingface’s transformers: State-of-the-art natural language processing. ArXiv, abs/1910.03771.
  64. 64.Kaiyu Yang, Jia Deng, and Danqi Chen. 2022. Generating natural language proofs with verifier-guided search. In Conference on Empirical Methods in Natural Language Processing.
  65. 65.Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018. Swag: A large-scale adversarial dataset for grounded commonsense inference. In Conference on Empirical Methods in Natural Language Processing.
  66. 66.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? In Annual Meeting of the Association for Computational Linguistics.
  67. 67.Sheng Zhang, Rachel Rudinger, Kevin Duh, and Benjamin Van Durme. 2017. Ordinal common-sense inference. Transactions of the Association for Computational Linguistics, 5:379–395.

Citation

MLA
Liu, J., et al. “Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1264–87, https://doi.org/10.18653/v1/2023.emnlp-main.81.
APA
Liu, J., Wang, W., Wang, D., Smith, N. A., Choi, Y., & Hajishirzi, H. (2023). Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1264–1287. https://doi.org/10.18653/v1/2023.emnlp-main.81
Chicago
Liu, J., W. Wang, D. Wang, N. A. Smith, Y. Choi, and H. Hajishirzi. 2023. “Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1264–87. https://doi.org/10.18653/v1/2023.emnlp-main.81.
Harvard
Liu, J. et al. (2023) “Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1264–1287. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.81.
Vancouver
1. Liu J, Wang W, Wang D, Smith NA, Choi Y, Hajishirzi H (2023) Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1264–1287

BibTeX

@inproceedings{liu-etal-2023-vera,
    title = "Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements",
    author = "Liu, Jiacheng  and
      Wang, Wenya  and
      Wang, Dianzhuo  and
      Smith, Noah  and
      Choi, Yejin  and
      Hajishirzi, Hannaneh",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.81/",
    doi = "10.18653/v1/2023.emnlp-main.81",
    pages = "1264--1287"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/