Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements
Jiacheng LiuWenya WangDianzhuo WangNoah A. SmithYejin ChoiHannaneh Hajishirzi
Presents VERA, a standalone commonsense verification model trained on millions of statements that outperforms systems like GPT-4 in estimating the plausibility of declarative claims and detecting errors in language model outputs.
Modern generative language models demonstrate remarkable capabilities across diverse tasks, yet they frequently generate text containing basic commonsense mistakes and lack built-in mechanisms to quantify uncertainty. Because these failures degrade user trust and pose operational risks in deployment, there is an urgent need for automated systems that can evaluate the commonsense plausibility of machine-generated statements.
The main objective of the article is to introduce and evaluate VERA, a general-purpose plausibility estimation model designed to estimate the commonsense correctness of declarative natural language statements without relying on external text retrieval.
To construct VERA, the researchers trained a 5-billion-parameter neural network using a two-stage process on approximately 7 million statements compiled from 19 question-answering datasets and two commonsense knowledge bases. The training pipeline integrated three distinct learning objectives: standard binary classification, multi-class ranking across statement groups, and supervised contrastive learning to separate similar true and false statements. To enhance robustness, the authors automatically augmented the training data with model-generated incorrect statements and implemented a post-training temperature calibration step to align confidence scores with actual correctness probabilities.
The evaluations yielded several key findings regarding model accuracy and utility. First, when applied to commonsense problem-solving benchmarks in a verification format, VERA achieved an average accuracy of 85.5% on seen datasets and 81.7% to 83.4% on unseen datasets, consistently outperforming leading general-purpose models such as GPT-3.5, ChatGPT, GPT-4, and Flan-T5. Second, when deployed to filter noisy knowledge generated by other language models, VERA enhanced the downstream accuracy boost of knowledge-augmented pipelines by 46% for GPT-3 and by 233% for Rainier. Third, in an evaluation on real-world errors generated by ChatGPT, VERA accurately identified mistakes with 91% precision and 74% recall, yielding an overall F1 score of 82%. Finally, performance scaling analyses revealed steady accuracy gains as model size increased, showing no signs of performance saturation at the 5-billion parameter scale.
These findings demonstrate that dedicated verification models can serve as effective, calibrated safety layers to supervise larger generative systems, reducing hallucination risks and improving automated reasoning quality. Unlike generative question-answering systems that require seeing multiple answer choices simultaneously, standalone verification models provide reliable confidence scores for individual declarative statements across diverse scientific, physical, and social domains.
Organizations developing or deploying generative language applications should consider integrating modular verification systems like VERA into their post-generation pipelines to filter knowledge and flag commonsense errors. When implementing such verifiers, system designers must account for trade-offs, such as the increased computational cost of evaluating multiple candidates individually. Further research and piloting are recommended to extend verification techniques to multi-sentence contexts, complex compositional claims, and domain-specific knowledge.
Readers should interpret these results with appropriate caution due to specific operational limitations. VERA is specifically designed for single-sentence commonsense assertions and is not trained to evaluate encyclopedic facts, long-form compositional text, or moral and ethical judgments. Additionally, while the model demonstrates strong general plausibility scoring, it remains susceptible to syntactic variations such as negations and paraphrases, and it is intended as a research prototype rather than an autonomous decision-making system.
- Paper: CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge, Alon Talmor et al. (2019). This paper establishes the CommonsenseQA benchmark, one of the foundational commonsense question-answering datasets converted and utilized by Vera for plausibility verification training.
- Paper: PIQA: Reasoning about Physical Commonsense in Natural Language, Yonatan Bisk et al. (2019). This work introduces the PIQA physical commonsense reasoning benchmark, providing key source datasets and formulations adapted by Vera into declarative plausibility statements.
- Paper: FEVER: a Large-scale Dataset for Fact Extraction and VERification, James Thorne et al. (2018). This foundational paper formalizes large-scale automated claim verification and evidence entailment, establishing the core verification paradigms that Vera generalizes to open commonsense statements.
- Paper: Training Verifiers to Solve Math Word Problems, Karl Cobbe et al. (2021). This study demonstrates the power of training dedicated verifier models to score and filter candidate language model outputs, inspiring Vera's retrospective verification architecture.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). This benchmark details how generative language models frequently imitate common human misconceptions and errors, framing the core challenge of detecting machine-generated implausibilities that Vera is designed to solve.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). This work analyzes factual consistency evaluation across multiple generation formats, establishing foundational evaluation criteria for standalone verification models.
- Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). This paper examines probing language models for factual and commonsense relations, motivating Vera’s transition from cloze-based extraction to direct statement plausibility estimation.
- Paper: Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge, Jiangjie Chen et al. (2023). This paper extends the study of commonsense errors by diagnosing language models' specific vulnerabilities in expressing negative commonsense knowledge, directly continuing Vera's evaluation of LLM commonsense flaws.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). This research builds upon automated claim verification methods by applying preference optimization and automated factuality scoring to systematically align language models against hallucinations.
- Paper: Factuality of Large Language Models: A Survey, Yuxia Wang et al. (2024). This comprehensive survey synthesizes downstream advances in large language model factuality verification, error categorization, and post-generation filtering techniques such as those pioneered by Vera.
- Paper: Generative Verifiers: Reward Modeling as Next-Token Prediction, Lunjun Zhang et al. (2025). This work advances beyond classification-based discriminative verifiers by developing generative next-token prediction reward models for complex verification and filtering.
- Paper: LM vs LM: Detecting Factual Errors via Cross Examination, Roi Cohen et al. (2023). This study presents an alternative multi-turn conversational cross-examination method to detect factual errors without external knowledge bases, complementing Vera's direct scoring mechanism.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). This paper explores the broader reliability of using automated model evaluators to judge output validity under adversarial conditions, extending the study of model-based verification.
