Built independently by an author, for readers. Read the story and support ChapterPal

keyword

natural language inference

Natural language inference is a core task in natural language processing that involves determining the directional semantic relationship between two text sequences, conventionally designated as a premise and a hypothesis. The system evaluates whether the hypothesis logically follows from the premise, directly contradicts it, or remains neutral because its truth value cannot be confirmed or disproven based solely on the premise. Also known as recognizing textual entailment, this task tests a model's capacity for semantic understanding, commonsense reasoning, and logical deduction. It serves as a standard benchmark for evaluating language representations and provides an essential framework for downstream applications such as automated fact-checking, question answering, and factual consistency verification in text generation.

41 items

VariErr NLI: Separating Annotation Error from Human Label Variation

VariErr NLI: Separating Annotation Error from Human Label Variation

Leon Weber-Genzel, Siyao Peng, Marie-Catherine de Marneffe, Barbara Plank

OrganizationsCENTALFNRSLudwig Maximilian University of MunichMaiNLP LabMunich Center for Machine LearningUniversité catholique de Louvain

Why you should read this

Introduces the VariErr benchmark and a two-round explanation-validation methodology to disentangle genuine human label variation from true annotation errors in natural language inference, exposing critical performance gaps in automatic error detection methods.

Human label variation arises when annotators assign different labels to the same item for valid reasons, while annotation errors occur when labels are assigned for invalid reasons. These two issues are prevalent in NLP benchmarks, yet existing research has studied them in isolation. To the best of our knowledge, there exists no prior work that focuses on teasing apart error from signal, especially in cases where signal is beyond black-and-white. To fill this gap, we introduce a systematic methodology and a new dataset, VariErr (variation versus error), focusing on the NLI task in English. We propose a 2-round annotation procedure with annotators explaining each label and subsequently judging the validity of label-explanation pairs. VariErr contains 7,732 validity judgments on 1,933 explanations for 500 re-annotated MNLI items. We assess the effectiveness of various automatic error detection (AED) methods and GPTs in uncovering errors versus human label variation. We find that state-of-the-art AED methods significantly underperform GPTs and humans. While GPT-4 is the best system, it still falls short of human performance. Our methodology is applicable beyond NLI, offering fertile ground for future research on error versus plausible variation, which in turn can yield better and more trustworthy NLP systems.

Added

2026-10-04

MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs

MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs

Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, Salman Avestimehr

OrganizationsAmazonUniversity of Southern California

Why you should read this

Presents Meaning-Aware Response Scoring (MARS), a framework that weights each token by its semantic contribution to the answer rather than applying uniform length normalization, substantially improving uncertainty estimation and error detection across multiple large language models and question-answering benchmarks.

Generative Large Language Models (LLMs) are widely utilized for their excellence in various tasks. However, their tendency to produce inaccurate or misleading outputs poses a potential risk, particularly in high-stakes environments. Therefore, estimating the correctness of generative LLM outputs is an important task for enhanced reliability. Uncertainty Estimation (UE) in generative LLMs is an evolving domain, where SOTA probability-based methods commonly employ length-normalized scoring. In this work, we propose Meaning-Aware Response Scoring (MARS) as an alternative to length-normalized scoring for UE methods. MARS is a novel scoring function that considers the semantic contribution of each token in the generated sequence in the context of the question. We demonstrate that integrating MARS into UE methods results in a universal and significant improvement in UE performance. We conduct experiments using three distinct closed-book question-answering datasets across five popular pre-trained LLMs. Lastly, we validate the efficacy of MARS on a Medical QA dataset. Code can be found here.

Added

2026-10-04

(QA)²: Question Answering with Questionable Assumptions

(QA)²: Question Answering with Questionable Assumptions

Najoung Kim, Phu Mon Htut, Samuel R. Bowman, Jackson Petty

OrganizationsAmazon Web ServicesBoston UniversityGoogleNew York University

Why you should read this

Introduces the (QA)² benchmark of naturally occurring search queries to evaluate how well language models identify false or unverifiable premises and appropriately correct them rather than generating misleading answers.

Naturally occurring information-seeking questions often contain questionable assumptions—assumptions that are false or unverifiable. Questions containing questionable assumptions are challenging because they require a distinct answer strategy that deviates from typical answers for information-seeking questions. For instance, the question When did Marie Curie discover Uranium? cannot be answered as a typical when question without addressing the false assumption Marie Curie discovered Uranium. In this work, we propose (QA)² (Question Answering with Questionable Assumptions), an open-domain evaluation dataset consisting of naturally occurring search engine queries that may or may not contain questionable assumptions. To be successful on (QA)², systems must be able to detect questionable assumptions and also be able to produce adequate responses for both typical information-seeking questions and ones with questionable assumptions. Through human rater acceptability on end-to-end QA with (QA)², we find that current models do struggle with handling questionable assumptions, leaving substantial headroom for progress.

Added

2026-10-03

Evaluating Factuality in Text Simplification

Evaluating Factuality in Text Simplification

Ashwin Devaraj, William Sheffield, Byron C. Wallace, Junyi Jessy Li

OrganizationsComputer ScienceLinguisticsMathematicsNortheastern UniversityUniversity of Texas at Austin

Why you should read this

Presents a taxonomy of factual errors in text simplification, revealing that standard evaluation metrics fail to detect frequent information insertions, deletions, and substitutions in both benchmark datasets and model outputs.

Automated simplification models aim to make input texts more readable. Such methods have the potential to make complex information accessible to a wider audience, e.g., providing access to recent medical literature which might otherwise be impenetrable for a lay reader. However, such models risk introducing errors into automatically simplified texts, for instance by inserting statements unsupported by the corresponding original text, or by omitting key information. Providing more readable but inaccurate versions of texts may in many cases be worse than providing no such access at all. The problem of factual accuracy (and the lack thereof) has received heightened attention in the context of summarization models, but the factuality of automatically simplified texts has not been investigated. We introduce a taxonomy of errors that we use to analyze both references drawn from standard simplification datasets and state-of-the-art model outputs. We find that errors often appear in both that are not captured by existing evaluation metrics, motivating a need for research into ensuring the factual accuracy of automated simplification models.

Added

2026-10-03

A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains

A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains

Alon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins, Roee Aharoni, Mor Geva

OrganizationsBar-Ilan UniversityGoogleTel Aviv University

Why you should read this

Presents REVEAL, a benchmark dataset equipped with step-level annotations for relevance, evidence attribution, and logical correctness to systematically evaluate how well automatic verifiers detect errors in language model reasoning chains.

Prompting language models to provide step-by-step answers (e.g., “Chain-of-Thought”) is the prominent approach for complex reasoning tasks, where more accurate reasoning chains typically improve downstream task performance. Recent literature discusses automatic methods to verify reasoning steps to evaluate and improve their correctness. However, no fine-grained step-level datasets are available to enable thorough evaluation of such verification methods, hindering progress in this direction. We introduce REVEAL: Reasoning Verification Evaluation, a new dataset to benchmark automatic verifiers of complex Chain-of-Thought reasoning in open-domain question answering settings. REVEAL includes comprehensive labels for the relevance, attribution to evidence passages, and logical correctness of each reasoning step in a language model’s answer, across a wide variety of datasets and state-of-the-art language models. Available at reveal-dataset.github.io.

Added

2026-10-03

Reasoning Like Program Executors

Reasoning Like Program Executors

Xinyu Pi, Qian Liu, Bei Chen, Morteza Ziyadi, Zeqi Lin, Qiang Fu, Yan Gao, Jian-Guang Lou, Weizhu Chen

OrganizationsMicrosoftSea AI LabUniversity of Illinois Urbana-Champaign

Why you should read this

Proposes a pre-training paradigm that teaches language models to predict program execution outputs, transferring formal symbolic reasoning capabilities directly into neural models for downstream natural language tasks.

Reasoning over natural language is a long-standing goal for the research community. However, studies have shown that existing language models are inadequate in reasoning. To address the issue, we present PoET, a novel reasoning pre-training paradigm. Through pre-training language models with programs and their execution results, PoET empowers language models to harvest the reasoning knowledge possessed by program executors via a data-driven approach. PoET is conceptually simple and can be instantiated by different kinds of program executors. In this paper, we showcase two simple instances PoET-Math and PoET-Logic, in addition to a complex instance, PoET-SQL. Experimental results on six benchmarks demonstrate that PoET can significantly boost model performance in natural language reasoning, such as numerical reasoning, logical reasoning, and multi-hop reasoning. PoET opens a new gate on reasoning-enhancement pre-training, and we hope our analysis would shed light on the future research of reasoning like program executors.

Added

2026-10-03

LAMBADA: Backward Chaining for Automated Reasoning in Natural Language

LAMBADA: Backward Chaining for Automated Reasoning in Natural Language

Mehran Kazemi, Najoung Kim, Deepti Bhatia, Xin Xu, Deepak Ramachandran

OrganizationsGoogle

Why you should read this

Proposes LAMBADA, a modular backward chaining framework that recursively decomposes natural language reasoning goals using few-shot prompted language models to significantly improve proof accuracy and query efficiency over forward reasoning methods.

Remarkable progress has been made on automated reasoning with natural text, by using Language Models (LMs) and methods such as Chain-of-Thought and Selection-Inference. These techniques search for proofs in the forward direction from axioms to the conclusion, which suffers from a combinatorial explosion of the search space, and thus high failure rates for problems requiring longer chains of reasoning. The classical automated reasoning literature has shown that reasoning in the backward direction (i.e. from the intended conclusion to supporting axioms) is significantly more efficient at proof-finding. Importing this intuition into the LM setting, we develop a Backward Chaining algorithm, called LAMBADA, that decomposes reasoning into four sub-modules. These sub-modules are simply implemented by few-shot prompted LM inference. We show that LAMBADA achieves sizable accuracy boosts over state-of-the-art forward reasoning methods on two challenging logical reasoning datasets, particularly when deep and accurate proof chains are required.

Added

2026-10-02

Mitigating Label Biases for In-context Learning

Mitigating Label Biases for In-context Learning

Yu Fei, Yifan Hou, Zeming Chen, Antoine Bosselut

OrganizationsÉcole Polytechnique Fédérale de LausanneETH ZurichICNLP LabUniversity of California, Irvine

Why you should read this

Proposes a domain-context calibration method that estimates and removes pre-existing task corpus biases from language models, improving in-context text classification performance by up to 37% Macro-F1.

Various design settings for in-context learning (ICL), such as the choice and order of the in-context examples, can bias the model’s predictions. While many studies discuss these design choices, there have been few systematic investigations into categorizing them and mitigating their impact. In this work, we define a typology for three types of label biases in ICL for text classification: vanilla-label bias, context-label bias, and domain-label bias (which we conceptualize and detect for the first time). Our analysis demonstrates that prior label bias calibration methods fall short of addressing all three types of biases. Specifically, domain-label bias restricts LLMs to random-level performance on many tasks regardless of the choice of in-context examples. To mitigate the effect of these biases, we propose a simple bias calibration method that estimates a language model’s label bias using random in-domain words from the task corpus. After controlling for this estimated bias when making predictions, our novel domain-context calibration significantly improves the ICL performance of GPT-J and GPT-3 on a wide range of tasks. The gain is substantial on tasks with large domain-label bias (up to 37% in Macro-F1). Furthermore, our results generalize to models with different scales, pretraining methods, and manually-designed task instructions, showing the prevalence of label biases in ICL.

Added

2026-10-01

Compositional Exemplars for In-context Learning

Compositional Exemplars for In-context Learning

Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu, Lingpeng Kong

OrganizationsDepartment of Computer ScienceSea AI LabShanghai Artificial Intelligence LaboratoryUniversity of Hong Kong

Why you should read this

Proposes a determinantal point process framework that optimizes in-context exemplar selection by modeling both input relevance and inter-example diversity through contrastive learning, achieving state-of-the-art performance across 12 diverse NLP benchmarks.

Large pretrained language models (LMs) have shown impressive In-Context Learning (ICL) ability, where the model learns to do an unseen task via a prompt consisting of input-output examples as the demonstration, without any parameter updates. The performance of ICL is highly dominated by the quality of the selected in-context examples. However, previous selection methods are mostly based on simple heuristics, leading to sub-optimal performance. In this work, we formulate in-context example selection as a subset selection problem. We propose CEIL (Compositional Exemplars for In-context Learning), which is instantiated by Determinantal Point Processes (DPPs) to model the interaction between the given input and in-context examples, and optimized through a carefully-designed contrastive learning objective to obtain preference from LMs. We validate CEIL on 12 classification and generation datasets from 7 distinct NLP tasks, including sentiment analysis, paraphrase detection, natural language inference, commonsense reasoning, open-domain question answering, code generation, and semantic parsing. Extensive experiments demonstrate not only the state-of-the-art performance but also the transferability and compositionality of CEIL, shedding new light on in-context learning. Our code is released at https://github.com/HKUNLP/icl-ceil.

Added

2026-09-30

Fact-Checking Complex Claims with Program-Guided Reasoning

Fact-Checking Complex Claims with Program-Guided Reasoning

Liangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, Preslav Nakov

OrganizationsMohamed bin Zayed University of Artificial IntelligenceNanyang Technological UniversityNational University of SingaporeUniversity of California, Santa Barbara

Why you should read this

Proposes a program-guided framework that decomposes complex claims into executable reasoning steps handled by specialized tools, delivering interpretable and data-efficient automated fact-checking that outperforms competitive baselines across multiple evidence settings.

Fact-checking real-world claims often requires collecting multiple pieces of evidence and applying complex multi-step reasoning. In this paper, we present Program-Guided Fact-Checking (PROGRAMFC), a novel fact-checking model that decomposes complex claims into simpler sub-tasks that can be solved using a shared library of specialized functions. We first leverage the in-context learning ability of large language models to generate reasoning programs to guide the verification process. Afterward, we execute the program by delegating each sub-task to the corresponding sub-task handler. This process makes our model both explanatory and data-efficient, providing clear explanations of its reasoning process and requiring minimal training data. We evaluate PROGRAMFC on two challenging fact-checking datasets and show that it outperforms seven fact-checking baselines across different settings of evidence availability, with explicit output programs that benefit human debugging.1

Added

2026-09-28

MetaICL: Learning to Learn In Context

MetaICL: Learning to Learn In Context

Sewon Min, Mike Lewis, Luke Zettlemoyer, Hannaneh Hajishirzi

OrganizationsAllen Institute for AIMetaUniversity of Washington

Why you should read this

Proposes a meta-training framework that teaches language models how to perform in-context learning across diverse tasks, allowing smaller models to generalize to unseen target tasks and rival fully finetuned baselines without requiring task-specific templates or parameter updates.

We introduce MetaICL (Meta-training for In-Context Learning), a new meta-training framework for few-shot learning where a pretrained language model is tuned to do in-context learning on a large set of training tasks. This meta-training enables the model to more effectively learn a new task in context at test time, by simply conditioning on a few training examples with no parameter updates or task-specific templates. We experiment on a large, diverse collection of tasks consisting of 142 NLP datasets including classification, question answering, natural language inference, paraphrase detection and more, across seven different meta-training/target splits. MetaICL outperforms a range of baselines including in-context learning without meta-training and multi-task learning followed by zero-shot transfer. We find that the gains are particularly significant for target tasks that have domain shifts from the meta-training tasks, and that using a diverse set of the meta-training tasks is key to improvements. We also show that MetaICL approaches (and sometimes beats) the performance of models fully finetuned on the target task, and outperforms much bigger models with nearly 8x parameters. Finally, we show that MetaICL is complementary to human-written instructions, and the best performance can be achieved by combining both approaches.

Added

2026-09-28

TRUE: Re-evaluating Factual Consistency Evaluation

TRUE: Re-evaluating Factual Consistency Evaluation

Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansky, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, Yossi Matias

OrganizationsGoogleMetaTel Aviv University

Why you should read this

Establishes a standardized benchmark and example-level evaluation protocol across eleven datasets to measure how reliably factual consistency metrics detect grounded text generation errors.

Grounded text generation systems often generate text that contains factual inconsistencies, hindering their real-world applicability. Automatic factual consistency evaluation may help alleviate this limitation by accelerating evaluation cycles, filtering inconsistent outputs and augmenting training data. While attracting increasing attention, such evaluation metrics are usually developed and evaluated in silo for a single task or dataset, slowing their adoption. Moreover, previous meta-evaluation protocols focused on system-level correlations with human annotations, which leave the example-level accuracy of such metrics unclear. In this work, we introduce TRUE: a comprehensive survey and assessment of factual consistency metrics on a standardized collection of existing texts from diverse tasks, manually annotated for factual consistency. Our standardization enables an example-level meta-evaluation protocol that is more actionable and interpretable than previously reported correlations, yielding clearer quality measures. Across diverse state-of-the-art metrics and 11 datasets we find that large-scale NLI and question generation-and-answering-based approaches achieve strong and complementary results. We recommend those methods as a starting point for model and metric developers, and hope TRUE will foster progress towards even better evaluation methods.

Added

2026-09-26

Consistency Analysis of ChatGPT

Consistency Analysis of ChatGPT

Myeongjun Jang, Thomas Lukasiewicz

OrganizationsTU WienUniversity of Oxford

Why you should read this

Demonstrates that ChatGPT and GPT-4 frequently violate fundamental logical properties such as semantic, negation, symmetric, and transitive consistency, proving that standard mitigation techniques like prompt engineering, few-shot prompting, and model scaling cannot fully resolve large language model inconsistency.

ChatGPT has gained a huge popularity since its introduction. Its positive aspects have been reported through many media platforms, and some analyses even showed that ChatGPT achieved a decent grade in professional exams, adding extra support to the claim that AI can now assist and even replace humans in industrial fields. Others, however, doubt its reliability and trustworthiness. This paper investigates the trustworthiness of ChatGPT and GPT-4 regarding logically consistent behaviour, focusing specifically on semantic consistency and the properties of negation, symmetric, and transitive consistency. Our findings suggest that while both models appear to show an enhanced language understanding and reasoning ability, they still frequently fall short of generating logically consistent predictions. We also ascertain via experiments that prompt designing, few-shot learning and employing larger large language models (LLMs) are unlikely to be the ultimate solution to resolve the inconsistency issue of LLMs.

Added

2026-09-26

NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks

NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks

Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Singh Sachdeva, Peter Clark, Chitta Baral, Ashwin Kalyan

OrganizationsAllen Institute for AIArizona State UniversityMicrosoft

Why you should read this

Presents NUMGLUE, an eight-task benchmark spanning roughly 100,000 problems that exposes large language models' severe arithmetic brittleness compared to human reasoning while showing that joint multi-task training significantly boosts numerical performance across diverse question formats.

Given the ubiquitous nature of numbers in text, reasoning with numbers to perform simple calculations is an important skill of AI systems. While many datasets and models have been developed to this end, state-of-the-art AI systems are brittle; failing to perform the underlying mathematical reasoning when they appear in a slightly different scenario. Drawing inspiration from GLUE (Wang et al., 2018) that was proposed in the context of natural language understanding, we propose NUMGLUE, a multi-task benchmark that evaluates the performance of AI systems on eight different tasks, that at their core require simple arithmetic understanding. We show that this benchmark is far from being solved with neural models including state-of-the-art large-scale language models performing significantly worse than humans (lower by 46.4%). Further, NUMGLUE promotes sharing knowledge across tasks, especially those with limited training data as evidenced by the superior performance (average gain of 3.4% on each task) when a model is jointly trained on all the tasks as opposed to task-specific modeling. Finally, we hope that NUMGLUE will encourage systems that perform robust and general arithmetic reasoning within language, a first step towards being able to perform more complex mathematical reasoning.

Added

2026-09-26

Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements

Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements

Jiacheng Liu, Wenya Wang, Dianzhuo Wang, Noah A. Smith, Yejin Choi, Hannaneh Hajishirzi

OrganizationsAllen Institute for AIHarvard UniversityNanyang Technological UniversityUniversity of Washington

Why you should read this

Presents VERA, a standalone commonsense verification model trained on millions of statements that outperforms systems like GPT-4 in estimating the plausibility of declarative claims and detecting errors in language model outputs.

Today’s language models can be remarkably intelligent yet still produce text that contains trivial commonsense errors. Therefore, we seek a retrospective verification approach that can reflect on the commonsense plausibility of the machine text, and introduce VERA, a general-purpose model that learns to estimate the commonsense plausibility of declarative statements. To support diverse commonsense domains, VERA is trained on ~7M commonsense statements that are automatically converted from 19 QA datasets and two commonsense knowledge bases, and using a combination of three training objectives. When applied to solving commonsense problems in the verification format, VERA substantially outperforms existing models that can be repurposed for commonsense verification, even including GPT-3.5/ChatGPT/GPT-4, and it further exhibits generalization capabilities to unseen tasks and provides well-calibrated outputs. We find that VERA excels at filtering machine-generated commonsense knowledge and is useful in detecting erroneous commonsense statements generated by models like ChatGPT in real-world settings.

Added

2026-09-26

Debiased Contrastive Learning of Unsupervised Sentence Representations

Debiased Contrastive Learning of Unsupervised Sentence Representations

Kun Zhou, Beichen Zhang, Wayne Xin Zhao, Ji-Rong Wen

OrganizationsRenmin University of China

Why you should read this

Proposes a debiased contrastive learning framework that improves unsupervised sentence embeddings by downweighting false negatives and generating optimized noise-based negative samples to overcome representation anisotropy.

Recently, contrastive learning has been shown to be effective in improving pre-trained language models (PLM) to derive high-quality sentence representations. It aims to pull close positive examples to enhance the alignment while push apart irrelevant negatives for the uniformity of the whole representation space. However, previous works mostly adopt in-batch negatives or sample from training data at random. Such a way may cause the sampling bias that improper negatives (e.g., false negatives and anisotropy representations) are used to learn sentence representations, which will hurt the uniformity of the representation space. To address it, we present a new framework DCLR (Debiased Contrastive Learning of unsupervised sentence Representations) to alleviate the influence of these improper negatives. In DCLR, we design an instance weighting method to punish false negatives and generate noise-based negatives to guarantee the uniformity of the representation space. Experiments on seven semantic textual similarity tasks show that our approach is more effective than competitive baselines. Our code and data are publicly available at the link: https://github.com/RUCAIBox/DCLR.

Added

2026-09-26