keyword
natural language inference
Natural language inference is a core task in natural language processing that involves determining the directional semantic relationship between two text sequences, conventionally designated as a premise and a hypothesis. The system evaluates whether the hypothesis logically follows from the premise, directly contradicts it, or remains neutral because its truth value cannot be confirmed or disproven based solely on the premise. Also known as recognizing textual entailment, this task tests a model's capacity for semantic understanding, commonsense reasoning, and logical deduction. It serves as a standard benchmark for evaluating language representations and provides an essential framework for downstream applications such as automated fact-checking, question answering, and factual consistency verification in text generation.
41 items

VariErr NLI: Separating Annotation Error from Human Label Variation
Leon Weber-Genzel, Siyao Peng, Marie-Catherine de Marneffe, Barbara Plank
Why you should read this
Introduces the VariErr benchmark and a two-round explanation-validation methodology to disentangle genuine human label variation from true annotation errors in natural language inference, exposing critical performance gaps in automatic error detection methods.
Human label variation arises when annotators assign different labels to the same item for valid reasons, while annotation errors occur when labels are assigned for invalid reasons. These two issues are prevalent in NLP benchmarks, yet existing research has studied them in isolation. To the best of our knowledge, there exists no prior work that focuses on teasing apart error from signal, especially in cases where signal is beyond black-and-white. To fill this gap, we introduce a systematic methodology and a new dataset, VariErr (variation versus error), focusing on the NLI task in English. We propose a 2-round annotation procedure with annotators explaining each label and subsequently judging the validity of label-explanation pairs. VariErr contains 7,732 validity judgments on 1,933 explanations for 500 re-annotated MNLI items. We assess the effectiveness of various automatic error detection (AED) methods and GPTs in uncovering errors versus human label variation. We find that state-of-the-art AED methods significantly underperform GPTs and humans. While GPT-4 is the best system, it still falls short of human performance. Our methodology is applicable beyond NLI, offering fertile ground for future research on error versus plausible variation, which in turn can yield better and more trustworthy NLP systems.
Added
2026-10-04

PAED: Zero-Shot Persona Attribute Extraction in Dialogues
Luyao Zhu, Wei Li, Rui Mao, Vlad Pandelea, Erik Cambria
Why you should read this
Introduces a refined persona attribute triplet dataset alongside a generation-based framework that uses a variational autoencoder hard negative sampling strategy to extract structured user personas in zero-shot dialogue settings.
Persona attribute extraction is critical for personalized human-computer interaction. Dialogue is an important medium that communicates and delivers persona information. Although there is a public dataset for triplet-based persona attribute extraction from conversations, its automatically generated labels present many issues, including unspecific relations and inconsistent annotations. We fix such issues by leveraging more reliable text-label matching criteria to generate high-quality data for persona attribute extraction. We also propose a contrastive learning- and generation-based model with a novel hard negative sampling strategy for generalized zero-shot persona attribute extraction. We benchmark our model with state-of-the-art baselines on our dataset and a public dataset, showing outstanding accuracy gains. Our sampling strategy also exceeds others by a large margin in persona attribute extraction.
Added
2026-10-04

MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs
Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, Salman Avestimehr
Why you should read this
Presents Meaning-Aware Response Scoring (MARS), a framework that weights each token by its semantic contribution to the answer rather than applying uniform length normalization, substantially improving uncertainty estimation and error detection across multiple large language models and question-answering benchmarks.
Generative Large Language Models (LLMs) are widely utilized for their excellence in various tasks. However, their tendency to produce inaccurate or misleading outputs poses a potential risk, particularly in high-stakes environments. Therefore, estimating the correctness of generative LLM outputs is an important task for enhanced reliability. Uncertainty Estimation (UE) in generative LLMs is an evolving domain, where SOTA probability-based methods commonly employ length-normalized scoring. In this work, we propose Meaning-Aware Response Scoring (MARS) as an alternative to length-normalized scoring for UE methods. MARS is a novel scoring function that considers the semantic contribution of each token in the generated sequence in the context of the question. We demonstrate that integrating MARS into UE methods results in a universal and significant improvement in UE performance. We conduct experiments using three distinct closed-book question-answering datasets across five popular pre-trained LLMs. Lastly, we validate the efficacy of MARS on a Medical QA dataset. Code can be found here.
Added
2026-10-04

(QA)²: Question Answering with Questionable Assumptions
Najoung Kim, Phu Mon Htut, Samuel R. Bowman, Jackson Petty
Why you should read this
Introduces the (QA)² benchmark of naturally occurring search queries to evaluate how well language models identify false or unverifiable premises and appropriately correct them rather than generating misleading answers.
Naturally occurring information-seeking questions often contain questionable assumptions—assumptions that are false or unverifiable. Questions containing questionable assumptions are challenging because they require a distinct answer strategy that deviates from typical answers for information-seeking questions. For instance, the question When did Marie Curie discover Uranium? cannot be answered as a typical when question without addressing the false assumption Marie Curie discovered Uranium. In this work, we propose (QA)² (Question Answering with Questionable Assumptions), an open-domain evaluation dataset consisting of naturally occurring search engine queries that may or may not contain questionable assumptions. To be successful on (QA)², systems must be able to detect questionable assumptions and also be able to produce adequate responses for both typical information-seeking questions and ones with questionable assumptions. Through human rater acceptability on end-to-end QA with (QA)², we find that current models do struggle with handling questionable assumptions, leaving substantial headroom for progress.
Added
2026-10-03

Evaluating Factuality in Text Simplification
Ashwin Devaraj, William Sheffield, Byron C. Wallace, Junyi Jessy Li
Why you should read this
Presents a taxonomy of factual errors in text simplification, revealing that standard evaluation metrics fail to detect frequent information insertions, deletions, and substitutions in both benchmark datasets and model outputs.
Automated simplification models aim to make input texts more readable. Such methods have the potential to make complex information accessible to a wider audience, e.g., providing access to recent medical literature which might otherwise be impenetrable for a lay reader. However, such models risk introducing errors into automatically simplified texts, for instance by inserting statements unsupported by the corresponding original text, or by omitting key information. Providing more readable but inaccurate versions of texts may in many cases be worse than providing no such access at all. The problem of factual accuracy (and the lack thereof) has received heightened attention in the context of summarization models, but the factuality of automatically simplified texts has not been investigated. We introduce a taxonomy of errors that we use to analyze both references drawn from standard simplification datasets and state-of-the-art model outputs. We find that errors often appear in both that are not captured by existing evaluation metrics, motivating a need for research into ensuring the factual accuracy of automated simplification models.
Added
2026-10-03

A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
Alon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins, Roee Aharoni, Mor Geva
Why you should read this
Presents REVEAL, a benchmark dataset equipped with step-level annotations for relevance, evidence attribution, and logical correctness to systematically evaluate how well automatic verifiers detect errors in language model reasoning chains.
Prompting language models to provide step-by-step answers (e.g., “Chain-of-Thought”) is the prominent approach for complex reasoning tasks, where more accurate reasoning chains typically improve downstream task performance. Recent literature discusses automatic methods to verify reasoning steps to evaluate and improve their correctness. However, no fine-grained step-level datasets are available to enable thorough evaluation of such verification methods, hindering progress in this direction. We introduce REVEAL: Reasoning Verification Evaluation, a new dataset to benchmark automatic verifiers of complex Chain-of-Thought reasoning in open-domain question answering settings. REVEAL includes comprehensive labels for the relevance, attribution to evidence passages, and logical correctness of each reasoning step in a language model’s answer, across a wide variety of datasets and state-of-the-art language models. Available at reveal-dataset.github.io.
Added
2026-10-03

Reasoning Like Program Executors
Xinyu Pi, Qian Liu, Bei Chen, Morteza Ziyadi, Zeqi Lin, Qiang Fu, Yan Gao, Jian-Guang Lou, Weizhu Chen
Why you should read this
Proposes a pre-training paradigm that teaches language models to predict program execution outputs, transferring formal symbolic reasoning capabilities directly into neural models for downstream natural language tasks.
Reasoning over natural language is a long-standing goal for the research community. However, studies have shown that existing language models are inadequate in reasoning. To address the issue, we present PoET, a novel reasoning pre-training paradigm. Through pre-training language models with programs and their execution results, PoET empowers language models to harvest the reasoning knowledge possessed by program executors via a data-driven approach. PoET is conceptually simple and can be instantiated by different kinds of program executors. In this paper, we showcase two simple instances PoET-Math and PoET-Logic, in addition to a complex instance, PoET-SQL. Experimental results on six benchmarks demonstrate that PoET can significantly boost model performance in natural language reasoning, such as numerical reasoning, logical reasoning, and multi-hop reasoning. PoET opens a new gate on reasoning-enhancement pre-training, and we hope our analysis would shed light on the future research of reasoning like program executors.
Added
2026-10-03

Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for Misinformation
Max Glockner, Yufang Hou, Iryna Gurevych
Why you should read this
Reveals that automated fact-checking models fail on real-world misinformation because they rely on leaked counter-evidence from post-hoc reports rather than disproving the underlying reasoning behind novel claims.
The task of misinformation detection has tremendous potential to make a significant contribution to society. However, current automatic fact-checking systems ignore an important aspect of fact-checking: the lack of counter-evidence. In this paper, we analyze the prevalence of counter-evidence in real-world fact-checking datasets. We find that counter-evidence is almost always present in the evidence used by human fact-checkers, but is missing in the evidence retrieved by current automatic fact-checking systems. We argue that this discrepancy is a major reason for the lack of robustness of current systems. To address this issue, we propose a new task setting that requires the system to identify the lack of counter-evidence and to abstain from making a prediction in such cases. We show that this is a challenging task for current systems, and that our proposed method can improve the robustness of fact-checking systems.
Added
2026-10-03

Entailer: Answering Questions with Faithful and Truthful Chains of Reasoning
Oyvind Tafjord, Bhavana Dalvi Mishra, Peter Clark
Why you should read this
Proposes a question-answering framework that couples backward-chaining entailment generation with self-query verification to construct multi-step reasoning trees directly reflecting a language model's internal beliefs.
Our goal is for a system to be able to answer questions and justify its answers with a faithful and truthful chain of reasoning. We present Entailer, a system that produces chains of reasoning by decomposing the question into subquestions and then answering each subquestion using a textual entailment model. The model is trained on a new dataset of 10K examples of multistep reasoning, and uses a novel method to generate the training data automatically from a large corpus. Entailer achieves state-of-the-art results on two multistep reasoning datasets, and its reasoning chains are more faithful and truthful than those of a large language model, as measured by human evaluation and a new automatic metric. We also show that Entailer can be used to improve the performance of a large language model by providing it with a faithful derivation of the answer.
Added
2026-10-03

LAMBADA: Backward Chaining for Automated Reasoning in Natural Language
Mehran Kazemi, Najoung Kim, Deepti Bhatia, Xin Xu, Deepak Ramachandran
Why you should read this
Proposes LAMBADA, a modular backward chaining framework that recursively decomposes natural language reasoning goals using few-shot prompted language models to significantly improve proof accuracy and query efficiency over forward reasoning methods.
Remarkable progress has been made on automated reasoning with natural text, by using Language Models (LMs) and methods such as Chain-of-Thought and Selection-Inference. These techniques search for proofs in the forward direction from axioms to the conclusion, which suffers from a combinatorial explosion of the search space, and thus high failure rates for problems requiring longer chains of reasoning. The classical automated reasoning literature has shown that reasoning in the backward direction (i.e. from the intended conclusion to supporting axioms) is significantly more efficient at proof-finding. Importing this intuition into the LM setting, we develop a Backward Chaining algorithm, called LAMBADA, that decomposes reasoning into four sub-modules. These sub-modules are simply implemented by few-shot prompted LM inference. We show that LAMBADA achieves sizable accuracy boosts over state-of-the-art forward reasoning methods on two challenging logical reasoning datasets, particularly when deep and accurate proof chains are required.
Added
2026-10-02

Mitigating Label Biases for In-context Learning
Yu Fei, Yifan Hou, Zeming Chen, Antoine Bosselut
Why you should read this
Proposes a domain-context calibration method that estimates and removes pre-existing task corpus biases from language models, improving in-context text classification performance by up to 37% Macro-F1.
Various design settings for in-context learning (ICL), such as the choice and order of the in-context examples, can bias the model’s predictions. While many studies discuss these design choices, there have been few systematic investigations into categorizing them and mitigating their impact. In this work, we define a typology for three types of label biases in ICL for text classification: vanilla-label bias, context-label bias, and domain-label bias (which we conceptualize and detect for the first time). Our analysis demonstrates that prior label bias calibration methods fall short of addressing all three types of biases. Specifically, domain-label bias restricts LLMs to random-level performance on many tasks regardless of the choice of in-context examples. To mitigate the effect of these biases, we propose a simple bias calibration method that estimates a language model’s label bias using random in-domain words from the task corpus. After controlling for this estimated bias when making predictions, our novel domain-context calibration significantly improves the ICL performance of GPT-J and GPT-3 on a wide range of tasks. The gain is substantial on tasks with large domain-label bias (up to 37% in Macro-F1). Furthermore, our results generalize to models with different scales, pretraining methods, and manually-designed task instructions, showing the prevalence of label biases in ICL.
Added
2026-10-01

Compositional Exemplars for In-context Learning
Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu, Lingpeng Kong
Why you should read this
Proposes a determinantal point process framework that optimizes in-context exemplar selection by modeling both input relevance and inter-example diversity through contrastive learning, achieving state-of-the-art performance across 12 diverse NLP benchmarks.
Large pretrained language models (LMs) have shown impressive In-Context Learning (ICL) ability, where the model learns to do an unseen task via a prompt consisting of input-output examples as the demonstration, without any parameter updates. The performance of ICL is highly dominated by the quality of the selected in-context examples. However, previous selection methods are mostly based on simple heuristics, leading to sub-optimal performance. In this work, we formulate in-context example selection as a subset selection problem. We propose CEIL (Compositional Exemplars for In-context Learning), which is instantiated by Determinantal Point Processes (DPPs) to model the interaction between the given input and in-context examples, and optimized through a carefully-designed contrastive learning objective to obtain preference from LMs. We validate CEIL on 12 classification and generation datasets from 7 distinct NLP tasks, including sentiment analysis, paraphrase detection, natural language inference, commonsense reasoning, open-domain question answering, code generation, and semantic parsing. Extensive experiments demonstrate not only the state-of-the-art performance but also the transferability and compositionality of CEIL, shedding new light on in-context learning. Our code is released at https://github.com/HKUNLP/icl-ceil.
Added
2026-09-30

Fact-Checking Complex Claims with Program-Guided Reasoning
Liangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, Preslav Nakov
Why you should read this
Proposes a program-guided framework that decomposes complex claims into executable reasoning steps handled by specialized tools, delivering interpretable and data-efficient automated fact-checking that outperforms competitive baselines across multiple evidence settings.
Fact-checking real-world claims often requires collecting multiple pieces of evidence and applying complex multi-step reasoning. In this paper, we present Program-Guided Fact-Checking (PROGRAMFC), a novel fact-checking model that decomposes complex claims into simpler sub-tasks that can be solved using a shared library of specialized functions. We first leverage the in-context learning ability of large language models to generate reasoning programs to guide the verification process. Afterward, we execute the program by delegating each sub-task to the corresponding sub-task handler. This process makes our model both explanatory and data-efficient, providing clear explanations of its reasoning process and requiring minimal training data. We evaluate PROGRAMFC on two challenging fact-checking datasets and show that it outperforms seven fact-checking baselines across different settings of evidence availability, with explicit output programs that benefit human debugging.1
Added
2026-09-28

MetaICL: Learning to Learn In Context
Sewon Min, Mike Lewis, Luke Zettlemoyer, Hannaneh Hajishirzi
Why you should read this
Proposes a meta-training framework that teaches language models how to perform in-context learning across diverse tasks, allowing smaller models to generalize to unseen target tasks and rival fully finetuned baselines without requiring task-specific templates or parameter updates.
We introduce MetaICL (Meta-training for In-Context Learning), a new meta-training framework for few-shot learning where a pretrained language model is tuned to do in-context learning on a large set of training tasks. This meta-training enables the model to more effectively learn a new task in context at test time, by simply conditioning on a few training examples with no parameter updates or task-specific templates. We experiment on a large, diverse collection of tasks consisting of 142 NLP datasets including classification, question answering, natural language inference, paraphrase detection and more, across seven different meta-training/target splits. MetaICL outperforms a range of baselines including in-context learning without meta-training and multi-task learning followed by zero-shot transfer. We find that the gains are particularly significant for target tasks that have domain shifts from the meta-training tasks, and that using a diverse set of the meta-training tasks is key to improvements. We also show that MetaICL approaches (and sometimes beats) the performance of models fully finetuned on the target task, and outperforms much bigger models with nearly 8x parameters. Finally, we show that MetaICL is complementary to human-written instructions, and the best performance can be achieved by combining both approaches.
Added
2026-09-28

TRUE: Re-evaluating Factual Consistency Evaluation
Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansky, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, Yossi Matias
Why you should read this
Establishes a standardized benchmark and example-level evaluation protocol across eleven datasets to measure how reliably factual consistency metrics detect grounded text generation errors.
Grounded text generation systems often generate text that contains factual inconsistencies, hindering their real-world applicability. Automatic factual consistency evaluation may help alleviate this limitation by accelerating evaluation cycles, filtering inconsistent outputs and augmenting training data. While attracting increasing attention, such evaluation metrics are usually developed and evaluated in silo for a single task or dataset, slowing their adoption. Moreover, previous meta-evaluation protocols focused on system-level correlations with human annotations, which leave the example-level accuracy of such metrics unclear. In this work, we introduce TRUE: a comprehensive survey and assessment of factual consistency metrics on a standardized collection of existing texts from diverse tasks, manually annotated for factual consistency. Our standardization enables an example-level meta-evaluation protocol that is more actionable and interpretable than previously reported correlations, yielding clearer quality measures. Across diverse state-of-the-art metrics and 11 datasets we find that large-scale NLI and question generation-and-answering-based approaches achieve strong and complementary results. We recommend those methods as a starting point for model and metric developers, and hope TRUE will foster progress towards even better evaluation methods.
Added
2026-09-26

Consistency Analysis of ChatGPT
Myeongjun Jang, Thomas Lukasiewicz
Why you should read this
Demonstrates that ChatGPT and GPT-4 frequently violate fundamental logical properties such as semantic, negation, symmetric, and transitive consistency, proving that standard mitigation techniques like prompt engineering, few-shot prompting, and model scaling cannot fully resolve large language model inconsistency.
ChatGPT has gained a huge popularity since its introduction. Its positive aspects have been reported through many media platforms, and some analyses even showed that ChatGPT achieved a decent grade in professional exams, adding extra support to the claim that AI can now assist and even replace humans in industrial fields. Others, however, doubt its reliability and trustworthiness. This paper investigates the trustworthiness of ChatGPT and GPT-4 regarding logically consistent behaviour, focusing specifically on semantic consistency and the properties of negation, symmetric, and transitive consistency. Our findings suggest that while both models appear to show an enhanced language understanding and reasoning ability, they still frequently fall short of generating logically consistent predictions. We also ascertain via experiments that prompt designing, few-shot learning and employing larger large language models (LLMs) are unlikely to be the ultimate solution to resolve the inconsistency issue of LLMs.
Added
2026-09-26

NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks
Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Singh Sachdeva, Peter Clark, Chitta Baral, Ashwin Kalyan
Why you should read this
Presents NUMGLUE, an eight-task benchmark spanning roughly 100,000 problems that exposes large language models' severe arithmetic brittleness compared to human reasoning while showing that joint multi-task training significantly boosts numerical performance across diverse question formats.
Given the ubiquitous nature of numbers in text, reasoning with numbers to perform simple calculations is an important skill of AI systems. While many datasets and models have been developed to this end, state-of-the-art AI systems are brittle; failing to perform the underlying mathematical reasoning when they appear in a slightly different scenario. Drawing inspiration from GLUE (Wang et al., 2018) that was proposed in the context of natural language understanding, we propose NUMGLUE, a multi-task benchmark that evaluates the performance of AI systems on eight different tasks, that at their core require simple arithmetic understanding. We show that this benchmark is far from being solved with neural models including state-of-the-art large-scale language models performing significantly worse than humans (lower by 46.4%). Further, NUMGLUE promotes sharing knowledge across tasks, especially those with limited training data as evidenced by the superior performance (average gain of 3.4% on each task) when a model is jointly trained on all the tasks as opposed to task-specific modeling. Finally, we hope that NUMGLUE will encourage systems that perform robust and general arithmetic reasoning within language, a first step towards being able to perform more complex mathematical reasoning.
Added
2026-09-26

Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements
Jiacheng Liu, Wenya Wang, Dianzhuo Wang, Noah A. Smith, Yejin Choi, Hannaneh Hajishirzi
Why you should read this
Presents VERA, a standalone commonsense verification model trained on millions of statements that outperforms systems like GPT-4 in estimating the plausibility of declarative claims and detecting errors in language model outputs.
Today’s language models can be remarkably intelligent yet still produce text that contains trivial commonsense errors. Therefore, we seek a retrospective verification approach that can reflect on the commonsense plausibility of the machine text, and introduce VERA, a general-purpose model that learns to estimate the commonsense plausibility of declarative statements. To support diverse commonsense domains, VERA is trained on ~7M commonsense statements that are automatically converted from 19 QA datasets and two commonsense knowledge bases, and using a combination of three training objectives. When applied to solving commonsense problems in the verification format, VERA substantially outperforms existing models that can be repurposed for commonsense verification, even including GPT-3.5/ChatGPT/GPT-4, and it further exhibits generalization capabilities to unseen tasks and provides well-calibrated outputs. We find that VERA excels at filtering machine-generated commonsense knowledge and is useful in detecting erroneous commonsense statements generated by models like ChatGPT in real-world settings.
Added
2026-09-26

Debiased Contrastive Learning of Unsupervised Sentence Representations
Kun Zhou, Beichen Zhang, Wayne Xin Zhao, Ji-Rong Wen
Why you should read this
Proposes a debiased contrastive learning framework that improves unsupervised sentence embeddings by downweighting false negatives and generating optimized noise-based negative samples to overcome representation anisotropy.
Recently, contrastive learning has been shown to be effective in improving pre-trained language models (PLM) to derive high-quality sentence representations. It aims to pull close positive examples to enhance the alignment while push apart irrelevant negatives for the uniformity of the whole representation space. However, previous works mostly adopt in-batch negatives or sample from training data at random. Such a way may cause the sampling bias that improper negatives (e.g., false negatives and anisotropy representations) are used to learn sentence representations, which will hurt the uniformity of the representation space. To address it, we present a new framework DCLR (Debiased Contrastive Learning of unsupervised sentence Representations) to alleviate the influence of these improper negatives. In DCLR, we design an instance weighting method to punish false negatives and generate noise-based negatives to guarantee the uniformity of the representation space. Experiments on seven semantic textual similarity tasks show that our approach is more effective than competitive baselines. Our code and data are publicly available at the link: https://github.com/RUCAIBox/DCLR.
Added
2026-09-26

A Decomposable Attention Model for Natural Language Inference
Ankur P. Parikh, Oscar Täckström, Dipanjan Das, Jakob Uszkoreit
Why you should read this
Introduces a lightweight, highly parallelizable decomposable attention model for natural language inference that achieves state-of-the-art accuracy on the SNLI benchmark with an order of magnitude fewer parameters and without relying on word-order information.
We propose a simple neural architecture for natural language inference. Our approach uses attention to decompose the problem into subproblems that can be solved separately, thus making it trivially parallelizable. On the Stanford Natural Language Inference (SNLI) dataset, we obtain state-of-the-art results with almost an order of magnitude fewer parameters than previous work and without relying on any word-order information. Adding intra-sentence attention that takes a minimum amount of order into account yields further improvements.
Added
2026-09-26

