Built independently by an author, for readers. Read the story and support ChapterPal

keyword

machine translation

Machine translation is the automated process of converting text or speech from one natural language into another using computational algorithms and software. As a foundational subfield of artificial intelligence and computational linguistics, it has evolved from early rule-based and statistical approaches to modern neural network architectures and multilingual language models that learn complex linguistic and semantic patterns from data. These systems analyze source-language vocabulary, grammar, and discourse context to generate equivalent expressions in a target language, playing a central role in cross-lingual communication, digital localization, and broader multilingual language processing tasks.

80 items

The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Sasko, Quentin Lhoest, Angelina McMillan-Major, Gérard Dupont, Stella Biderman, Anna Rogers, Loubna Ben Allal, Francesco De Toni, Giada Pistilli, Olivier Nguyen, Somaieh Nikpoor, Maraim Masoud, Pierre Colombo, Javier de la Rosa, Paulo Villegas, Tristan Thrush, Shayne Longpre, Sebastian Nagel, Leon Weber, Manuel Muñoz, Jian Zhu, Daniel van Strien, Zaid Alyafeai, Khalid Almubarak, Minh Chien Vu, Itziar Gonzalez-Dios, Aitor Soroa, Kyle Lo, Manan Dey, Pedro Ortiz Suarez, Aaron Gokaslan, Shamik Bose, David Ifeoluwa Adelani, Long Phan, Hieu Tran, Ian Yu, Suhas Pai, Jenny Chim, Violette Lepercq, Suzana Ilic, Margaret Mitchell, Alexandra Sasha Luccioni, Yacine Jernite

Why you should read this

Presents the construction, governance framework, and filtering methodology behind the 1.6TB multilingual ROOTS dataset, providing a transparent foundation and reusable tools for training open language models across 59 languages.

As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop, a 1-year international and multidisciplinary initiative, was formed with the goal of researching and training large language models as a values-driven undertaking, putting issues of ethics, harm, and governance in the foreground. This paper documents the data creation and curation efforts undertaken by BigScience to assemble the Responsible Open-science Open-collaboration Text Sources (ROOTS) corpus, a 1.6TB dataset spanning 59 languages that was used to train the 176-billion-parameter BigScience Large Open-science Open-access Multilingual (BLOOM)(BigScience Workshop, 2022) language model. We further release a large initial subset of the corpus and analyses thereof, and hope to empower large-scale monolingual and multilingual modeling projects with both the data and the processing tools, as well as stimulate research around this large multilingual corpus.

Added

2026-10-05

HMoE: Heterogeneous Mixture of Experts for Language Modeling

HMoE: Heterogeneous Mixture of Experts for Language Modeling

An Wang, Xingwu Sun, Ruobing Xie, Shuaipeng Li, Jiaqi Zhu, Zhen Yang, Pinxue Zhao, Weidong Han, Zhanhui Kang, Di Wang, Naoaki Okazaki, Cheng-Zhong Xu

OrganizationsInstitute of Science TokyoTencentUniversity of Macau

Why you should read this

Proposes a heterogeneous Mixture of Experts framework featuring differently sized expert networks paired with a parameter-penalty loss, reducing activated parameter counts while surpassing standard homogeneous architectures across language modeling benchmarks.

Mixture of Experts (MoE) offers remarkable performance and computational efficiency by selectively activating subsets of model parameters. Traditionally, MoE models use homogeneous experts, each with identical capacity. However, varying complexity in input data necessitates experts with diverse capabilities, which prevents homogeneous MoE from effective expert specialization and efficient parameter utilization. In this study, we propose a novel Heterogeneous HMoE framework, where experts differ in size and thus possess diverse capacities. This heterogeneity allows for more specialized experts to handle varying token complexities more effectively. To address the imbalance in expert activation, we propose a novel training objective that encourages frequent activation of smaller experts so as to improve computational efficiency and parameter utilization. Extensive experiments demonstrate that HMoE achieves a lower loss rate with fewer activated parameters and outperforms conventional homogeneous MoE models on various pre-training evaluation benchmarks. Our codes are available at https://github.com/AnWang-AI/HMoE.

Added

2026-10-05

Rethinking Multimodal Entity and Relation Extraction from a Translation Point of View

Rethinking Multimodal Entity and Relation Extraction from a Translation Point of View

Changmeng Zheng, Junhao Feng, Yi Cai, Xiaoyong Wei, Qing Li

OrganizationsHong Kong Polytechnic UniversitySichuan UniversitySouth China University of Technology

Why you should read this

Proposes a translation-inspired framework that resolves text-image misalignment in multimodal entity and relation extraction by combining diffusion-based back-translation with divergence estimation to outperform 14 state-of-the-art baselines.

We revisit the multimodal entity and relation extraction from a translation point of view. Special attention is paid on the misalignment issue in text-image datasets which may mislead the learning. We are motivated by the fact that the cross-modal misalignment is a similar problem of cross-lingual divergence issue in machine translation. The problem can then be transformed and existing solutions can be borrowed by treating a text and its paired image as the translation to each other. We implement a multimodal back-translation using diffusion-based generative models for pseudo-paralleled pairs and a divergence estimator by constructing a high-resource corpora as a bridge for low-resource learners. Fine-grained confidence scores are generated to indicate both types and degrees of alignments with which better representations are obtained. The method has been validated in the experiments by outperforming 14 state-of-the-art methods in both entity and relation extraction tasks. The source code is available at https://github.com/thecharm/TMR.

Added

2026-10-05

Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers

Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers

Manuel Mager, Elisabeth Mager, Katharina Kann, Ngoc Thang Vu

OrganizationsAmazon Web ServicesUniversidad Nacional Autónoma de MéxicoUniversity of Colorado BoulderUniversity of Stuttgart

Why you should read this

Examines the ethical challenges of developing machine translation for Indigenous languages through direct interviews with community leaders and language activists, establishing practical guidance for respectful data collection and collaborative system design.

In recent years, machine translation has become very successful for high-resource language pairs. This has also sparked new interest in research on the automatic translation of low-resource languages, including Indigenous languages. However, the latter are deeply related to the ethnic and cultural groups that speak (or used to speak) them. The data collection, modeling and deploying machine translation systems thus result in new ethical questions that must be addressed. Motivated by this, we first survey the existing literature on ethical considerations for the documentation, translation, and general natural language processing for Indigenous languages. Afterward, we conduct and analyze an interview study to shed light on the positions of community leaders, teachers, and language activists regarding ethical concerns for the automatic translation of their languages. Our results show that the inclusion, at different degrees, of native speakers and community members is vital to performing better and more ethical research on Indigenous languages.

Added

2026-10-05

IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models

IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models

David Ifeoluwa Adelani, Jessica Ojo, Israel Abebe Azime, Jian Yun Zhuang, Jesujoba Oluwadara Alabi, Xuanli He, Millicent Ochieng, Sara Hooker, Andiswa Bukula, En-Shiun Annie Lee, Chiamaka Ijeoma Chukwuneke, Happy Buzaaba, Blessing K. Sibanda, Godson Koffi Kalipe, Jonathan Mukiibi, Salomon Kabongo Kabenamualu, Foutse Yuehgoh, Mmasibidi Setaka, Lolwethu Ndolela, Nkiruka Odu, Rooweither Mabuya, Salomey Osei, Shamsuddeen Hassan Muhammad, Sokhar Samb, Tadesse Kebede Guge, Tombekai Vangoni Sherman, Pontus Stenetorp

OrganizationsCIFARCohereConservatoire national des arts et métiersDakar American University of Science and TechnologyHaramaya UniversityImperial College LondonLancaster UniversityLeibniz Universität HannoverLelapa AIMakerere UniversityMasakhaneMcGill UniversityMicrosoftMilaOntario Tech UniversityPrinceton UniversitySaarland UniversitySADiLaRUniversidad de DeustoUniversity College LondonUniversity of Toronto

Why you should read this

Introduces IrokoBench, a human-translated benchmark across 17 African languages that evaluates 16 open and proprietary language models on natural language inference, mathematical reasoning, and question answering to measure performance disparities against high-resource languages.

Despite the widespread adoption of Large language models (LLMs), their remarkable capabilities remain limited to a few high-resource languages. Additionally, many low-resource languages (\eg African languages) are often evaluated only on basic text classification tasks due to the lack of appropriate or comprehensive benchmarks outside of high-resource languages. In this paper, we introduce IrokoBench -- a human-translated benchmark dataset for 17 typologically-diverse low-resource African languages covering three tasks: natural language inference~(AfriXNLI), mathematical reasoning~(AfriMGSM), and multi-choice knowledge-based question answering~(AfriMMLU). We use IrokoBench to evaluate zero-shot, few-shot, and translate-test settings~(where test sets are translated into English) across 10 open and six proprietary LLMs. Our evaluation reveals a significant performance gap between high-resource languages~(such as English and French) and low-resource African languages. We observe a significant performance gap between open and proprietary models, with the highest performing open model, Gemma 2 27B only at 63\% of the best-performing proprietary model GPT-4o performance. In addition, machine translating the test set to English before evaluation helped to close the gap for larger models that are English-centric, such as Gemma 2 27B and LLaMa 3.1 70B. These findings suggest that more efforts are needed to develop and adapt LLMs for African languages.

Added

2026-10-04

When Does Translation Require Context? A Data-driven, Multilingual Exploration

When Does Translation Require Context? A Data-driven, Multilingual Exploration

Patrick Fernandes, Kayo Yin, Emmy Liu, André F. T. Martins, Graham Neubig

OrganizationsCarnegie Mellon UniversityInstituto de TelecomunicaçõesInstituto Superior TécnicoLUMLIS (Lisbon ELLIS Unit)UnbabelUniversity of California Berkeley

Why you should read this

Introduces an information-theoretic methodology and multilingual benchmark spanning 14 language pairs to systematically identify context-dependent translation phenomena and assess how effectively context-aware models resolve discourse-level ambiguities.

Although proper handling of discourse significantly contributes to the quality of machine translation (MT), these improvements are not adequately measured in common translation quality metrics. Recent works in context-aware MT attempt to target a small set of discourse phenomena during evaluation, however not in a fully systematic way. In this paper, we develop the Multilingual Discourse-Aware (MuDA) benchmark, a series of taggers that identify and evaluate model performance on discourse phenomena in any given dataset. The choice of phenomena is inspired by a novel methodology to systematically identify translations requiring context. We confirm the difficulty of previously studied phenomena while uncovering others that were previously unaddressed. We find that common context-aware MT models make only marginal improvements over context-agnostic models, which suggests these models do not handle these ambiguities effectively. We release code and data for 14 language pairs to encourage the MT community to focus on accurately capturing discourse phenomena.

Added

2026-10-04

MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li

OrganizationsCarnegie Mellon UniversityNanyang Technological UniversityNational University of SingaporeNew York UniversityNorthwestern UniversityPolytechnique MontréalSingapore Management UniversitySmartor LLCUniversity College DublinUniversity of AlbertaUniversity of California BerkeleyUniversity of GenevaUniversity of New South WalesUniversity of TokyoWaseda UniversityYale University

Why you should read this

Presents MMLU-ProX, an expert-verified benchmark spanning 29 languages with parallel 10-choice questions, exposing critical cross-lingual reasoning performance drops in low-resource languages across 36 state-of-the-art large language models.

Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities. This dual limitation makes it challenging to assess LLMs’ performance in the multilingual setting comprehensively. To fill this gap, we introduce MMLU-ProX, a comprehensive benchmark covering 29 languages, built on an English benchmark. Each language version consists of 11,829 identical questions, enabling direct cross-lingual comparisons. Additionally, to meet efficient evaluation needs, we provide a lite version containing 658 questions per language. To ensure the high quality of MMLU-ProX, we employ a rigorous development process that involves multiple powerful LLMs for translation, followed by expert review to ensure accurate expression, consistent terminology, and cultural relevance. Building on this, we systematically evaluate 36 state-of-the-art LLMs, including reasoning-enhanced and multilingual-optimized LLMs. The results reveal significant disparities in the multilingual capabilities of LLMs: While they perform well in high-resource languages, their performance declines markedly in low-resource languages, particularly for African languages. Through MMLU-ProX, we aim to advance the development of more inclusive AI systems and promote equitable access to technology across global contexts.

Added

2026-10-03

The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm

The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm

Aakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant, Julia Kreutzer, Marzieh Fadaee, Sara Hooker

OrganizationsCohere

Why you should read this

Presents the Aya Red-teaming dataset across eight languages and demonstrates that Direct Preference Optimization effectively reduces both universally recognized and culturally specific harms across multiple languages without degrading general model capabilities.

A key concern with the concept of alignment is the implicit question of alignment to what? AI systems are increasingly used across the world, yet safety alignment is often focused on homogeneous monolingual settings. Additionally, preference training and safety measures often overfit to harms common in Western-centric datasets. Here, we explore the viability of different alignment approaches when balancing dual objectives: addressing and optimizing for a non-homogeneous set of languages and cultural preferences while minimizing both global and local harms. We collect the first set of human-annotated red-teaming prompts¹ in different languages distinguishing between global and local harm, which serve as a laboratory for understanding the reliability of alignment techniques when faced with preference distributions that are non-stationary across geographies and languages. While this setting is seldom covered by the literature to date, which primarily centers on English harm mitigation, it captures real-world interactions with AI systems around the world. We establish a new precedent for state-of-the-art alignment techniques across 6 languages with minimal degradation in general performance. Our work provides important insights into cross-lingual transfer and novel optimization approaches to safeguard AI systems designed to serve global populations.

Added

2026-10-03

IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages

IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages

Mohammed Safi Ur Rahman Khan, Priyam Mehta, Ananth Sankar, Umashankar Kumaravelan, Sumanth Doddapaneni, Suriyaprasaad B, Varun Balan G, Sparsh Jain, Anoop Kunchukuttan, Pratyush Kumar, Raj Dabre, Mitesh M. Khapra

Why you should read this

Presents a comprehensive open-source data processing pipeline and resource suite spanning 22 Indian languages, featuring 251 billion pre-training tokens and 74.8 million instruction-tuning pairs to support large language model development for under-resourced languages.

Despite the considerable advancements in English LLMs, the progress in building comparable models for other languages has been hindered due to the scarcity of tailored resources. Our work aims to bridge this divide by introducing an expansive suite of resources specifically designed for the development of Indic LLMs, covering 22 languages, containing a total of 251B tokens and 74.8M instruction-response pairs. Recognizing the importance of both data quality and quantity, our approach combines highly curated manually verified data, unverified yet valuable data, and synthetic data. We build a clean, open-source pipeline for curating pre-training data from diverse sources, including websites, PDFs, and videos, incorporating best practices for crawling, cleaning, flagging, and deduplication. For instruction-fine tuning, we amalgamate existing Indic datasets, translate/transliterate English datasets into Indian languages, and utilize LLaMa2 and Mixtral models to create conversations grounded in articles from Indian Wikipedia and Wikihow. Additionally, we address toxicity alignment by generating toxic prompts for multiple scenarios and then generate non-toxic responses by feeding these toxic prompts to an aligned LLaMa2 model. We hope that the datasets, tools, and resources released as a part of this work will not only propel the research and development of Indic LLMs but also establish an open-source blueprint for extending such efforts to other languages. The data and other artifacts created as part of this work are released with permissive licenses at https://github.com/AI4Bharat/IndicLLMSuite

Added

2026-10-03

Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability

Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability

Eleftheria Briakou, Colin Cherry, George F. Foster

Why you should read this

Reveals that large language models encounter millions of unintended parallel translation pairs during pre-training, demonstrating that this incidental bilingualism directly drives their zero-shot translation performance and provides data-driven prompts that boost translation quality.

Large language models exhibit impressive translation capabilities despite being trained primarily on monolingual data. In this work we study how much of their performance comes from accidental ingestion of parallel sentences present in web-scale training corpora—that include non-English text. Our extensive analysis characterizes when such exposure could lead to improved machine translation: i.e., languages seen during pre-training [6] \emph{and} content matching at inference time (§3). Our large- scale experiments reveal two major findings about target-side priming: First, exposing LLMs via naturalistic triggers – i.e., using cross-language exemplars at test-time prompts – significantly improves lexical choices, as evidenced by gains up to +4.5 chrF++†_{\text{++}}^\dagger. Second, inducing domain-specific multilingual datasets mined from real-world MT outputs to train sentence-level cross-attention adapters confirms our intuition

Added

2026-10-02

Multilingual Large Language Models Are Not (Yet) Code-Switchers

Multilingual Large Language Models Are Not (Yet) Code-Switchers

Ruochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Indra Winata, Alham Fikri Aji

Why you should read this

Reveals through systematic benchmarking across four distinct NLP tasks that prompted multilingual large language models consistently underperform substantially smaller fine-tuned models on code-switched text, demonstrating that standard multilingual pretraining fails to confer proficiency in mixed-language communication.

Multilingual Large Language Models (LLMs) have recently shown great capabilities in a wide range of tasks, exhibiting state-of-the-art performance through zero-shot or few-shot prompting methods. While there have been extensive studies on their abilities in monolingual tasks, the investigation of their potential in the context of code-switching (CSW), the practice of alternating languages within an utterance, remains relatively uncharted. In this paper, we provide a comprehensive empirical analysis of various multilingual LLMs, benchmarking their performance across four tasks: sentiment analysis, machine translation, summarization and word-level language identification. Our results indicate that despite multilingual LLMs exhibiting promising outcomes in certain tasks using zero or few-shot prompting, they still underperform in comparison to fine-tuned models of much smaller scales. We argue that current “multilingualism" in LLMs does not inherently imply proficiency with code-switching texts, calling for future research to bridge this discrepancy.

Added

2026-10-02

GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP

GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP

Md. Tawkat Islam Khondaker, Abdul Waheed, El Moatez Billah Nagoudi, Muhammad Abdul-Mageed

Why you should read this

Demonstrates through a large-scale evaluation across 44 tasks and over 60 datasets that ChatGPT and GPT-4 consistently lag behind smaller, fine-tuned dedicated models on Arabic natural language processing, particularly when handling regional dialectal varieties compared to Modern Standard Arabic.

ChatGPT’s emergence heralds a transformative phase in NLP, particularly demonstrated through its excellent performance on many English benchmarks. However, the model’s efficacy across diverse linguistic contexts remains largely uncharted territory. This work aims to bridge this knowledge gap, with a primary focus on assessing ChatGPT’s capabilities on Arabic languages and dialectal varieties. Our comprehensive study conducts a large-scale automated and human evaluation of ChatGPT, encompassing 44 distinct language understanding and generation tasks on over 60 different datasets. To our knowledge, this marks the first extensive performance analysis of ChatGPT’s deployment in Arabic NLP. Our findings indicate that, despite its remarkable performance in English, ChatGPT is consistently surpassed by smaller models that have undergone finetuning on Arabic. We further undertake a meticulous comparison of ChatGPT and GPT-4’s Modern Standard Arabic (MSA) and Dialectal Arabic (DA), unveiling the relative shortcomings of both models in handling Arabic dialects compared to MSA. Although we further explore and confirm the utility of employing GPT-4 as a potential alternative for human evaluation, our work adds to a growing body of research underscoring the limitations of ChatGPT.

Added

2026-10-02

Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better

Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better

David Dale, Elena Voita, Loïc Barrault, Marta R. Costa-jussà

Why you should read this

Demonstrates that measuring source token contribution from a translation model's internal mechanics doubles the detection accuracy of severe hallucinations without external tools, while cross-lingual sentence similarity models achieve an 80% precision gain across all hallucination types.

While the problem of hallucinations in neural machine translation has long been recognized, so far the progress on its alleviation is very little. Indeed, recently it turned out that without artificially encouraging models to hallucinate, previously existing methods fall short and even the standard sequence log-probability is more informative. It means that internal characteristics of the model can give much more information than we expect, and before using external models and measures, we first need to ask: how far can we go if we use nothing but the translation model itself ? We propose to use a method that evaluates the percentage of the source contribution to a generated translation. Intuitively, hallucinations are translations “detached” from the source, hence they can be identified by low source contribution. This method improves detection accuracy for the most severe hallucinations by a factor of 2 and is able to alleviate hallucinations at test time on par with the previous best approach that relies on external models. Next, if we move away from internal model characteristics and allow external tools, we show that using sentence similarity from cross-lingual embeddings further improves these results. We release the code of our experiments.

Added

2026-10-02

MolXPT: Wrapping Molecules with Text for Generative Pre-training

MolXPT: Wrapping Molecules with Text for Generative Pre-training

Zequn Liu, Wei Zhang, Yingce Xia, Lijun Wu, Shufang Xie, Tao Qin, Ming Zhang, Tie-Yan Liu

OrganizationsMicrosoftPeking UniversityRenmin University of ChinaUniversity of Science and Technology of China

Why you should read this

Proposes a unified generative language model pre-trained on biomedical text where molecule names are replaced with SMILES sequences, achieving superior MoleculeNet property prediction and competitive text-molecule translation with fewer parameters.

Generative pre-trained Transformer (GPT) has demonstrates its great success in natural language processing and related techniques have been adapted into molecular modeling. Considering that text is the most important record for scientific discovery, in this paper, we propose MolXPT, a unified language model of text and molecules pre-trained on SMILES (a sequence representation of molecules) wrapped by text. Briefly, we detect the molecule names in each sequence and replace them to the corresponding SMILES. In this way, the SMILES could leverage the information from surrounding text, and vice versa. The above wrapped sequences, text sequences from PubMed and SMILES sequences from PubChem are all fed into a language model for pre-training. Experimental results demonstrate that MolXPT outperforms strong baselines of molecular property prediction on MoleculeNet, performs comparably to the best model in text-molecule translation while using less than half of its parameters, and enables zero-shot molecular generation without finetuning.

Added

2026-10-01

Towards a Unified Multi-Dimensional Evaluator for Text Generation

Towards a Unified Multi-Dimensional Evaluator for Text Generation

Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, Jiawei Han

OrganizationsCarnegie Mellon UniversityMicrosoftUniversity of California, Los AngelesUniversity of Illinois Urbana-Champaign

Why you should read this

Proposes UniEval, a unified text generation evaluator that reframes multi-dimensional assessment as Boolean question answering, markedly improving correlation with human judgments across summarization and dialogue while enabling zero-shot generalization to unseen criteria.

Multi-dimensional evaluation is the dominant paradigm for human evaluation in Natural Language Generation (NLG), i.e., evaluating the generated text from multiple explainable dimensions, such as coherence and fluency. However, automatic evaluation in NLG is still dominated by similarity-based metrics, and we lack a reliable framework for a more comprehensive evaluation of advanced models. In this paper, we propose a unified multi-dimensional evaluator UniEval for NLG. We re-frame NLG evaluation as a Boolean Question Answering (QA) task, and by guiding the model with different questions, we can use one evaluator to evaluate from multiple dimensions. Furthermore, thanks to the unified Boolean QA format, we are able to introduce an intermediate learning phase that enables UniEval to incorporate external knowledge from multiple related tasks and gain further improvement. Experiments on three typical NLG tasks show that UniEval correlates substantially better with human judgments than existing metrics. Specifically, compared to the top-performing unified evaluators, UniEval achieves a 23% higher correlation on text summarization, and over 43% on dialogue response generation. Also, UniEval demonstrates a strong zero-shot learning ability for unseen evaluation dimensions and tasks. Source code, data and all pre-trained evaluators are available on our GitHub repository (this https URL).

Added

2026-09-29

LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Dongfu Jiang, Xiang Ren, Bill Yuchen Lin

OrganizationsAllen Institute for AIUniversity of Southern CaliforniaZhejiang University

Why you should read this

Proposes LLM-Blender, an ensembling framework that combines multiple open-source language models by using cross-attention pairwise ranking to select top candidate outputs and a generative fusion module to merge them into superior responses.

We present LLM-BLENDER, an ensembling framework designed to attain consistently superior performance by leveraging the diverse strengths of multiple open-source large language models (LLMs). Our framework consists of two modules: PAIRRANKER and GENFUSER, addressing the observation that optimal LLMs for different examples can significantly vary. PAIRRANKER employs a specialized pairwise comparison method to distinguish subtle differences between candidate outputs. It jointly encodes the input text and a pair of candidates, using cross-attention encoders to determine the superior one. Our results demonstrate that PAIRRANKER exhibits the highest correlation with ChatGPT-based ranking. Then, GENFUSER aims to merge the top-ranked candidates, generating an improved output by capitalizing on their strengths and mitigating their weaknesses. To facilitate large-scale evaluation, we introduce a benchmark dataset, MixInstruct, which is a mixture of multiple instruction datasets featuring oracle pairwise comparisons. Our LLM-BLENDER significantly outperform individual LLMs and baseline methods across various metrics, establishing a substantial performance gap.

Added

2026-09-28