Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multitask language understanding

Multitask language understanding is an artificial intelligence capability and evaluation paradigm where a single computational model processes, interprets, and solves a broad variety of distinct linguistic and cognitive tasks. Instead of specializing in a single isolated objective like translation or sentiment classification, a system with multitask language understanding relies on shared representations and cross-domain knowledge to perform diverse activities such as reasoning, question answering, and factual problem-solving across multiple academic, professional, and cultural subjects. This framework is widely used to assess the generalization, zero-shot or few-shot adaptability, and holistic world knowledge of modern language models across varying levels of complexity and input modalities.

3 items

In-Context Impersonation Reveals Large Language Models' Strengths and Biases

In-Context Impersonation Reveals Large Language Models' Strengths and Biases

Leonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz, Zeynep Akata

Why you should read this

Demonstrates that prompting large language models to impersonate specific personas not only boosts task performance through simulated domain expertise and age-appropriate exploration strategies, but also exposes latent demographic biases in multimodal and reasoning evaluations.

In everyday conversations, humans can take on different roles and adapt their vocabulary to their chosen roles. We explore whether LLMs can take on, that is impersonate, different roles when they generate text in-context. We ask LLMs to assume different personas before solving vision and language tasks. We do this by prefixing the prompt with a persona that is associated either with a social identity or domain expertise. In a multi-armed bandit task, we find that LLMs pretending to be children of different ages recover human-like developmental stages of exploration. In a language-based reasoning task, we find that LLMs impersonating domain experts perform better than LLMs impersonating non-domain experts. Finally, we test whether LLMs’ impersonations are complementary to visual information when describing different categories. We find that impersonation can improve performance: an LLM prompted to be a bird expert describes birds better than one prompted to be a car expert. However, impersonation can also uncover LLMs’ biases: an LLM prompted to be a man describes cars better than one prompted to be a woman. These findings demonstrate that LLMs are capable of taking on diverse roles and that this in-context impersonation can be used to uncover their strengths and hidden biases. Our code is available at https://github.com/ExplainableML/in-context-impersonation.

Added

2026-10-05

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning

Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, Adam Roberts

OrganizationsGoogle

Why you should read this

Demonstrates how training with mixed zero-shot, few-shot, and chain-of-thought prompts drives superior instruction tuning performance, releasing the Flan collection to provide more effective and computationally efficient starting checkpoints for downstream tasks.

We study the design decisions of publicly available instruction tuning methods, and break down the development of Flan 2022 (Chung et al., 2022). Through careful ablation studies on the Flan Collection of tasks and methods, we tease apart the effect of design decisions which enable Flan-T5 to outperform prior work by 3-17%+ across evaluation settings. We find task balancing and enrichment techniques are overlooked but critical to effective instruction tuning, and in particular, training with mixed prompt settings (zero-shot, few-shot, and chain-of-thought) actually yields stronger (2%+) performance in all settings. In further experiments, we show Flan-T5 requires less finetuning to converge higher and faster than T5 on single downstream tasks, motivating instruction-tuned models as more computationally-efficient starting checkpoints for new tasks. Finally, to accelerate research on instruction tuning, we make the Flan 2022 collection of datasets, templates, and methods publicly available at this https URL.

Added

2026-10-05

EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models

EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models

Rocktim Jyoti Das, Simeon Emilov Hristov, Haonan Li, Dimitar Dimitrov, Ivan Koychev, Preslav Nakov

OrganizationsFMIMohamed bin Zayed University of Artificial IntelligenceSofia University “St. Kliment Ohridski”

Why you should read this

Introduces EXAMS-V, a benchmark of over 20,000 real-world exam questions across 20 academic disciplines and 11 languages to evaluate how effectively vision-language models perform joint visual and textual reasoning on culturally diverse school subjects.

We introduce EXAMS-V, a new challenging multi-discipline multimodal multilingual exam benchmark for evaluating vision language models. It consists of 20,932 multiple-choice questions across 20 school disciplines covering natural science, social science, and other miscellaneous studies, e.g., religion, fine arts, business, etc. EXAMS-V includes a variety of multimodal features such as text, images, tables, figures, diagrams, maps, scientific symbols, and equations. The questions come in 11 languages from 7 language families. Unlike existing benchmarks, EXAMS-V is uniquely curated by gathering school exam questions from various countries, with a variety of education systems. This distinctive approach calls for intricate reasoning across diverse languages and relies on region-specific knowledge. Solving the problems in the dataset requires advanced perception and joint reasoning over the text and the visual content of the image. Our evaluation results demonstrate that this is a challenging dataset, which is difficult even for advanced vision–text models such as GPT-4V and Gemini; this underscores the inherent complexity of the dataset and its significance as a future benchmark.

Added

2026-10-03