Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
Avni MittalShanu KumarSandipan DandapatMonojit Choudhury
Introduces a 1,500-question benchmark and a DAG-orchestrated agentic system to accurately predict multilingual model performance across target languages and tasks when direct evaluation data is missing from published literature.
Deploying multilingual artificial intelligence models often requires selecting models for tasks and languages where direct evaluation data is missing. Direct benchmark results are frequently scattered, inconsistent, or prohibitively costly to obtain, particularly for lower-resource languages. Consequently, practitioners face the challenge of estimating missing performance without full experimental coverage.
The article addresses this problem by developing a controlled evaluation benchmark to test how systems estimate missing performance from incomplete scientific literature, alongside introducing LITMUS (RE)AGENT, a graph-orchestrated multi-agent system designed for predictive multilingual evaluation.
The authors constructed a controlled benchmark containing 1,500 questions across six tasks and five distinct evidence scenarios. To assess predictive reasoning under partial information, the benchmark provides systems with access only to a reduced corpus of research papers while validating answers against a larger ground-truth corpus. Systems were evaluated on both numeric score prediction and comparative reasoning. LITMUS (RE)AGENT was evaluated alongside five baseline systems by decomposing complex queries into parallel hypotheses, retrieving relevant literature, extracting typological linguistic features, and running lightweight regression models.
The findings demonstrate that LITMUS (RE)AGENT achieved the lowest overall error in numeric prediction, recording a mean absolute error of 10.4 compared to 12.7 for its predecessor and 16.5 for direct model prompting. The system also attained the highest comparative reasoning accuracy across scenarios at 21.6%. Performance gains were largest in transfer-heavy scenarios where direct evidence was absent, such as transferring knowledge from related languages. In a human evaluation study, participants using the system achieved higher confidence, actionability, and justification ratings compared to standard tools.
These results indicate that structured, multi-agent hypothesis decomposition combined with linguistic feature modeling can substantially improve the reliability of model selection under sparse data. For decision-makers, this approach reduces the risk and expense of running exhaustive benchmark suites for new deployments. However, comparative reasoning remains challenging across all automated systems, and predictive estimates show higher variability in complex tasks like code generation and mathematical reasoning.
Organizations should use agentic predictive systems as decision-support tools for prioritization and early-stage planning rather than as direct replacements for empirical testing in high-stakes deployments. Future work should focus on reducing automated code-execution failures, decreasing system response latency, and incorporating task-specific score calibration to handle disparate performance metric ranges.
- Paper: General Scales Unlock AI Evaluation with Explanatory and Predictive Power, Lexin Zhou et al. (2025). Its task-demand scales and instance-level assessors establish a closely related approach to predicting model performance beyond observed benchmark results.
- Paper: BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer, Akari Asai et al. (2024). BUFFET makes cross-lingual transfer across diverse tasks and languages measurable, grounding the transfer-heavy scenarios that Litmus (Re)Agent predicts.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). MEGA provides a broad multilingual evaluation baseline across tasks and languages, clarifying the coverage gaps that motivate predictive evaluation.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). XNLI established a standardized cross-lingual transfer benchmark, giving useful context for evaluating multilingual performance when target-language evidence is sparse.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). MEGAVERSE’s wide multilingual benchmark and documented performance disparities help situate the uneven evidence that the source’s predictions must synthesize.
No sufficiently relevant recommendations were found.
