Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models

Avni MittalShanu KumarSandipan DandapatMonojit Choudhury

article2026arXiv0 citations

Introduces a 1,500-question benchmark and a DAG-orchestrated agentic system to accurately predict multilingual model performance across target languages and tasks when direct evaluation data is missing from published literature.

Listen

Deploying multilingual artificial intelligence models often requires selecting models for tasks and languages where direct evaluation data is missing. Direct benchmark results are frequently scattered, inconsistent, or prohibitively costly to obtain, particularly for lower-resource languages. Consequently, practitioners face the challenge of estimating missing performance without full experimental coverage.

The article addresses this problem by developing a controlled evaluation benchmark to test how systems estimate missing performance from incomplete scientific literature, alongside introducing LITMUS (RE)AGENT, a graph-orchestrated multi-agent system designed for predictive multilingual evaluation.

The authors constructed a controlled benchmark containing 1,500 questions across six tasks and five distinct evidence scenarios. To assess predictive reasoning under partial information, the benchmark provides systems with access only to a reduced corpus of research papers while validating answers against a larger ground-truth corpus. Systems were evaluated on both numeric score prediction and comparative reasoning. LITMUS (RE)AGENT was evaluated alongside five baseline systems by decomposing complex queries into parallel hypotheses, retrieving relevant literature, extracting typological linguistic features, and running lightweight regression models.

The findings demonstrate that LITMUS (RE)AGENT achieved the lowest overall error in numeric prediction, recording a mean absolute error of 10.4 compared to 12.7 for its predecessor and 16.5 for direct model prompting. The system also attained the highest comparative reasoning accuracy across scenarios at 21.6%. Performance gains were largest in transfer-heavy scenarios where direct evidence was absent, such as transferring knowledge from related languages. In a human evaluation study, participants using the system achieved higher confidence, actionability, and justification ratings compared to standard tools.

These results indicate that structured, multi-agent hypothesis decomposition combined with linguistic feature modeling can substantially improve the reliability of model selection under sparse data. For decision-makers, this approach reduces the risk and expense of running exhaustive benchmark suites for new deployments. However, comparative reasoning remains challenging across all automated systems, and predictive estimates show higher variability in complex tasks like code generation and mathematical reasoning.

Organizations should use agentic predictive systems as decision-support tools for prioritization and early-stage planning rather than as direct replacements for empirical testing in high-stakes deployments. Future work should focus on reducing automated code-execution failures, decreasing system response latency, and incorporating task-specific score calibration to handle disparate performance metric ranges.

arXiv: 2604.08970

No sufficiently relevant recommendations were found.

Cover for Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models

Abstract

We study predictive multilingual evaluation: estimating how well a model will perform on a task in a target language when direct benchmark results are missing. This problem is common in multilingual deployment, where evaluation coverage is sparse and published evidence is uneven across languages, tasks, and model families. We introduce a controlled benchmark of 1,500 questions spanning six tasks and five evidence scenarios. The benchmark separates accessible evidence from ground truth, enabling evaluation of systems that must infer missing results from incomplete literature evidence. We also present Litmus (Re)Agent, a DAG-orchestrated agentic system that decomposes queries into hypotheses, retrieves evidence, and synthesises predictions through feature-aware aggregation. Across six systems, Litmus (Re)Agent achieves the best overall performance, with the largest gains in transfer-heavy scenarios where direct evidence is weak or absent. These results show that structured agentic reasoning is a promising approach to multilingual performance estimation under incomplete evidence.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Multilingual Evaluation and Benchmarks
  • 2.2 Predictive Multilingual Analysis
  • 2.3 Agentic and Literature-Grounded Systems
  • 3 Benchmark Details
  • 3.1 Benchmark Scope
  • 3.2 Query Types: Numeric Prediction and Comparative Reasoning
  • 3.3 Scenario Design
  • 3.4 Question Construction
  • 3.5 Ground Truth and Metric Normalisation
  • 3.6 Language Similarity for Transfer Scenarios
  • 3.7 Dataset Statistics
  • 4 Litmus (Re)Agent
  • 4.1 Base Architecture
  • 4.2 Improvements in Litmus (Re)Agent
  • 4.3 Evaluation Setup
  • 5 Results
  • 5.1 Overall Performance
  • 5.2 Per-Task Analysis
  • 5.3 Per-Scenario Analysis
  • 5.4 Agent Reasoning Quality
  • 5.5 Ablation Study: Backbone LLM
  • 5.6 Human Evaluation Study
  • 6 Conclusion
  • 7 Limitations
  • References
  • A Agent Lifecycle Details
  • B Systems Compared
  • B.1 LLM Hyperparameters
  • C Implementation Details
  • D Knowledge Base Curation
  • E Detailed Task–Scenario Results
  • E.1 Per-Metric MAE Analysis
  • E.2 Per-Scenario MAE Trends
  • E.3 Detailed Error Analysis by Task and Scenario
  • F Human Evaluation Details
  • F.1 Participant Demographics
  • F.2 Study Interface and Evaluation Form
  • F.3 Evaluation Questions
  • F.4 Per-Question Form Structure
  • F.5 Mid-Survey and Final Survey
  • F.6 Quality Metric Profiles
  • F.7 Self-Assessed Prediction Quality
  • F.8 Strategy Shifts
  • F.9 Challenge Distribution
  • F.10 Calibration Confidence
  • F.11 Workload and Final Survey
  • F.12 Qualitative Feedback
  • G Coder Agent Analysis
  • H Benchmark Creation
  • H.1 Scenario-wise Coverage
  • H.2 Language Distribution Across Questions
  • H.3 Model Family Distribution Across Questions
  • H.4 Language Similarity Computation Details
  • H.5 Prompt Template for Question Generation
  • H.6 LLM-as-Judge Prompt for Quality Metrics
  • H.7 Prediction Extraction Prompt
  • H.8 Thought Creator Evaluation Prompt
  • H.9 Generated Code Evaluation Prompt
  • H.10 Tool Call Relevance Evaluation Prompt
  • I Use of AI Assistants

Knowls

  1. Knowl 1 — Controlled benchmark for predicting missing multilingual results

    experimental setup

    The benchmark evaluates Task–Model–Language (TML) prediction: estimating a model’s performance on a task in a language when the relevant result is unavailable at inference time. It contains 1,500 questions spanning six tasks—code generation, mathematical reasoning, question answering/visual question answering (QA/VQA), text classification including natural language inference (NLI), text summarisation, and machine translation—and five evidence scenarios. Each task–scenario block has 50 questions: 25 numeric PREDSET questions and 25 comparative QNASET questions, for 750 questions in each subset. Questions are formed from language-to-model-family mappings; question generation is restricted to languages and model families in those mappings and excludes programming languages for code generation. The reduced paper corpus is the evidence available to systems, while ground-truth answers are extracted from a larger combined corpus and manually validated. This separation makes it possible to assess prediction when known results have been deliberately withheld.

  2. Knowl 2 — Evidence scenarios and language-transfer partition

    definition

    The benchmark’s five scenarios specify which target-language and target-model-family evidence is present in the reduced corpus. S1 (same language, same model family) provides direct evidence for the target combination; S2 (same language, different model family) supports cross-model transfer within an observed language; S3 (similar language, same model family) and S4 (distant language, same model family) support transfer to an unobserved language; S5 (different language, different model family) withholds both target dimensions and requires transfer across both. “Same model” means the same model family, not necessarily the same checkpoint. For S3 and S4, language similarity is based on cosine distance between L2-normalised lang2vec vectors using syntax_knn+fam+geo features. The close-language cutoff is the 10th percentile of pairwise distances, 0.1069: distances at or below it are close (S3), and larger distances are distant (S4). These scenarios represent different evidence configurations, not a strictly increasing difficulty scale.

  3. Knowl 3 — DAG-based architecture for evidence-grounded predictions

    model/method

    LITMUS (RE)AGENT answers a new query by having a MainAgent identify it as new or a follow-up, then using a ThoughtCreatorAgent to decompose it into hypotheses. A separate ThoughtAgent investigates each hypothesis as a node in a dynamic directed acyclic graph (DAG). Each ThoughtAgent coordinates a Research Planner, Web Search and Crawl, Expert Knowledge, Coder, and Reporter agent: respectively, these plan evidence gathering, retrieve and parse documents, consult multilingual knowledge, execute analytical scripts, and report findings. A ThoughtAnalyzerAgent monitors branches, creates new hypotheses when gaps arise, and prunes irrelevant, redundant, or divergent branches; agents can be active, completed, or discarded. After the hypotheses are resolved, a ResponseAnalyzerAgent aggregates their findings into a final answer with a prediction, supporting citations, and structured rationale. Independent hypothesis trails are intended to keep evidence traceable while limiting context-window saturation.

  4. Knowl 4 — System additions: multilingual knowledge, features, and prompting

    model/method

    Relative to the earlier LITMUS++ DAG system, LITMUS (RE)AGENT makes three changes. First, it expands the Expert Knowledge base from a small task-specific collection to broader multilingual, cross-task evaluation literature and expert-oriented notes covering language–task–model observations, transfer strategies, and failure modes; the knowledge base is guidance for analysis, not ground truth. Second, the Coder agent gains access to lang2vec and URIEL resources, supplying typological, syntactic, phonological, geographic, and family-level features for more than 7,000 languages for distance calculation, feature extraction, and feature-informed regression. Third, prompts encourage expert-style hypothesis generation, evidence retrieval, feature-based analysis, citation of evidence, and careful tool use, with the goal of reducing invalid tool calls, brittle code, and unsupported citations.

  5. Knowl 5 — Comparative evaluation protocol and baselines

    experimental setup

    All systems are evaluated against the same benchmark ground truth while retrieval is restricted to the reduced paper corpus. The comparison includes LITMUS (RE)AGENT, its predecessor LITMUS++, a ThoughtAgent group chat without DAG decomposition, a single ReAct-style agent with tools but no multi-agent coordination, direct GPT-4.1 prompting without agentic scaffolding, and Magentic-One, a general-purpose multi-agent framework. All agentic systems use GPT-4.1 as their underlying model. Numeric PREDSET answers are evaluated by mean absolute error (MAE), where lower is better; comparative QNASET answers are evaluated by accuracy, where higher is better. Ground-truth scores are placed on a 0–100 scale: values in [0,1] are multiplied by 100, other metrics receive task-specific linear scaling, and comparisons are restricted to compatible metric families within each task.

  6. Knowl 6 — Overall prediction and comparison results

    empirical result

    In the paper’s task-level results, LITMUS (RE)AGENT has the lowest overall PREDSET MAE and highest QNASET accuracy among the six systems. The reported overall figures are: LITMUS (RE)AGENT, MAE 10.4 and accuracy 21.6%; LITMUS++, 12.7 and 15.1%; ThoughtAgent, 15.1 and 15.9%; Single Agent, 14.5 and 20.1%; GPT-4.1 direct, 16.5 and 15.9%; and Magentic-One, 17.2 and 10.0%. Thus the full system improves both numeric prediction and comparative reasoning over the evaluated baselines. The table reports a 2.3-point MAE reduction relative to LITMUS++ (12.7 to 10.4). The single-agent system’s lower MAE than the non-DAG ThoughtAgent group chat also indicates that group chat alone did not consistently help numeric prediction.

  7. Knowl 7 — Performance varies with evidence configuration

    empirical result

    Across S1–S5, LITMUS (RE)AGENT’s PREDSET MAEs are 9.4, 10.3, 9.0, 11.7, and 11.9, respectively; its QNASET accuracies are 29.3%, 14.7%, 33.3%, 13.8%, and 16.6%. It has the lowest MAE in every scenario and the highest QNASET accuracy in four of five scenarios; Single Agent leads QNASET S2 with 22.0%. The system’s largest comparative-reasoning result is 33.3% in S3, transfer from a similar language, while its strongest PREDSET MAE is also in S3 (9.0), not S1 (9.4). In S4, the system’s MAE is 11.7, compared with 23.5 for Magentic-One; Magentic-One’s QNASET accuracy is 2.1% in S4 and 3.6% in S5. These outcomes show that evidence scenarios are not ordered monotonically by difficulty and that the system’s advantage includes transfer settings, particularly where linguistic signals can support prediction.

  8. Knowl 8 — Task-level strengths and remaining hard cases

    empirical result

    LITMUS (RE)AGENT’s PREDSET MAEs by task are 14.2 for code generation, 20.0 for mathematical reasoning, 9.5 for QA/VQA, 7.9 for classification/NLI, 5.1 for text summarisation, and 6.2 for machine translation. It has the lowest MAE in the reported task table for code generation, classification/NLI, and machine translation; GPT-4.1 direct is best on mathematical reasoning (18.2), LITMUS++ on QA/VQA (8.7), and Single Agent on summarisation (4.7). LITMUS (RE)AGENT’s QNASET accuracy is highest for code generation (55.3%) and classification/NLI (16.0%), but the best systems on the other tasks are GPT-4.1 for mathematical reasoning (18.7%), LITMUS++ for QA/VQA (22.4%), Single Agent for summarisation (21.8%), and ThoughtAgent for machine translation (19.2%). The paper’s metric-level analysis also reports higher LITMUS (RE)AGENT MAE for accuracy-based metrics (12.7) than for BLEU (5.0), chrF (7.5), or ROUGE (6.5). Code generation and mathematical reasoning, especially in transfer scenarios, are highlighted as comparatively difficult.

  9. Knowl 9 — Diagnostics identify code execution as a bottleneck

    empirical result

    Across 1,730 LITMUS (RE)AGENT conversations and 5,184 generated hypotheses, the reported diagnostic rates are 93.9% for thought faithfulness to expert guidance, 70.0% for capability compliance, 83.5% for web-search relevance, 91.6% for feature correctness, and 30.6% for successful code execution. Permitting retries raises code-execution success to 54.6%, though execution remains a weakness. In a separate LLM-judge assessment using 1–5 scores, LITMUS (RE)AGENT receives the highest reported averages for predictive plausibility (4.6) and feature selection (4.8); GPT-4.1 direct leads coherence (5.0) and citation emphasis (4.3). The diagnostics therefore show strong hypothesis alignment, retrieval relevance, and feature selection alongside a substantial execution constraint.

  10. Knowl 10 — Small human study finds better-rated predictions with the system

    empirical result

    A within-subjects study involved eight participants (four novices and four experts) making code-generation performance predictions under the same constrained paper-corpus setting. Participants answered five questions using general LLM tools and five using LITMUS (RE)AGENT, then rated prediction confidence, interpretability, explainability, justification, and actionability. In the aggregate, mean ratings increased from baseline to the system condition for confidence (2.6 to 3.4), interpretability (3.1 to 3.7), explainability (3.2 to 3.9), justification (2.8 to 3.8), and actionability (2.6 to 3.5), on the study’s Likert scales. Time per question increased overall from 8.6 to 12.0 minutes; novice time rose from 5.9 to 13.7 minutes, while expert time changed from 10.7 to 10.3 minutes. Five of seven participants in the final survey said the system improved prediction accuracy. The small sample and study setting limit how broadly these results can be generalised.

  11. Knowl 11 — Scope and reliability limitations

    limitation

    The authors identify limitations in the system, benchmark, and baseline comparison. All agentic experiments use GPT-4.1, so generalisation to smaller or open-source backbones is unverified; code execution succeeds in only 30.6% of conversations without retries, and end-to-end latency is higher than direct prompting. The benchmark covers six tasks but excludes multimodal reasoning, dialogue, and safety, depends on a fixed curated corpus, and uses paper-reported ground truth that may contain measurement noise and cross-study inconsistencies. Differences in coordination strategies among LITMUS++, ThoughtAgent, and Magentic-One mean that measured gains over them do not isolate architecture from prompt-level effects. The paper positions predictions as decision-support for prioritisation and early analysis, not as substitutes for direct benchmark measurements in high-stakes decisions, and notes that estimates may inherit publication and language-coverage biases.

Coverage note — Full prompt templates, implementation and deployment details, participant-level records, and supplementary per-question plots are omitted because they support the main benchmark, system, and evaluation findings rather than adding separate load-bearing contributions.

References

  1. 1.Alessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran, Subhransu Maji, Charless C Fowlkes, Stefano Soatto, and Pietro Perona. 2019. Task2vec: Task embedding for meta-learning. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6430–6439.
  2. 2.Kabir Ahuja, Sandipan Dandapat, Sunayana Sitaram, and Monojit Choudhury. 2022a. Beyond static models and test sets: Benchmarking the potential of pretrained models across tasks and languages. In Proceedings of NLP Power! The First Workshop on Efficient Benchmarking in NLP, pages 64–74.
  3. 3.Kabir Ahuja, Shanu Kumar, Sandipan Dandapat, and Monojit Choudhury. 2022b. Multi task learning for zero shot performance prediction of multilingual models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5454–5467.
  4. 4.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, and 1 others. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  5. 5.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, and 1 others. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  6. 6.Błazej Dolicki and Gerasimos Spanakis. 2021. Analysing the impact of linguistic features on cross-lingual transfer. arXiv preprint arXiv:2105.05975.
  7. 7.Adam Fourney, Gyu Hong, Lance J.Y. Liew, Jianping Chi, Ziyang Wang, Ryen White, Eric Horvitz, Mounia Lalmas, Faraaz Ansari, Chris Ackerson, Ben Suh, Javier Diaz, Yoni Halpern, Mark Dredze, Jacob Eisenstein, Monica S. Lam, and Nataniel Ruiz. 2024. Magentic-one: A generalist multi-agent system for solving complex tasks. arXiv preprint arXiv:2410.04468.
  8. 8.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation. In International conference on machine learning, pages 4411–4421. PMLR.
  9. 9.Shanu Kumar, Soujanya Abbaraju, Sandipan Dandapat, Sunayana Sitaram, and Monojit Choudhury. 2023. Ditto: A feature representation imitation approach for improving cross-lingual transfer. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 385–406.
  10. 10.Anne Lauscher, Vinit Ravishankar, Ivan Vulic, and Goran Glavaš. 2020. From zero to hero: On the limitations of zero-shot language transfer with multilingual transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4483–4499.
  11. 11.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K"uttler, Mike Lewis, Wen-tau Yih, Tim Rockt"aschel, and Sebastian Riedel. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33:9459–9474.
  12. 12.Zihao Li, Yucheng Shi, Zirui Liu, Fan Yang, Ali Payani, Ninghao Liu, and Mengnan Du. 2025. Language ranker: A metric for quantifying llm performance across high and low-resource languages. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 28186–28194.
  13. 13.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, and 1 others. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.
  14. 14.Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, and 1 others. 2020. Xglue: A new benchmark dataset for cross-lingual pre-training, understanding and generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6008–6018.
  15. 15.Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, and 1 others. 2019. Choosing transfer languages for cross-lingual learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3125–3135.
  16. 16.Patrick Littell, David R Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin. 2017. Uriel and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 8–14.
  17. 17.Qian Liu, Xiaoting Wang, Yining Hou, Bohan Lyu, Huanliang Liu, Yunhao Liu, Haowen Jiang, Chaojia Gu, Hanqing Lu, Xintong Han, Weiliang Dai, Cailong Cai, Yizhou Wang, Zhiyuan Liu, Lei Hou, and Juanzi Li. 2023. Paperqa: Citation-aware question answering over scientific literature. arXiv preprint arXiv:2310.12321.
  18. 18.Yixin Liu, Guibin Zhang, Kun Wang, Shiyuan Li, and Shirui Pan. 2025. Graph-augmented large language model agents: Current progress and future prospects. arXiv preprint arXiv:2507.21407.
  19. 19.Avni Mittal, Shanu Kumar, Sandipan Dandapat, and Monojit Choudhury. 2025. Litmus++: An agentic system for predictive analysis of low-resource languages across tasks and models. In Proceedings of The 14th International Joint Conference on Natural Language Processing and The 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics: System Demonstrations, pages 47–54.
  20. 20.Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 conference on empirical methods in natural language processing, pages 1797–1807.
  21. 21.Cuong Nguyen, Tal Hassner, Matthias Seeger, and Cedric Archambeau. 2020. Leep: A new measure to evaluate transferability of learned representations. In International conference on machine learning, pages 7294–7305. PMLR.
  22. 22.NLLB Team. 2022. No language left behind: Scaling human-centered machine translation. https://arxiv.org/abs/2207.04672.
  23. 23.OpenAI. 2023. Gpt-4 technical report. https://cdn.openai.com/papers/gpt-4.pdf.
  24. 24.Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and 1 others. 2021. Xtreme-r: Towards more challenging and nuanced multilingual evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10215–10245.
  25. 25.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. Advances in neural information processing systems, 36:68539–68551.
  26. 26.Viktoria Schram, Daniel Beck, and Trevor Cohn. 2023. Performance prediction via bayesian matrix factorisation for multilingual natural language processing tasks. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 1790–1801.
  27. 27.Anirudh Srinivasan, Gauri Kholkar, Rahul Kejriwal, Tanuja Ganu, Sandipan Dandapat, Sunayana Sitaram, Balakrishnan Santhanam, Somak Aditya, Kalika Bali, and Monojit Choudhury. 2022. Litmus predictor: An ai assistant for building reliable, high-performing and fair multilingual nlp systems. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 13227–13229.
  28. 28.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, and 1 others. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on machine learning research.
  29. 29.Soumava Tan, Lipika Dey, Kushagra Chawla, Sumeet Kumar, Mihir Rajora, Rishabh Mehrotra, Ahmed Awadallah, and Mitesh M. Khapra. 2024. Judgebench: A benchmark for evaluating llm-based judges. https://openreview.net/forum?id=G0dksFayVq.
  30. 30.Alexander Tsvetkov and Alon Kipnis. 2024. Information parity: Measuring and predicting the multilingual capabilities of language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 7971–7989.
  31. 31.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. In The eleventh international conference on learning representations.
  32. 32.Kaichao You, Yong Liu, Jianmin Wang, and Mingsheng Long. 2021. Logme: Practical assessment of pretrained models for transfer learning. In International conference on machine learning, pages 12133–12143. PMLR.
  33. 33.Tianhao Zhang, Hui Lu, Yilin Zhang, Saibin Zhang, Yuhang Jiang, Huanxin Wang, and Peiyan Liu. 2024a. Deepresearcher: Autonomous agents for end-to-end literature review and reasoning. arXiv preprint arXiv:2402.06789.
  34. 34.Yifan Zhang, Yang Yuan, and Andrew Chi-Chih Yao. 2024b. On the diagram of thought. arXiv preprint arXiv:2409.10038.
  35. 35.Mingchen Zhuge, Wenyi Wang, Louis Kirsch, Francesco Faccio, Dmitrii Khizbullin, and Jürgen Schmidhuber. 2024. Gptswarm: Language agents as optimizable graphs. In Forty-first International Conference on Machine Learning.

Citation

MLA
Mittal, A., et al. “Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models”. arXiv, 2026, http://arxiv.org/abs/2604.08970v1.
APA
Mittal, A., Kumar, S., Dandapat, S., & Choudhury, M. (2026). Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models. arXiv. http://arxiv.org/abs/2604.08970v1
Chicago
Mittal, A., S. Kumar, S. Dandapat, and M. Choudhury. 2026. “Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models”. arXiv. http://arxiv.org/abs/2604.08970v1.
Harvard
Mittal, A. et al. (2026) “Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.08970v1.
Vancouver
1. Mittal A, Kumar S, Dandapat S, Choudhury M (2026) Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models. arXiv

BibTeX

@article{mittal2026litmus,
  title = {Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models},
  author = {Mittal, Avni and Kumar, Shanu and Dandapat, Sandipan and Choudhury, Monojit},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.08970v1},
  eprint = {2604.08970}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/