PromptBench: A Unified Library for Evaluation of Large Language Models

Kaijie ZhuQinlin ZhaoHao ChenJindong WangXing Xie

article2024JMLR75 citations

Introduces PromptBench, an open-source evaluation framework that unifies model benchmarking, prompt engineering, multi-level adversarial attacks, and dynamic testing protocols to rigorously assess large language models.

Listen

As large language models become increasingly integrated into critical areas like healthcare and education, evaluating their true capabilities and safety has become urgent. Existing evaluation tools are fragmented and struggle with issues like prompt sensitivity, vulnerability to adversarial inputs, and benchmark data contamination, which introduce major security and reliability risks.

The article presents PromptBench, an open-source, unified Python library designed to evaluate language models across diverse capabilities, prompt engineering techniques, adversarial attacks, and dynamic testing protocols.

The developers designed an extensible, modular architecture supporting both open-source models (such as Llama 2, Mistral, and Mixtral) and proprietary systems (including ChatGPT and GPT-4). The framework covers 12 broad task categories across 22 standard datasets, spanning text understanding, reasoning, mathematics, and multimodal processing. It implements testing protocols that apply four levels of prompt manipulation—character, word, sentence, and semantic variations—alongside dynamic data generation to assess robustness and prevent memorization.

Empirical assessments conducted through the library yield several key findings. First, all evaluated language models demonstrate sensitivity and performance drops when subjected to adversarial prompt modifications, though advanced proprietary systems like GPT-4 and ChatGPT maintain higher overall resilience. Second, prompt engineering techniques such as Chain-of-Thought, EmotionPrompt, and Expert Prompting provide notable performance improvements in targeted domains, but no single technique consistently outperforms standard baselines across all tasks. Third, dynamic evaluation reveals that even leading frontier models like GPT-4 still experience substantial performance degradation on complex algorithmic problems, abductive logic, and linear equation tasks.

These findings indicate that relying on standard, static benchmarks creates a false sense of security regarding model reliability and production readiness. Organizations deploying language models risk unexpected operational failures or security vulnerabilities if systems are not tested against noisy or adversarial real-world inputs. For decision-makers, comprehensive benchmarking must become an integral part of risk management and compliance rather than an afterthought.

Organizations should adopt standardized, multi-faceted evaluation pipelines that include adversarial stress-testing and dynamic prompt variations before deploying models into production. Teams must avoid one-size-fits-all prompt engineering and instead validate specific prompting techniques against concrete target tasks. Developers and researchers should leverage extensible frameworks like PromptBench to continuously assess new models and contribute domain-specific benchmarks.

While PromptBench provides broad coverage, its findings depend on the quality of underlying datasets, and some evaluation metrics may fail to capture subtle qualitative variations in model outputs. Decision-makers can have high confidence in the framework's broad robustness trends, but should exercise caution and conduct domain-specific testing when deploying systems in specialized, high-stakes environments.

No sufficiently relevant recommendations were found.

Cover for PromptBench: A Unified Library for Evaluation of Large Language Models

Abstract

The evaluation of large language models (LLMs) is crucial to assess their performance and mitigate potential security risks. In this paper, we introduce PromptBench, a unified library to evaluate LLMs. It consists of several key components that can be easily used and extended by researchers: prompt construction, prompt engineering, dataset and model loading, adversarial prompt attack, dynamic evaluation protocols, and analysis tools. PromptBench is designed as an open, general, and flexible codebase for research purpose. It aims to facilitate original study in creating new benchmarks, deploying downstream applications, and designing new evaluation protocols. The code is available at: https://github.com/microsoft/promptbench and will be continuously supported.

Table of Contents

  • 1. Introduction
  • 2. PromptBench
  • 2.1 Components
  • 2.2 Evaluation pipeline
  • 2.3 Supported research topics
  • 3. Conclusion and Discussion
  • References
  • Appendix A. Comparison with Related Code Libraries
  • Appendix B. Details of PromptBench
  • B.1 Models
  • B.2 Tasks and Datasets
  • B.3 Evaluation protocols
  • B.4 Prompts
  • B.4.1 Prompts
  • B.4.2 Prompt Engineering
  • B.4.3 Adversarial Prompt Attacks
  • B.5 Pipeline
  • Appendix C. Benchmark Results
  • C.1 Adversarial prompt robustness
  • C.2 Prompt engineering
  • C.3 Dynamic evaluation
  • Appendix D. Extensibility
  • D.1 Add new datasets
  • D.2 Add new models
  • D.3 Add new prompt engineering methods
  • D.4 Add new metrics and input/output process functions

Knowls

  1. Knowl 1 — PromptBench is a modular research toolkit for LLM evaluation

    model/method

    PromptBench is an open-source Python library for building and extending evaluation pipelines for large language models (LLMs) and vision-language models (VLMs). Its modules cover model and dataset loading, prompt construction and engineering, adversarial prompt attacks, evaluation protocols, and result analysis. The framework is designed to let researchers combine these components for standard evaluation as well as research on prompt robustness and dynamic evaluation.

  2. Knowl 2 — Evaluation protocols include direct, dynamic, and semantic testing

    model/method

    PromptBench supports three evaluation protocols. Standard evaluation runs models directly on a fixed test dataset. Dynamic evaluation uses DyVal to generate complexity-tailored samples on the fly, including seven reasoning tasks: arithmetic, linear equations, boolean logic, deductive logic, abductive logic, reachability, and maximum-sum paths. Semantic evaluation uses MSTemp to generate out-of-distribution examples with evaluator LLMs and word replacement. The framework is open to adding further protocols.

  3. Knowl 3 — A four-stage API constructs an evaluation pipeline

    algorithm

    A PromptBench evaluation pipeline consists of four stages: load a task dataset with DatasetLoader; instantiate a model through LLMModel; define one or more prompts with the Prompt interface, using dataset-specific defaults if none are supplied; and define input processing, output processing, and an evaluation metric. The input and output functions are provided through InputProcess and OutputProcess, and metrics through Eval. The model interface supports generation settings such as maximum new tokens and temperature. Users can pass a list of prompts to compare their performance on the same dataset and model.

  4. Knowl 4 — Adversarial prompt testing covers four perturbation levels

    model/method

    PromptBench provides attack interfaces for testing how LLM performance changes under prompt perturbations intended to resemble user errors or variation in expression. Character-level attacks introduce typos or other character changes; word-level attacks replace words with synonyms or contextually similar alternatives; sentence-level attacks append irrelevant or redundant sentences; and semantic-level attacks imitate linguistic styles associated with different regions. The framework also accepts curated adversarial prompts for robustness evaluation.

  5. Knowl 5 — Prompt formats and engineering methods support controlled comparisons

    model/method

    PromptBench provides task-oriented prompts, which state the requested task, and role-oriented prompts, which assign the model a role such as expert or translator. Either format can be used in zero-shot or few-shot settings; the paper's few-shot examples randomly select three training examples per task. The library also implements six prompt-engineering methods: Chain-of-Thought, Zero-Shot Chain-of-Thought, EmotionPrompt, Expert Prompting, Generated Knowledge, and Least-to-Most. These respectively elicit intermediate reasoning, append “Let's think step by step,” add emotional stimuli, condition answers on a generated expert identity, generate knowledge before answering, and decompose a problem into sequential subproblems. Users can also define custom prompts.

  6. Knowl 6 — The library spans diverse models, tasks, and datasets

    experimental setup

    PromptBench provides unified interfaces for open-source and proprietary LLMs and VLMs, including custom models, and lets users specify generation settings such as maximum new tokens and temperature. The authors report support for 12 tasks and 22 public datasets. Task coverage ranges from sentiment analysis, grammatical acceptability, duplicate-sentence detection, and natural-language inference to reading comprehension, translation, mathematics, and logical, commonsense, symbolic, and algorithmic reasoning. The included dataset collection also covers multimodal tasks such as visual question answering, image captioning, and diagram or chart reasoning. A unified dataset loader provides customizable dataset loading and processing.

  7. Knowl 7 — Analysis utilities and leaderboards support result interpretation and comparison

    model/method

    PromptBench includes sweep-running utilities for collecting benchmark results, attention visualization, and word-frequency analysis of words used in attacks. It also supports defense analysis by integrating word-correction tools. The platform provides leaderboards for adversarial prompt attacks, prompt engineering, and dynamic evaluation so that results can be compared and new results submitted.

  8. Knowl 8 — Reported evaluations show prompt sensitivity and task-dependent method effects

    empirical result

    In the paper's adversarial robustness evaluations, all tested models are reported as vulnerable to adversarial prompts; ChatGPT and GPT-4 show the strongest robustness among the compared models. In prompt-engineering comparisons, methods can help on particular datasets, but no method outperforms the baseline on every dataset. In dynamic evaluations across seven reasoning tasks, GPT-4 outperforms the other compared models overall, while the authors identify linear equations, abductive logic, and maximum-sum paths as areas where performance could improve.

  9. Knowl 9 — Each major module has a defined extension path

    algorithm

    PromptBench modules can be extended by implementing a component and registering it with the relevant interface. A new dataset class inherits from Dataset, loads its data, and is registered with DataLoader. A new model inherits from LLMModel and implements its tokenizer and model; it may define a custom prediction function, otherwise the inherited default is used, and the model is registered with the model-creation interface. A new prompt-engineering method implements init and query, inherits from the shared base class, stores its prompts in the designated prompt collection, and is registered in the method map. New metrics and input/output processors are added as static functions to Eval, InputProcess, or OutputProcess, respectively.

  10. Knowl 10 — Evaluation coverage and measurement quality limit conclusions

    limitation

    The authors caution that PromptBench may not cover every evaluation scenario and that some metrics may fail to capture nuanced performance differences. Its effectiveness also depends on the quality and diversity of the datasets and prompts used. These limitations qualify conclusions drawn from evaluations conducted with the framework.

Coverage note — The exhaustive inventory of individual supported models and datasets, and the paper's example prompt wordings, are omitted as catalog detail; the framework's coverage and configurable interfaces are captured above.

References

  1. 1.Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson. Nocaps: Novel object captioning at scale. In Proceedings of the IEEE/CVF international conference on computer vision, pages 8948–8957, 2019.
  2. 2.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  3. 3.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023.
  4. 4.BIG bench authors. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=uyTL5Bvosj.
  5. 5.BerriAI. https://docs.litellm.ai/, 2023.
  6. 6.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. Gpt-neox-20b: An open-source autoregressive language model, 2022. URL https://arxiv.org/abs/2204.06745.
  7. 7.Google Brain. A new open source flan 20b with ul2, 2023. URL https://www.yitay.net/blog/flan-ul2-20b.
  8. 8.Mauro Cettolo, Marcello Federico, Luisa Bentivogli, Jan Niehues, Sebastian St¨uker, Katsuhito Sudoh, Koichiro Yoshino, and Christian Federmann. Overview of the IWSLT 2017 evaluation campaign. In Proceedings of the 14th International Conference on Spoken Language Translation, pages 2–14, Tokyo, Japan, December 14-15 2017. International Workshop on Spoken Language Translation. URL https://aclanthology.org/2017.iwslt-1.1.
  9. 9.Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. A survey on evaluation of large language models. arXiv preprint arXiv:2307.03109, 2023.
  10. 10.Harrison Chase. Langchain. https://github.com/langchain-ai/langchain, 2022. Date released: 2022-10-17.
  11. 11.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023. URL https://lmsys.org/blog/2023-03-30-vicuna/.
  12. 12.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. Scaling instruction-finetuned language models, 2022. URL https://arxiv.org/abs/2210.11416.
  13. 13.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  14. 14.Contributors. promptfoo: test your llm app, 2023a. URL https://github.com/promptfoo/promptfoo.
  15. 15.Evals Contributors. Openai evals, 2023b. URL https://github.com/openai/evals.
  16. 16.OpenCompass Contributors. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023c.
  17. 17.Databricks. Hello dolly: Democratizing the magic of chatgpt with open models, 2023. URL https://www.databricks.com/blog/2023/03/24/hello-dolly-democratizing-magic-chatgpt-open-models.html.
  18. 18.Nolan Dey, Gurpreet Gosal, Zhiming, Chen, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, and Joel Hestness. Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster, 2023.
  19. 19.Shizhe Diao, Zhichao Huang, Ruijia Xu, Xuechun Li, Yong Lin, Xiao Zhou, and Tong Zhang. Black-box prompt learning for pre-trained language models. arXiv preprint arXiv:2201.08531, 2022.
  20. 20.William B. Dolan and Chris Brockett. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005), 2005. URL https://aclanthology.org/I05-5002.
  21. 21.Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Xilin Wei, Songyang Zhang, Haodong Duan, Maosong Cao, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Xinyue Zhang, Wei Li, Jingwen Li, Kai Chen, Conghui He, Xingcheng Zhang, Yu Qiao, Dahua Lin, and Jiaqi Wang. Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model. arXiv preprint arXiv:2401.16420, 2024.
  22. 22.Andreas Eisele and Yu Chen. MultiUN: A multilingual corpus from united nation documents. In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10), Valletta, Malta, May 2010. European Language Resources Association (ELRA). URL http://www.lrec-conf.org/proceedings/lrec2010/pdf/686_Paper.pdf.
  23. 23.Michael Eisenstein. A test of artificial intelligence. Nature Outlook: Robotics and artificial intelligence, 2023.
  24. 24.J. Gao, J. Lanchantin, M. L. Soffa, and Y. Qi. Black-box generation of adversarial text sequences to evade deep learning classifiers. In 2018 IEEE Security and Privacy Workshops (SPW), pages 50–56, May 2018. doi: 10.1109/SPW.2018.00016.
  25. 25.Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework for few-shot language model evaluation, 12 2023. URL https://zenodo.org/records/10256836.
  26. 26.Google. https://deepmind.google/technologies/gemini/#introduction, 2023.
  27. 27.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913, 2017.
  28. 28.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=d7KBjmI3GmQ.
  29. 29.Namgyu Ho, Laura Schmid, and Se-Young Yun. Large language models are reasoning teachers. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, pages 14852–14882. Association for Computational Linguistics, 2023. URL https://aclanthology.org/2023.acl-long.830.
  30. 30.Jordan Hoffmann et al. Training compute-optimal large language models, 2022.
  31. 31.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, et al. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. arXiv preprint arXiv:2305.08322, 2023.
  32. 32.HuggingFace. Open-source large language models leaderboard. https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard, 2023.
  33. 33.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L´elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timoth´ee Lacroix, and William El Sayed. Mistral 7b, 2023.
  34. 34.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, L´elio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Th´eophile Gervet, Thibaut Lavril, Thomas Wang, Timoth´ee Lacroix, and William El Sayed. Mixtral of experts, 2024.
  35. 35.Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. Is bert really robust? natural language attack on text classification and entailment. arXiv preprint arXiv:1907.11932, 2019.
  36. 36.Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14, pages 235–251. Springer, 2016.
  37. 37.Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal. Qasc: A dataset for question answering via sentence composition, 2020.
  38. 38.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=e2TBb5y0yFf.
  39. 39.Hector Levesque, Ernest Davis, and Leora Morgenstern. The winograd schema challenge. In Thirteenth international conference on the principles of knowledge representation and reasoning, 2012.
  40. 40.Cheng Li, Jindong Wang, Yixuan Zhang, Kaijie Zhu, Wenxin Hou, Jianxun Lian, Fang Luo, Qiang Yang, and Xing Xie. Large language models understand and can be enhanced by emotional stimuli, 2023a.
  41. 41.Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. TextBugger: Generating adversarial text against real-world applications. In Proceedings 2019 Network and Distributed System Security Symposium. Internet Society, 2019. doi: 10.14722/ndss.2019.23138. URL https://doi.org/10.14722%2Fndss.2019.23138.
  42. 42.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR, 2023b.
  43. 43.Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. BERT-ATTACK: Adversarial attack against BERT using BERT. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6193–6202, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.500. URL https://aclanthology.org/2020.emnlp-main.500.
  44. 44.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. Alpacaeval: An automatic evaluator of instruction-following models. Github repository, 2023c.
  45. 45.Yuanzhi Li, S´ebastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee. Textbooks are all you need ii: phi-1.5 technical report. arXiv preprint arXiv:2309.05463, 2023d.
  46. 46.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022.
  47. 47.Bill Yuchen Lin, Seyeon Lee, Rahul Khanna, and Xiang Ren. Birds have four legs?! numersense: Probing numerical commonsense knowledge of pre-trained language models, 2020.
  48. 48.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning, 2023a.
  49. 49.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36, 2024.
  50. 50.Jerry Liu. LlamaIndex, 11 2022. URL https://github.com/jerryjliu/llama_index.
  51. 51.Jiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck, Peter West, Ronan Le Bras, Yejin Choi, and Hannaneh Hajishirzi. Generated knowledge prompting for commonsense reasoning, 2022.
  52. 52.Yachuan Liu, Liang Chen, Jindong Wang, Qiaozhu Mei, and Xing Xie. Meta semantic template for evaluation of large language models. arXiv preprint arXiv:2310.01448, 2023b.
  53. 53.Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. In The 36th Conference on Neural Information Processing Systems (NeurIPS), 2022.
  54. 54.Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023.
  55. 55.Ahmed Masry, Do Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, pages 2263–2279, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.findings-acl.177. URL https://aclanthology.org/2022.findings-acl.177.
  56. 56.Daniel McDuff, Mike Schaekermann, Tao Tu, Anil Palepu, Amy Wang, Jake Garrison, Karan Singhal, Yash Sharma, Shekoofeh Azizi, Kavita Kulkarni, Le Hou, Yong Cheng, Yun Liu, S Sara Mahdavi, Sushant Prakash, Anupam Pathak, Christopher Semturs, Shwetak Patel, Dale R Webster, Ewa Dominowska, Juraj Gottweis, Joelle Barral, Katherine Chou, Greg S Corrado, Yossi Matias, Jake Sunshine, Alan Karthikesalingam, and Vivek Natarajan. Towards accurate differential diagnosis with large language models, 2023.
  57. 57.Microsoft. Semantic kernel. https://github.com/microsoft/semantic-kernel, 2023.
  58. 58.Mixtral. Mixtral, 2023. URL https://mistral.ai/news/mixtral-of-experts/.
  59. 59.Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. Stress test evaluation for natural language inference. In ACL, pages 2340–2353, Santa Fe, New Mexico, USA, August 2018. Association for Computational Linguistics. URL https://aclanthology.org/C18-1198.
  60. 60.OpenAI. https://chat.openai.com.chat, 2023a.
  61. 61.OpenAI. Gpt-4 technical report, 2023b.
  62. 62.Rui Pan, Shuo Xing, Shizhe Diao, Xiang Liu, Kashun Shum, Jipeng Zhang, and Tong Zhang. Plum: Prompt learning using metaheuristic. arXiv preprint arXiv:2311.08364, 2023.
  63. 63.Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal. Grips: Gradient-free, edit-based instruction search for prompting large language models. arXiv preprint arXiv:2203.07281, 2022.
  64. 64.Pranav Rajpurkar, Robin Jia, and Percy Liang. Know what you don’t know: Unanswerable questions for SQuAD. In ACL, pages 784–789, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-2124. URL https://aclanthology.org/P18-2124.
  65. 65.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. Beyond accuracy: Behavioral testing of NLP models with CheckList. In ACL, pages 4902–4912, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.442. URL https://aclanthology.org/2020.acl-main.442.
  66. 66.David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. Analysing mathematical reasoning abilities of neural models. In ICLR, 2019. URL https://openreview.net/forum?id=H1gR5iR5FX.
  67. 67.Gabriel Simmons. Moral mimicry: Large language models produce moral rationalizations tailored to political identity. arXiv preprint arXiv:2209.12106, 2022.
  68. 68.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP, pages 1631–1642, Seattle, Washington, USA, October 2013. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/D13-1170.
  69. 69.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge, 2019.
  70. 70.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023.
  71. 71.AJ Thirunavukarasu, S Mahmood, A Malem, WP Foster, R Sanghera, and R Hassan. Large language models approach expert-level clinical knowledge and reasoning in ophthalmology: A head-to-head cross-sectional study. PLOS Digital Health, 3(4):e0000341, 2024. doi: 10.1371/journal.pdig.0000341.
  72. 72.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023a.
  73. 73.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  74. 74.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. GLUE: A multi-task benchmark and analysis platform for natural language understanding. 2019. In the Proceedings of ICLR.
  75. 75.Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.
  76. 76.Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al. Decodingtrust: A comprehensive assessment of trustworthiness in gpt models. arXiv preprint arXiv:2306.11698, 2023a.
  77. 77.Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, et al. On the robustness of chatgpt: An adversarial and out-of-distribution perspective. In International conference on learning representations (ICLR) workshop on Trustworthy and Reliable Large-Scale Machine Learning Models, 2023b.
  78. 78.Zhiguo Wang, Wael Hamza, and Radu Florian. Bilateral multi-perspective matching for natural language sentences, 2017.
  79. 79.Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. Neural network acceptability judgments. arXiv preprint arXiv:1805.12471, 2018.
  80. 80.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models, 2023.
  81. 81.Adina Williams, Nikita Nangia, and Samuel Bowman. A broad-coverage challenge corpus for sentence understanding through inference. In NAACL HLT, pages 1112–1122. Association for Computational Linguistics, 2018. URL http://aclweb.org/anthology/N18-1101.
  82. 82.Moritz Willig, Matej Zecevic, Devendra Singh Dhami, and Kristian Kersting. Causal parrots: Large language models may talk causality but are not causal. Transactions on machine learning research (TMLR), 8, 2023.
  83. 83.Benfeng Xu, An Yang, Junyang Lin, Quan Wang, Chang Zhou, Yongdong Zhang, and Zhendong Mao. Expertprompting: Instructing large language models to be distinguished experts, 2023.
  84. 84.Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, Fei Deng, Feng Wang, Feng Liu, Guangwei Ai, Guosheng Dong, Haizhou Zhao, Hang Xu, Haoze Sun, Hongda Zhang, Hui Liu, Jiaming Ji, Jian Xie, JunTao Dai, Kun Fang, Lei Su, Liang Song, Lifeng Liu, Liyun Ru, Luyao Ma, Mang Wang, Mickel Liu, MingAn Lin, Nuolan Nie, Peidong Guo, Ruiyang Sun, Tao Zhang, Tianpeng Li, Tianyu Li, Wei Cheng, Weipeng Chen, Xiangrong Zeng, Xiaochuan Wang, Xiaoxi Chen, Xin Men, Xin Yu, Xuehai Pan, Yanjun Shen, Yiding Wang, Yiyu Li, Youxin Jiang, Yuchen Gao, Yupeng Zhang, Zenan Zhou, and Zhiying Wu. Baichuan 2: Open large-scale language models, 2023.
  85. 85.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601, 2023.
  86. 86.Yi. Yi, 2023. URL https://github.com/01-ai/Yi.
  87. 87.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arXiv preprint arXiv:2311.16502, 2023.
  88. 88.Zeno. https://zenoml.com/, 2023.
  89. 89.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023.
  90. 90.Denny Zhou, Nathanael Sch¨arli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, and Ed Chi. Least-to-most prompting enables complex reasoning in large language models, 2023a.
  91. 91.Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han. Don’t make your llm an evaluation benchmark cheater. arXiv preprint arXiv:2311.01964, 2023b.
  92. 92.Kaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong, Diyi Yang, and Xing Xie. Dyval: Graph-informed dynamic evaluation of large language models. arXiv preprint arXiv:2309.17167, 2023a.
  93. 93.Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, et al. Promptbench: Towards evaluating the robustness of large language models on adversarial prompts. arXiv preprint arXiv:2306.04528, 2023b.

Citation

MLA
Zhu, K., et al. “PromptBench: A Unified Library for Evaluation of Large Language Models”. Journal of Machine Learning Research, vol. 25, no. 254, 2024, pp. 1–2, https://www.jmlr.org/papers/v25/24-0023.html.
APA
Zhu, K., Zhao, Q., Chen, H., Wang, J., & Xie, X. (2024). PromptBench: A Unified Library for Evaluation of Large Language Models. Journal of Machine Learning Research, 25(254), 1–22. https://www.jmlr.org/papers/v25/24-0023.html
Chicago
Zhu, K., Q. Zhao, H. Chen, J. Wang, and X. Xie. 2024. “PromptBench: A Unified Library for Evaluation of Large Language Models”. Journal of Machine Learning Research 25 (254): 1–22. https://www.jmlr.org/papers/v25/24-0023.html.
Harvard
Zhu, K. et al. (2024) “PromptBench: A Unified Library for Evaluation of Large Language Models”, Journal of Machine Learning Research, 25(254), pp. 1–22. Available at: https://www.jmlr.org/papers/v25/24-0023.html.
Vancouver
1. Zhu K, Zhao Q, Chen H, Wang J, Xie X (2024) PromptBench: A Unified Library for Evaluation of Large Language Models. Journal of Machine Learning Research 25:1–22

BibTeX

@article{JMLR:v25:24-0023,
  author  = {Kaijie Zhu and Qinlin Zhao and Hao Chen and Jindong Wang and Xing Xie},
  title   = {PromptBench: A Unified Library for Evaluation of Large Language Models},
  journal = {Journal of Machine Learning Research},
  year    = {2024},
  volume  = {25},
  number  = {254},
  pages   = {1--22},
  url     = {http://jmlr.org/papers/v25/24-0023.html}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/