Dynamic Evaluation of Large Language Models by Meta Probing Agents

Kaijie ZhuJindong WangQinlin ZhaoRuochen XuXing Xie

article2024ICML64 citations

Proposes a psychometrics-inspired dynamic evaluation framework using collaborative probing and judging agents to systematically transform existing benchmarks, exposing widespread performance drops from data contamination and uncovering strong correlations among underlying cognitive abilities in large language models.

Listen

Assessing large language models through static public benchmarks increasingly misleads decision-makers due to data contamination, where models inadvertently memorize publicly available test questions during training rather than demonstrating genuine reasoning. Furthermore, standard aggregate test scores fail to isolate underlying cognitive competencies. The article introduces and evaluates Meta Probing Agents, an automated, psychometrically grounded dynamic evaluation framework designed to measure true model capabilities and break the static nature of standard benchmarks.

The framework deploys a collaborative agent architecture: a probing agent transforms standard benchmark questions across three core cognitive dimensions—language understanding, problem solving, and domain knowledge—using techniques such as paraphrasing, adding extraneous context, and introducing plausible distractor choices. A judge agent then adversarially validates whether each revised question retains semantic equivalence and correctness relative to the original. The evaluation tested leading proprietary models (such as GPT-4-Turbo and Gemini-Pro) and open-source models (such as Llama2-70b-chat and Mixtral-8x7b-Instruct) across established benchmarks including MMLU, ARC-C, GSM8K, and subsets of BigBench-Hard. Human expert verification confirmed the validity of the dynamically generated questions with a 94% semantic equivalence rate and a 97% correctness rate.

The evaluation revealed several critical findings. First, all models suffered substantial performance declines on dynamic benchmarks compared to their baseline scores; for instance, GPT-4-Turbo experienced a 15.54 percentage point drop on MMLU and an 11.49 point drop on ARC-C, indicating that standard benchmark performance is inflated by memorization. Second, open-source models exhibited higher rates of performance degradation when questions were altered, showing greater vulnerability to contamination. Third, standard prompt engineering techniques, such as Chain-of-Thought and In-Context Learning, provided only marginal improvements and failed to recover lost accuracy. Fourth, fine-grained analysis revealed strong cross-ability correlations, particularly between language understanding and problem solving, alongside a Matthew effect where larger models demonstrated tighter integration across all three cognitive dimensions. Finally, a pilot study demonstrated that using questions generated by the framework for fine-tuning improved model accuracy by an average of 2% on test benchmarks.

These findings imply that relying on published benchmark figures introduces significant risk when deploying models into complex, dynamic environments, as static scores mask core capability gaps in professional and ethical domains such as law, ethics, and psychology. Organizations should avoid relying solely on static public leaderboards for model selection and procurement. Instead, technical leaders should adopt dynamic, multi-agent probing frameworks for pre-deployment audits and explore probing-based data augmentation to fine-tune and strengthen model reasoning.

While the findings are well supported across thousands of test instances, current implementation depends heavily on advanced models like GPT-4-Turbo to reliably generate and judge questions, as weaker models introduce unintended drift in question meaning. Decision-makers should treat current benchmark rankings with healthy skepticism and consider expanding dynamic probing across broader, domain-specific evaluation sets before executing high-stakes deployments.

Cover for Dynamic Evaluation of Large Language Models by Meta Probing Agents

Abstract

Evaluation of large language models (LLMs) has raised great concerns in the community due to the issue of data contamination. Existing work designed evaluation protocols using well-defined algorithms for specific tasks, which cannot be easily extended to diverse scenarios. Moreover, current evaluation benchmarks can only provide the overall benchmark results and cannot support a fine-grained and multifaceted analysis of LLMs' abilities. In this paper, we propose meta probing agents (MPA), a general dynamic evaluation protocol inspired by psychometrics to evaluate LLMs. MPA designs the probing and judging agents to automatically transform an original evaluation problem into a new one following psychometric theory on three basic cognitive abilities: language understanding, problem solving, and domain knowledge. These basic abilities are also dynamically configurable, allowing multifaceted analysis. We conducted extensive evaluations using MPA and found that most LLMs achieve poorer performance, indicating room for improvement. Our multifaceted analysis demonstrated the strong correlation between the basic abilities and an implicit Matthew effect on model size, i.e., larger models possess stronger correlations of the abilities. MPA can also be used as a data augmentation approach to enhance LLMs. Code is available at: https://github.com/microsoft/promptbench.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Overview
  • 3.2. Probing Agent
  • 3.3. Judge Agent
  • 3.4. Psychometric principles
  • 3.4.1. LANGUAGE UNDERSTANDING
  • 3.4.2. PROBLEM SOLVING
  • 3.4.3. DOMAIN KNOWLEDGE
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Main Results
  • 4.3. Effect of Different Probing Principles
  • 4.4. Ablation Study on Prompt Engineering
  • 4.5. Albation Study on Data Contamination
  • 4.6. Weak LLMs as Probing and Judge Agents
  • 5. Multifaceted Analysis of the Basic Abilities
  • 5.1. Analysis on Benchmark Complexity
  • 5.2. Relationship of the Basic Abilities
  • 5.3. Analysis on Model Size
  • 5.4. Error Analysis
  • 6. MPA as Data Augmentation for LLMs
  • 7. Conclusion and Discussion
  • Impact Statement
  • References
  • A. Datasets
  • B. Evaluation Prompts
  • C. Detailed Results
  • C.1. Standard Deviation of Main Results
  • C.2. Results of Different Modular Principles
  • C.3. Results of Relationship of the Basic Abilities
  • C.4. Top topics of MMLU
  • C.5. Examples of wrong/ambiguous evaluation samples
  • D. Examples generated by MPA
  • E. Human verification results

Knowls

  1. Knowl 1 — Meta Probing Agents Evaluation Protocol

    model/method

    Meta Probing Agents (MPA) is an agent-based dynamic evaluation framework for large language models (LLMs) that transforms existing evaluation questions based on psychometric cognitive ability principles. The architecture consists of two interacting LLM agents:

    1. Probing Agent: Given an original evaluation question QQ from an existing benchmark, the probing agent transforms QQ into a semantically equivalent but structurally altered question according to a specified psychometric principle pip_i, preserving the core concept, knowledge domain, and correct answer.
    2. Judge Agent: Evaluates the transformed question against the original question QQ in an adversarial verification step using a prompt-based binary response ('Yes' or 'No'). It assesses whether the new question assesses the exact same underlying concept/knowledge area without unintentional drift or ambiguity.

    If the judge agent outputs 'No', a feedback loop prompts the probing agent to regenerate the question. When the judge agent outputs 'Yes', the transformed question is accepted into the dynamic evaluation set. In experimental implementations, GPT-4-Turbo is deployed as both the probing agent (sampling temperature 0.70.7) and the judge agent (temperature 00), with a maximum generation length of 10001000 tokens per agent.

  2. Knowl 2 — Psychometric Probing Principles for LLM Ability Decomposition

    definition

    The Meta Probing Agents (MPA) framework establishes five transformation principles mapped directly to the three basic cognitive abilities identified in psychometric theory:

    1. Language Understanding (LU): Evaluates whether an LLM maintains semantic comprehension when linguistic representations vary:
      • Principle 1 (p1p_1) - Paraphrasing Questions: Rephrases the prompt text of a question while preserving its underlying concepts.
      • Principle 2 (p2p_2) - Paraphrasing Choices: Rephrases candidate multiple-choice options while preserving their original meanings and distractors.
      • Principle 3 (p3p_3) - Permuting Choices: Randomly shuffles the order of multiple-choice options to test for position bias (implemented deterministically via code without an agent).
    2. Problem Solving (PS): Evaluates critical thinking, deduction, and filtering irrelevant information:
      • Principle 4 (p4p_4) - Adding Extra Context into Questions: Inserts non-essential background details into the question that are topical but unhelpful for deriving the solution, testing whether the model can filter out noise.
    3. Domain Knowledge (DK): Evaluates domain depth and fine-grained conceptual differentiation:
      • Principle 5 (p5p_5) - Adding a New Choice: Introduces an additional plausible, topic-relevant, but incorrect option (e.g., a 5th option E in 4-choice items) requiring domain knowledge to rule out.
  3. Knowl 3 — Relative Effectiveness of Probing Principles

    equation

    To evaluate the isolated impact of individual probing principles relative to the cumulative performance drop under all principles combined, the Relative Effectiveness (RERE) of principle pip_i is computed as:

    RE=Accpi−AccAccpall−AccRE = \frac{\text{Acc}_{p_i} - \text{Acc}}{\text{Acc}_{p_{\text{all}}} - \text{Acc}}

    where:

    • Acc\text{Acc} is the evaluated model's accuracy on the vanilla (original) benchmark dataset.
    • Accpi\text{Acc}_{p_i} is the model's accuracy when only principle pip_i is applied by the probing agent.
    • Accpall\text{Acc}_{p_{\text{all}}} is the model's accuracy on the probing benchmark when all applicable principles are applied simultaneously.

    A higher RERE value denotes that principle pip_i is responsible for a larger portion of the overall benchmark performance drop.

  4. Knowl 4 — Performance Degradation of LLMs Under MPA Dynamic Benchmarks

    empirical result

    Evaluating LLMs across dynamic probing benchmarks constructed with Meta Probing Agents (MPA) causes significant accuracy decreases compared to vanilla benchmark scores on MMLU, GSM8K, ARC-Challenge (ARC-C), and BigBench-Hard (BBH partial subsets: Formal Fallacies, Object Counting, Temporal Sequences), indicating benchmark memorization and potential data contamination. Evaluation settings use zero-shot generation with temperature 00 (except Gemini-Pro on MMLU at 0.70.7) and maximum generation length of 10001000 tokens.

    Dataset GPT-4-Turbo GPT-3.5-Turbo Gemini Pro Yi-34b Mixtral-8x7b Llama2-70b-chat
    Vanilla Ours (Δ\Delta) Vanilla Ours (Δ\Delta) Vanilla Ours (Δ\Delta) Vanilla Ours (Δ\Delta) Vanilla Ours (Δ\Delta) Vanilla Ours (Δ\Delta)
    MMLU 84.40 68.86 (-15.54) 68.12 56.15 (-11.97) 67.04 55.55 (-11.49) 67.31 63.30 (-4.01) 66.49 55.24 (-11.25) 56.85 49.70 (-7.15)
    GSM8K 95.22 88.50 (-6.72) 77.71 71.54 (-6.17) 22.97 20.39 (-2.58) 73.54 68.54 (-5.00) 61.56 47.18 (-14.38) 52.92 51.50 (-1.42)
    ARC-C 96.16 84.67 (-11.49) 85.41 74.60 (-10.81) 86.18 75.91 (-10.27) 86.78 74.03 (-12.75) 84.47 70.36 (-14.11) 73.55 64.19 (-9.36)
    BBH 88.53 87.78 (-0.75) 54.67 49.73 (-4.94) 65.47 60.00 (-5.47) 55.47 52.49 (-2.98) 53.47 40.53 (-12.94) 38.53 38.22 (-0.31)

    Knowledge-intensive benchmarks (MMLU and ARC-C) exhibit more severe accuracy drops (e.g., GPT-4-Turbo dropping 15.54%15.54\% on MMLU and 11.49%11.49\% on ARC-C) than mathematical and logical reasoning benchmarks. Confusion matrix analysis reveals a high proportion of 'Original True / Probing False' (OT/PF) responses, with open-source models (Llama2-70b-chat at 28%28\%, Mixtral-8x7b at 25%25\%, Yi-34b at 23%23\%) exhibiting higher OT/PF rates than GPT-4-Turbo (13%13\%).

  5. Knowl 5 — Correlation Between Basic Cognitive Abilities in LLMs

    empirical result

    Evaluating large language models across separate probing benchmarks dedicated to each individual cognitive ability—Language Understanding (LU: p1,p2,p3p_1, p_2, p_3), Problem Solving (PS: p4p_4), and Domain Knowledge (DK: p5p_5) on MMLU and ARC-C—shows that all three cognitive dimensions are strongly correlated across models.

    Ability Pair Pearson Correlation Spearman Correlation Kendall Correlation
    LU PS 0.994 0.986 0.939
    LU DK 0.986 0.972 0.909
    PS DK 0.986 0.979 0.909

    The correlation between Language Understanding and Problem Solving is the highest across all metrics (r=0.994r = 0.994, ρ=0.986\rho = 0.986, τ=0.939\tau = 0.939), indicating that language comprehension and problem-solving reasoning in LLMs are strongly predictive of each other.

  6. Knowl 6 — Model Scaling and the Matthew Effect on Cognitive Ability Correlations

    empirical result

    Analyzing model size scaling on MPA probing benchmarks reveals two distinct structural effects:

    1. Uniform Scaling Slopes: Testing Llama2 variants (7B, 13B, and 70B parameters) on LU, PS, and DK benchmarks demonstrates that performance across each basic cognitive ability increases with model size along almost identical slopes.
    2. Implicit Matthew Effect on Inter-Ability Correlation: Categorizing models into three capacity groups:
      • Small: Llama2-7b-chat, Llama2-13b-chat
      • Mid: Yi-34b-chat, Mixtral-8x7b-Instruct, Llama2-70b-chat
      • Large: Gemini-Pro, GPT-3.5-Turbo, GPT-4-Turbo

    shows that larger and stronger models possess systematically higher correlation coefficients among their basic cognitive abilities (LU-PS, LU-DK, PS-DK correlations rise from ≈0.96\approx 0.96 in small models to ≈0.98\approx 0.98 in mid-size models and >0.99> 0.99 in large models). This mirrors Spearman's psychometric theory of the general intelligence gg-factor.

  7. Knowl 7 — Monotonic Accuracy Degradation Under Compounded Probing Complexity

    empirical result

    Constructing dynamic probing benchmarks with increasing levels of operational complexity leads to monotonic decreases in accuracy across all evaluated LLMs (GPT-4-Turbo, GPT-3.5-Turbo, Gemini-Pro) on ARC-C and MMLU:

    1. Level 1 (LU): Applying Language Understanding principle 1 (paraphrasing questions).
    2. Level 2 (LU+PS): Jointly applying Language Understanding principle 1 and Problem Solving principle 4 (adding extra context).
    3. Level 3 (LU+PS+DK): Jointly applying Language Understanding principle 1, Problem Solving principle 4, and Domain Knowledge principle 5 (adding a new choice).

    On both ARC-C and MMLU, model accuracy drops monotonically from Vanilla →\to Level 1 →\to Level 2 →\to Level 3. On ARC-C, GPT-4-Turbo accuracy declines from 96.16%96.16\% (Vanilla) to ≈85%\approx 85\% (Level 3); GPT-3.5-Turbo declines from 85.41%85.41\% to ≈75%\approx 75\%; and Gemini-Pro declines from 86.18%86.18\% to ≈76%\approx 76\%. GPT-4-Turbo consistently achieves the highest precision across all complexity levels.

  8. Knowl 8 — Prompt Engineering Inefficacy on Dynamic Probing Benchmarks

    empirical result

    Standard prompt engineering techniques—Chain-of-Thought (CoT) prompting and 5-shot In-Context Learning using original training exemplars (ICLo\text{ICL}_o) or MPA-transformed exemplars (ICLt\text{ICL}_t)—fail to consistently recover the accuracy losses induced by MPA dynamic probing on GSM8K and ARC-C.

    Dataset Model Original Probing CoT ICLo_o ICLt_t
    GSM8K GPT-4-Turbo 88.50 89.31 89.39 88.78
    GSM8K GPT-3.5-Turbo 71.54 70.58 65.73 64.90
    GSM8K Gemini-Pro 20.39 24.49 75.28 73.01
    ARC-C GPT-4-Turbo 84.67 82.85 85.32 85.67
    ARC-C GPT-3.5-Turbo 74.60 75.94 74.32 74.49
    ARC-C Gemini-Pro 75.91 76.02 76.71 79.10

    Except for Gemini-Pro on GSM8K (where few-shot formatting compensates for poor zero-shot instruction following), prompting methods yield negligible improvements (e.g., GPT-4-Turbo improves by less than 1%1\% on GSM8K and drops from 84.67%84.67\% to 82.85%82.85\% with CoT on ARC-C), indicating that the performance degradation reflects underlying cognitive difficulty rather than prompt format sensitivity.

  9. Knowl 9 — Meta Probing Agents as a Data Augmentation Method for Fine-Tuning

    empirical result

    Synthesizing transformed training examples with Meta Probing Agents (MPA) serves as an effective data augmentation technique for fine-tuning LLMs. Fine-tuning GPT-3.5-Turbo using the original training sets of MMLU and ARC-C augmented with transformed questions generated across the five probing principles yields an average accuracy improvement of approximately 2%2\% on both vanilla test sets and MPA probing test sets:

    • MMLU Vanilla Test: Accuracy increases from 68.12%68.12\% (Vanilla GPT-3.5-Turbo) to ≈70%\approx 70\% (Fine-Tuned model).
    • MMLU Probing Test: Accuracy increases from 56.15%56.15\% to ≈58%\approx 58\%.
    • ARC-C Vanilla Test: Accuracy increases from 85.41%85.41\% to ≈87%\approx 87\%.
    • ARC-C Probing Test: Accuracy increases from 74.60%74.60\% to ≈77%\approx 77\%.
  10. Knowl 10 — Dependency on Frontier LLMs for Probing and Judge Agents

    limitation

    The reliability of the Meta Probing Agents (MPA) framework depends on the cognitive and instruction-following capability of the underlying agent models:

    1. Judging Shortfalls in Weaker LLMs: When GPT-3.5-Turbo or Gemini-Pro are used as judge agents alongside a GPT-4-Turbo probing agent on GSM8K, manual verification reveals they frequently fail to detect semantic divergence and allow corrupted questions to pass verification.
    2. Probing Inaccuracies in Weaker LLMs: When GPT-3.5-Turbo or Gemini-Pro serve as probing agents alongside a GPT-4-Turbo judge agent, they often alter nuances, distort constraints, or introduce unintended factual errors into questions. Even GPT-4-Turbo as a judge agent fails to consistently detect these subtle discrepancies.

    Consequently, high-capability frontier models (such as GPT-4-Turbo) are necessary to execute the MPA framework reliably without degrading benchmark validity.

  11. Knowl 11 — Human Verification of MPA Question Equivalence and Correctness

    empirical result

    A human evaluation study conducted by 30 expert annotators (divided into 3 groups of 10) evaluating 1,800 question pairs (500 from MMLU and 100 from ARC-C across each of the three cognitive abilities) verified the validity of MPA transformations under two binary criteria:

    1. Semantic Equivalence: Whether the transformed question assessed the same concept as the original question.
    2. Answer Correctness: Whether the ground-truth answer remained accurate and unique for the transformed question.
    Metric (Equivalence / Correctness) Language Understanding Problem Solving Domain Knowledge Average
    Group 1 0.88 / 0.95 0.91 / 0.93 1.00 / 1.00 0.93 / 0.96
    Group 2 0.92 / 0.99 0.94 / 0.98 1.00 / 0.97 0.95 / 0.98
    Group 3 0.91 / 0.95 0.92 / 0.99 1.00 / 1.00 0.94 / 0.98
    Average 0.90 / 0.96 0.92 / 0.97 1.00 / 0.99 0.94 / 0.97

    The overall average equivalence was 94%94\% and the overall average correctness was 97%97\%, confirming high structural and semantic validity.

Coverage note — Detailed topic-level frequency rankings across individual MMLU sub-tasks (e.g. professional law, moral scenarios) and specific narrative question examples from GSM8K/BBH were omitted as they serve as qualitative illustrations of the aggregate empirical results.

References

  1. 1.01-ai. Yi: A series of large language models. https://github.com/01-ai/Yi, 2024.
  2. 2.Bai, Y., Ying, J., Cao, Y., Lv, X., He, Y., Wang, X., Yu, J., Zeng, K., Xiao, Y., Lyu, H., et al. Benchmarking foundation models with language-model-as-an-examiner. arXiv preprint arXiv:2306.04181, 2023.
  3. 3.bench authors, B. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research, 2023. ISSN 2835-8856.
  4. 4.Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. On the dangers of stochastic parrots: Can language models be too big? FAccT 2021, pp. 610–623, New York, NY, USA, 2021. Association for Computing Machinery. ISBN 9781450383097.
  5. 5.Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O. The reversal curse: Llms trained on “a is b” fail to learn “b is a”. arXiv preprint arXiv:2309.12288, 2023.
  6. 6.Biderman, S., Prashanth, U. S., Sutawika, L., Schoelkopf, H., Anthony, Q., Purohit, S., and Raf, E. Emergent and predictable memorization in large language models. arXiv preprint arXiv:2304.11158, 2023.
  7. 7.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  8. 8.Burnell, R., Hao, H., Conway, A. R., and Orallo, J. H. Revealing the structure of language model capabilities. arXiv preprint arXiv:2306.10062, 2023.
  9. 9.Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. arXiv preprint arXiv:2312.09390, 2023.
  10. 10.Chang, Y., Wang, X., Wang, J., Wu, Y., Zhu, K., Chen, H., Yang, L., Yi, X., Wang, C., Wang, Y., et al. A survey on evaluation of large language models. arXiv preprint arXiv:2307.03109, 2023.
  11. 11.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  12. 12.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  13. 13.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  14. 14.Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P., and Hashimoto, T. B. Alpacafarm: A simulation framework for methods that learn from human feedback. arXiv preprint arXiv:2305.14387, 2023.
  15. 15.Fan, L., Hua, W., Li, L., Ling, H., Zhang, Y., and Hemphill, L. Nphardeval: Dynamic benchmark on reasoning ability of large language models via complexity classes. arXiv preprint arXiv:2312.14890, 2023.
  16. 16.Fernandes, P., Deutsch, D., Finkelstein, M., Riley, P., Martins, A. F., Neubig, G., Garg, A., Clark, J. H., Freitag, M., and Firat, O. The devil is in the errors: Leveraging large language models for fine-grained machine translation evaluation. arXiv preprint arXiv:2308.07286, 2023.
  17. 17.Gao, I., Ilharco, G., Lundberg, S., and Ribeiro, M. T. Adaptive testing of computer vision models. arXiv preprint arXiv:2212.02774, 2022.
  18. 18.GeminiTeam. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  19. 19.Golchin, S. and Surdeanu, M. Data contamination quiz: A tool to detect and estimate contamination in large language models. arXiv preprint arXiv:2311.06233, 2023a.
  20. 20.Golchin, S. and Surdeanu, M. Time travel in llms: Tracing data contamination in large language models. arXiv preprint arXiv:2308.08493, 2023b.
  21. 21.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021.
  22. 22.Hong, S., Zheng, X., Chen, J., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., et al. Metagpt: Meta programming for multi-agent collaborative framework. arXiv preprint arXiv:2308.00352, 2023.
  23. 23.HuggingFace. Open-source large language models leaderboard. https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard, 2023.
  24. 24.Kendall, M. G. A new measure of rank correlation. Biometrika, 30(1/2):81–93, 1938.
  25. 25.Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., Ma, Z., Thrush, T., Riedel, S., Waseem, Z., Stenetorp, P., Jia, R., Bansal, M., Potts, C., and Williams, A. Dynabench: Rethinking benchmarking in NLP. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 4110–4124, June 2021.
  26. 26.Kocoń, J., Cichecki, I., Kaszyca, O., Kochanek, M., Szydło, D., Baran, J., Bielaniewicz, J., Gruza, M., Janz, A., Kanclerz, K., et al. Chatgpt: Jack of all trades, master of none. Information Fusion, pp. 101861, 2023.
  27. 27.Lei, F., Liu, Q., Huang, Y., He, S., Zhao, J., and Liu, K. S3eval: A synthetic, scalable, systematic evaluation suite for large language models. arXiv preprint arXiv:2310.15147, 2023.
  28. 28.Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval, 2023a.
  29. 29.Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B. Alpacaeval: An automatic evaluator of instruction-following models, 2023b.
  30. 30.Li, Y. An open source data contamination report for llama series models. arXiv preprint arXiv:2310.17589, 2023.
  31. 31.Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C. A., Manning, C. D., Re, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., WANG, J., Santhanam, K., Orr, L., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N., Guha, N., Chatterji, N. S., Khattab, O., Henderson, P., Huang, Q., Chi, R. A., Xie, S. M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., and Koreeda, Y. Holistic evaluation of language models. Transactions on Machine Learning Research, 2023. ISSN 2835-8856.
  32. 32.Liu, H. and Yao, A. C.-C. Augmenting math word problems via iterative question composing. arXiv preprint arXiv:2401.09003, 2024.
  33. 33.Lovin, B. Gpt-4 performs significantly worse on coding problems not in its training data. https://brianlovin.com/hn/35297067, 2023.
  34. 34.Ma, Z., Ethayarajh, K., Thrush, T., Jain, S., Wu, L., Jia, R., Potts, C., Williams, A., and Kiela, D. Dynaboard: An evaluation-as-a-service platform for holistic next-generation benchmarking. Advances in Neural Information Processing Systems, 34:10351–10367, 2021.
  35. 35.Merton, R. K. The matthew effect in science: The reward and communication systems of science are considered. Science, 159(3810):56–63, 1968.
  36. 36.MistralAITeam. Mixtral-8x7b-v0.1. https://huggingface.co/mistralai/Mixtral-8x7B-v0.1, 2023.
  37. 37.OpenAI. https://chat.openai.com.chat, 2023a.
  38. 38.OpenAI. Gpt-4 technical report, 2023b.
  39. 39.Oren, Y., Meister, N., Chatterji, N., Ladhak, F., and Hashimoto, T. B. Proving test set contamination in black box language models. arXiv preprint arXiv:2310.17623, 2023.
  40. 40.Pearson, K. The history and theory of correlation. Biometrika Office, 1920.
  41. 41.Raykov, T. and Marcoulides, G. A. Introduction to psychometric theory. Routledge, 2011.
  42. 42.Ribeiro, M. T. and Lundberg, S. Adaptive testing and debugging of nlp models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3253–3267, 2022.
  43. 43.Ribeiro, M. T., Wu, T., Guestrin, C., and Singh, S. Beyond accuracy: Behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4902–4912, Online, July 2020. Association for Computational Linguistics.
  44. 44.Schaeffer, R., Miranda, B., and Koyejo, S. Are emergent abilities of large language models a mirage? In NeurIPS, 2023.
  45. 45.Significant-Gravitas. Autogpt. https://github.com/Significant-Gravitas/AutoGPT, 2023.
  46. 46.Spearman, C. The proof and measurement of association between two things. The American Journal of Psychology, 15(1):72–101, 1904.
  47. 47.Spearman, C. ” general intelligence” objectively determined and measured. 1961.
  48. 48.Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022.
  49. 49.Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., , and Wei, J. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
  50. 50.Tian, Y., Pei, K., Jana, S., and Ray, B. Deeptest: Automated testing of deep-neural-network-driven autonomous cars. In Proceedings of the 40th international conference on software engineering, pp. 303–314, 2018.
  51. 51.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  52. 52.Wang, B., Chen, W., Pei, H., Xie, C., Kang, M., Zhang, C., Xu, C., Xiong, Z., Dutta, R., Schaeffer, R., et al. Decodingtrust: A comprehensive assessment of trustworthiness in gpt models. arXiv preprint arXiv:2306.11698, 2023.
  53. 53.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560, 2022.
  54. 54.Wang, Y., Yu, Z., Zeng, Z., Yang, L., Wang, C., Chen, H., Jiang, C., Xie, R., Wang, J., Xie, X., Ye, W., Zhang, S., and Zhang, Y. Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization. 2024.
  55. 55.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35: 24824–24837, 2022.
  56. 56.Wei, T., Zhao, L., Zhang, L., Zhu, B., Wang, L., Yang, H., Li, B., Cheng, C., Lu, W., Hu, R., et al. Skywork: A more open bilingual foundation model. arXiv preprint arXiv:2310.19341, 2023.
  57. 57.Yang, L., Zhang, S., Qin, L., Li, Y., Wang, Y., Liu, H., Wang, J., Xie, X., and Zhang, Y. Glue-x: Evaluating natural language understanding models from an out-of-distribution generalization perspective. arXiv preprint arXiv:2211.08073, 2022.
  58. 58.Yang, S., Chiang, W.-L., Zheng, L., Gonzalez, J. E., and Stoica, I. Rethinking benchmark and contamination for language models with rephrased samples. arXiv preprint arXiv:2311.04850, 2023.
  59. 59.Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J. T., Li, Z., Weller, A., and Liu, W. Metamath: Bootstrap your own mathematical questions for large language models. arXiv preprint arXiv:2309.12284, 2023.
  60. 60.Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J. Self-rewarding language models. arXiv preprint arXiv:2401.10020, 2024.
  61. 61.Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N. Agieval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364, 2023.
  62. 62.Zhou, K., Zhu, Y., Chen, Z., Chen, W., Zhao, W. X., Chen, X., Lin, Y., Wen, J.-R., and Han, J. Don’t make your llm an evaluation benchmark cheater. arXiv preprint arXiv:2311.01964, 2023.
  63. 63.Zhu, K., Chen, J., Wang, J., Gong, N. Z., Yang, D., and Xie, X. Dyval: Graph-informed dynamic evaluation of large language models. arXiv preprint arXiv:2309.17167, 2023a.
  64. 64.Zhu, K., Wang, J., Zhou, J., Wang, Z., Chen, H., Wang, Y., Yang, L., Ye, W., Gong, N. Z., Zhang, Y., et al. Promptbench: Towards evaluating the robustness of large language models on adversarial prompts. arXiv preprint arXiv:2306.04528, 2023b.
  65. 65.Zong, Y., Yu, T., Zhao, B., Chavhan, R., and Hospedales, T. Fool your (vision and) language model with embarrassingly simple permutations. arXiv preprint arXiv:2310.01651, 2023.

Citation

MLA
Zhu, K., et al. “Dynamic Evaluation of Large Language Models by Meta Probing Agents”. arXiv, 2024, http://arxiv.org/abs/2402.14865v2.
APA
Zhu, K., Wang, J., Zhao, Q., Xu, R., & Xie, X. (2024). Dynamic Evaluation of Large Language Models by Meta Probing Agents. arXiv. http://arxiv.org/abs/2402.14865v2
Chicago
Zhu, K., J. Wang, Q. Zhao, R. Xu, and X. Xie. 2024. “Dynamic Evaluation of Large Language Models by Meta Probing Agents”. arXiv. http://arxiv.org/abs/2402.14865v2.
Harvard
Zhu, K. et al. (2024) “Dynamic Evaluation of Large Language Models by Meta Probing Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.14865v2.
Vancouver
1. Zhu K, Wang J, Zhao Q, Xu R, Xie X (2024) Dynamic Evaluation of Large Language Models by Meta Probing Agents. arXiv

BibTeX

@article{zhu2024dynamic,
  title = {Dynamic Evaluation of Large Language Models by Meta Probing Agents},
  author = {Zhu, Kaijie and Wang, Jindong and Zhao, Qinlin and Xu, Ruochen and Xie, Xing},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.14865v2},
  eprint = {2402.14865}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/