The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Seungone KimJuyoung SukJi Yong ChoShayne LongpreChaeeun KimDongkeun YoonGuijin SonYejin Choi 0001Sheikh ShafayatJinheon Baek

article2025NAACL82 citationsBest Paper Award

Presents a generation benchmark that evaluates 103 language models across nine core capabilities and 77 tasks by replacing generic metrics with instance-specific grading rubrics for automated model-based assessment.

Listen

As language models become more capable across diverse domains, evaluating their real-world generation quality has become a major bottleneck. Existing benchmarks often rely on simplistic proxy tasks, narrow evaluation of instruction following, or vague criteria such as general helpfulness, which fail to capture the nuanced, context-dependent judgments typical of human evaluation.

The article introduces BIGGEN BENCH, a principled generation benchmark designed to provide fine-grained, instance-specific evaluation of language models using evaluator language models. Its main objective is to establish a comprehensive framework that evaluates 103 frontier language models across nine core capabilities—instruction following, grounding, planning, reasoning, refinement, safety, theory of mind, tool usage, and multilingualism—spanning 77 distinct tasks and 765 human-validated instances.

The researchers developed the benchmark using a human-in-the-loop, top-down approach where each test prompt is paired with a specific five-point rubric and reference answer. They evaluated 103 models—comprising pre-trained base models, post-trained chat models, and commercial proprietary models—using five automated evaluator models alongside a rigorous human validation study comprising 3,236 ratings across 27 qualified evaluators.

The evaluation yielded several key findings regarding model capabilities and evaluation methodologies. First, automated evaluator models show statistically significant alignment with human judges across all capabilities, with GPT-4-Turbo achieving an average Pearson correlation of 0.623 and a jury majority vote among five evaluator models achieving 0.627. Second, instance-specific scoring rubrics significantly outperformed coarse-grained and domain-level criteria in correlating with human judgment, while direct five-point scoring eliminated verbosity bias (correlation between length and score was negligible at 0.05). Third, pre-trained base model performance scales smoothly and predictably with parameter size (R² of 0.47), whereas post-trained chat models exhibit weaker variance explained by size alone (R² of 0.22), showing that post-training techniques heavily dictate downstream quality. Fourth, while larger base models narrow the performance gap with chat models in basic instruction following, wide disparities remain in complex capabilities such as reasoning, tool usage, and refinement, where proprietary models still lead significantly over open-source alternatives.

These findings indicate that organizations cannot rely solely on parameter scaling or generic post-training to achieve high performance in advanced cognitive tasks. Furthermore, the article demonstrates that high-quality automated evaluation does not require expensive closed APIs; open-source evaluator models such as Prometheus-2, when continually trained on benchmark feedback and paired with self-consistency decoding, can achieve human correlation (0.607) on par with leading proprietary evaluators. This provides organizations with a cost-effective, reproducible method to build internal evaluation pipelines and safely track model development.

For future implementation, developers should adopt instance-specific evaluation rubrics and ensemble or specialized evaluator juries, particularly for complex capabilities like theory of mind and tool usage where automated evaluators show lower human agreement. Organizations should also conduct periodic human spot-checks of automated feedback. The primary limitations include the inherent sampling variability of open-ended generation benchmarks and the exclusion of pre-trained models from multilingual evaluation due to translation artifacts.

Cover for The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Abstract

As language models (LMs) become capable of handling a wide range of tasks, their evaluation is becoming as challenging as their development. Most generation benchmarks currently assess LMs using abstract evaluation criteria—like helpfulness and harmlessness—which often lack the flexibility and granularity of human assessment. Additionally, these benchmarks tend to focus disproportionately on specific capabilities such as instruction following, leading to coverage bias. To overcome these limitations, we introduce the BIGGEN BENCH, a principled generation benchmark designed to thoroughly evaluate nine distinct capabilities of LMs across 77 diverse tasks. A key feature of the BIGGEN BENCH is its use of instance-specific evaluation criteria, closely mirroring the nuanced discernment of human evaluation. We apply this benchmark to assess 103 frontier LMs using five evaluator LMs. Our code, data, and evaluation results are all publicly available.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 BIGGEN BENCH: The BiG Generation Benchmark
  • 3.1 Evaluation protocol
  • 3.2 Construction process
  • 4 Main results and analyses
  • 5 Can we rely on language models to evaluate other language models?
  • 5.1 Can evaluator LMs effectively simulate human evaluation across all capabilities?
  • 5.2 Are fine-grained evaluation criteria crucial to obtain more accurate judgments?
  • 5.3 Analysis of verbosity bias
  • 6 Conclusion
  • Acknowledgements
  • Limitations, Potential Risks, and Licenses
  • References
  • A Capabilities, tasks, and evaluation criteria in the BIGGEN BENCH
  • B Human evaluation
  • B.1 Cross-validation
  • B.2 Acquiring human judgments
  • B.3 Inter-human Agreement Statistics
  • C Ranking on each capability
  • C.1 Performance comparison with selected combinations of models
  • D Comparative analysis of LMs
  • D.1 Model group and model parameter size
  • D.2 Open-source base LMs vs. Open-source chat LMs
  • D.3 Open-source chat LMs vs. Proprietary LMs
  • E How can we improve open-source evaluator LMs for accessible evaluations?
  • E.1 Self-consistency decoding
  • E.2 Continual feedback training
  • F Prompt template
  • F.1 Prometheus prompt template
  • F.2 FLASK rubrics
  • F.3 MT-Bench prompt template
  • F.4 URIAL prompt template
  • G Evaluation Procedures & Full evaluation results
  • G.1 List of Evaluator LMs
  • G.2 Evaluation results with GPT-4-1106 as a judge
  • G.3 Evaluation results with GPT-4-2024-04-09 as a judge
  • G.4 Evaluation results with Prometheus-2-8x7B as a judge
  • G.5 Evaluation results with Prometheus-2-8x7B-BGB as a judge
  • G.6 Evaluation results with Claude-3-Opus as a judge

Knowls

  1. Knowl 1 — BiGGen Bench Taxonomy and Instance Architecture

    definition

    The BiGGen Bench is a generation evaluation benchmark structured in a top-down hierarchy: capability →\rightarrow task →\rightarrow instance →\rightarrow instance-specific evaluation rubric. It evaluates language models across 9 core capabilities spanning 77 tasks and 765 instances:

    • Instruction Following (10 tasks, 100 instances): Evaluates the ability to process ambiguous instructions, handle false presuppositions, follow lexical and semantic constraints, generate factual explanations, produce executable actions, solve compositional instructions, align with personalized values, and generate instructional/educational content.
    • Grounding (10 tasks, 100 instances): Measures strict adherence and sensitivity to system prompts, user instructions, file formats (JSON, CSV, XML), temporal bounds, and counterfactual contexts, with system messages defined as highest priority.
    • Reasoning (10 tasks, 100 instances): Evaluates inductive, deductive, and abductive reasoning, first-order logic, high-school and Olympiad-level mathematics, theorem proving, table reasoning, legal reasoning, and hypothesis generation.
    • Planning (7 tasks, 70 instances): Tests goal-oriented sequence generation, constrained scheduling, multi-level plan composition, next-state environment prediction, and reward function coding.
    • Refinement (8 tasks, 80 instances): Tests the enhancement and correction of responses given external supervision (compiler logs, API tools, human critiques, peer reviews) as well as unassisted self-refinement and judging other models.
    • Multilingualism (7 tasks, 70 instances): Evaluates culturally sensitive translation, poem writing, humor explanation, multilingual reasoning, historical interpretation, and cultural conventions across 10 non-English languages (Arabic, Bengali, French, German, Indonesian, Kazakh, Korean, Russian, Spanish, and Thai).
    • Safety (8 tasks, 80 instances): Evaluates adherence to ethical norms, controversy explanation, refusal of malicious/biased requests, secret preservation, harm identification, and in-context concept unlearning.
    • Theory of Mind (10 tasks, 100 instances): Measures belief and mental state tracking, opponent emotion and intention inference, dialogue turn prediction, perspective-taking for historical/fictional agents, and audience-adapted persuasive speech.
    • Tool Usage (8 tasks, 80 instances): Measures understanding of tool documentation, creating tools from existing code, sequential multi-step tool calling, search API usage, and web environment navigation.

    Each benchmark instance contains four core fields: (1) an optional system message, (2) a user input prompt, (3) a score-5 reference answer, and (4) an instance-specific scoring rubric providing explicit qualitative criteria for every integer score from 1 to 5.

  2. Knowl 2 — Evaluation Protocol for Base and Post-Trained Models on BiGGen Bench

    model/method

    The evaluation pipeline evaluates both pre-trained (base) and post-trained (chat) language models using a direct assessment protocol:

    • Post-Trained (Chat) Models: Evaluated via zero-shot prompting with the instance system prompt and input.
    • Pre-Trained (Base) Models: Evaluated using the URIAL in-context learning framework, which applies cross-task 3-shot demonstrations without instruction fine-tuning (inst_1k_v4 template for safety tasks, and inst_1k_v4.help for non-safety tasks).
    • Multilingual Tasks: Evaluated solely on chat models because base models exhibit "accidental translation" (generating responses in English even when prompted in another language). Multilingual scores are excluded from aggregate averages when benchmarking base models.
    • Evaluator LM Direct Assessment: An evaluator language model receives the prompt instruction, the evaluated response, the gold reference answer, and the 5-point instance rubric via the Prometheus prompt template. The evaluator generates explicit verbal feedback followed by a single integer Likert score s∈{1,2,3,4,5}s \in \{1, 2, 3, 4, 5\}.
    • Inference Configuration: All response and evaluator language models are sampled with temperature 0.00.0, nucleus sampling threshold top_p=0.9\text{top\_p} = 0.9, and repetition penalty 1.031.03.
  3. Knowl 3 — Superiority of Instance-Specific Evaluation Rubrics Over Coarse and Domain-Specific Criteria

    empirical result

    Comparing evaluator language model performance across varying levels of evaluation rubric granularity demonstrates that instance-specific rubrics achieve the highest alignment with human evaluations:

    • Coarse-Grained Rubrics (e.g., MT-Bench format): A single generic prompt evaluating general attributes such as helpfulness, relevance, accuracy, depth, creativity, and level of detail.
    • Domain-Specific Rubrics (e.g., FLASK format): High-level category rubrics (such as logical robustness, factuality, comprehension, insightfulness, harmlessness, metacognition) mapped across all instances in a capability.
    • Instance-Specific Rubrics (BiGGen Bench format): Tailored rubrics specifying the exact required steps, factual milestones, constraints, and edge cases for that specific instance.

    When evaluated against human annotations:

    • Instance-specific evaluation criteria achieve consistently higher Pearson correlations with human ratings than both coarse-grained and domain-specific criteria for both proprietary judges (GPT-4-Turbo-2024-04-09) and open-source judges (Prometheus-2-BGB 8x7B).
    • Open-source Prometheus-2-BGB 8x7B using instance-specific rubrics outperforms GPT-4-Turbo-2024-04-09 using coarse-grained rubrics in human correlation.
    • Domain-specific rubrics show lower correlation with humans than coarse-grained rubrics, because static category-level criteria fail to account for the fine-grained variations and task constraints present across heterogeneous generation prompts.
  4. Knowl 4 — Evaluator Language Model Correlation with Human Judgments Across Capabilities

    data/table

    Pearson correlation coefficients between scores produced by evaluator language models and human ground-truth ratings across 3,236 responses sampled from 765 BiGGen Bench inputs (6 evaluated responses per input for multilingual tasks, 4 evaluated responses per input for all other capabilities):

    Evaluator LM Inst. Follow. Ground. Reason. Plan. Refine. Multi. Safety ToM Tool. Average
    Prometheus-2 8x7B 0.413 0.526 0.517 0.607 0.421 0.459 0.516 0.371 0.412 0.471
    + Self-Consistency (N=3N=3) 0.432 0.583 0.549 0.590 0.455 0.502 0.571 0.371 0.469 0.502
    + Self-Consistency (N=5N=5) 0.465 0.577 0.539 0.593 0.436 0.484 0.593 0.392 0.452 0.503
    Prometheus-2-BGB 8x7B 0.620 0.661 0.626 0.642 0.516 0.554 0.691 0.441 0.441 0.577
    + Self-Consistency (N=3N=3) 0.643 0.699 0.665 0.701 0.585 0.540 0.678 0.501 0.455 0.607
    + Self-Consistency (N=5N=5) 0.619 0.689 0.659 0.716 0.577 0.545 0.672 0.533 0.455 0.607
    Claude-3-Opus 0.624 0.694 0.588 0.634 0.561 0.554 0.634 0.463 0.446 0.578
    GPT-4-1106 0.641 0.683 0.643 0.678 0.578 0.583 0.653 0.420 0.496 0.597
    GPT-4-Turbo-2024-04-09 0.647 0.718 0.695 0.678 0.578 0.574 0.692 0.478 0.551 0.623
    Majority Voting 0.646 0.715 0.674 0.708 0.575 0.611 0.687 0.497 0.529 0.627

    Among individual models, GPT-4-Turbo-2024-04-09 achieves the highest correlation with human judgments (r=0.623r = 0.623). Taking a majority vote over all five evaluator LMs achieves the highest overall alignment (r=0.627r = 0.627). Across all evaluators, Theory of Mind (r=0.371–0.533r = 0.371\text{--}0.533) and Tool Usage (r=0.412–0.551r = 0.412\text{--}0.551) exhibit systematically lower correlations with human scoring than other capabilities.

  5. Knowl 5 — Log-Linear Performance Scaling of Base Language Models Across Capabilities

    empirical result

    Evaluating 28 base language models (spanning 1B to 72B parameters) using URIAL prompting with GPT-4-Turbo-2024-04-09 as the direct assessment judge reveals a smooth, log-linear relationship between parameter size and model generation performance:

    Score=Slope⋅log⁡10(Parameters)+Intercept\text{Score} = \text{Slope} \cdot \log_{10}(\text{Parameters}) + \text{Intercept}

    Across capabilities, base model linear regression and Pearson correlation statistics are:

    • Average: Slope =0.68= 0.68, Intercept =1.98= 1.98, R2=0.47R^2 = 0.47, r=0.68r = 0.68, p=1.64×10−5p = 1.64 \times 10^{-5}
    • Grounding: Slope =0.72= 0.72, Intercept =2.16= 2.16, R2=0.45R^2 = 0.45, r=0.67r = 0.67, p=2.77×10−5p = 2.77 \times 10^{-5}
    • Instruction Following: Slope =0.64= 0.64, Intercept =2.24= 2.24, R2=0.37R^2 = 0.37, r=0.60r = 0.60, p=2.45×10−4p = 2.45 \times 10^{-4}
    • Planning: Slope =0.76= 0.76, Intercept =1.88= 1.88, R2=0.43R^2 = 0.43, r=0.66r = 0.66, p=4.46×10−5p = 4.46 \times 10^{-5}
    • Reasoning: Slope =0.74= 0.74, Intercept =1.71= 1.71, R2=0.51R^2 = 0.51, r=0.72r = 0.72, p=4.08×10−6p = 4.08 \times 10^{-6}
    • Refinement: Slope =0.56= 0.56, Intercept =2.08= 2.08, R2=0.44R^2 = 0.44, r=0.66r = 0.66, p=3.87×10−5p = 3.87 \times 10^{-5}
    • Safety: Slope =0.57= 0.57, Intercept =2.37= 2.37, R2=0.38R^2 = 0.38, r=0.61r = 0.61, p=1.80×10−4p = 1.80 \times 10^{-4}
    • Theory of Mind: Slope =0.64= 0.64, Intercept =2.07= 2.07, R2=0.34R^2 = 0.34, r=0.58r = 0.58, p=4.58×10−4p = 4.58 \times 10^{-4}
    • Tool Usage: Slope =0.78= 0.78, Intercept =1.29= 1.29, R2=0.62R^2 = 0.62, r=0.79r = 0.79, p=9.36×10−8p = 9.36 \times 10^{-8}

    Model parameter scaling explains nearly half (R2=0.47R^2 = 0.47) of the performance improvement variance in base models, driving qualitative gains in complex cognitive capabilities such as Tool Usage and Reasoning.

    In contrast, for 61 post-trained (chat) models, scaling accounts for less variance (Average Slope =0.44= 0.44, Intercept =2.87= 2.87, R2=0.22R^2 = 0.22, r=0.47r = 0.47, p=2.33×10−4p = 2.33 \times 10^{-4}), demonstrating that downstream chat performance is heavily determined by the post-training protocol rather than parameter scale alone.

  6. Knowl 6 — Diminishing Performance Gap Between Base and Chat Models with Parameter Scaling

    empirical result

    To analyze how model scaling affects the performance difference between base LMs and post-trained (chat) LMs, a linear mixed-effects model was fitted across matched base and chat variants sharing identical model specifications:

    Score=β0+β1⋅Group+β2⋅Size+β3⋅(Group⋅Size)+(1∣Model name)\text{Score} = \beta_0 + \beta_1 \cdot \text{Group} + \beta_2 \cdot \text{Size} + \beta_3 \cdot (\text{Group} \cdot \text{Size}) + (1|\text{Model name})

    where Group\text{Group} is an indicator variable (00 for base LM, 11 for chat LM), Size\text{Size} is the standardized model parameter count, Model name\text{Model name} is a random effect accounting for model-family variation, and β3\beta_3 represents the interaction between model type and parameter size.

    The interaction coefficient β3\beta_3 is negative and statistically significant across all capabilities (p<0.001p < 0.001):

    • Average: β3=−0.08∗∗∗\beta_3 = -0.08^{***} (Intercept β0=−0.24∗∗∗\beta_0 = -0.24^{***}, β1=0.44∗∗∗\beta_1 = 0.44^{***}, β2=0.23∗∗∗\beta_2 = 0.23^{***})
    • Refinement: β3=−0.05∗∗∗\beta_3 = -0.05^{***}
    • Reasoning: β3=−0.07∗∗∗\beta_3 = -0.07^{***}
    • Grounding: β3=−0.07∗∗∗\beta_3 = -0.07^{***}
    • Planning: β3=−0.07∗∗∗\beta_3 = -0.07^{***}
    • Tool Usage: β3=−0.08∗∗∗\beta_3 = -0.08^{***}
    • Instruction Following: β3=−0.09∗∗∗\beta_3 = -0.09^{***}
    • Safety: β3=−0.09∗∗∗\beta_3 = -0.09^{***}
    • Theory of Mind: β3=−0.14∗∗∗\beta_3 = -0.14^{***}

    The negative interaction indicates that as base models grow larger, tuning-free prompting allows them to close the performance gap with chat models. However, the gap narrows the least in Refinement (−0.05-0.05), Reasoning (−0.07-0.07), and Grounding (−0.07-0.07), showing that larger base model scale alone is insufficient for self-correction and structured reasoning, where post-training remains essential.

  7. Knowl 7 — Disparity Gap Between Open-Source Chat and Proprietary Language Models Across Capabilities

    data/table

    Welch's t-test and Hedges's gg effect size analysis comparing open-source chat language models (N=43,592N=43,592 judgments) against proprietary language models (N=10,708N=10,708 judgments) across 9 capabilities evaluated by GPT-4-Turbo-2024-04-09 on BiGGen Bench:

    Open-source Chat LMs Proprietary LMs Effect Size
    Capability NN Mean Std. NN Mean Std. tt Hedges's gg
    Average 43592 3.25 1.30 10708 3.89 1.06 53.32∗∗∗^{***} 0.51
    Instruction Following 5700 3.57 1.19 1400 4.01 0.98 14.48∗∗∗^{***} 0.38
    Grounding 5700 3.50 1.34 1400 4.13 1.05 19.04∗∗∗^{***} 0.49
    Planning 3990 3.49 1.14 980 4.11 0.73 20.91∗∗∗^{***} 0.58
    Reasoning 5700 3.09 1.35 1400 3.93 1.10 24.41∗∗∗^{***} 0.65
    Refinement 4331 3.20 1.26 1062 3.76 1.08 14.74∗∗∗^{***} 0.46
    Safety 4503 3.59 1.36 1106 4.06 1.14 12.01∗∗∗^{***} 0.36
    Theory of Mind 5699 3.43 0.91 1400 3.99 0.58 25.57∗∗∗^{***} 0.59
    Tool Usage 3990 2.92 1.21 980 3.60 1.04 17.85∗∗∗^{***} 0.58
    Multilingual 3989 2.08 1.26 980 3.16 1.39 22.10∗∗∗^{***} 0.84

    Proprietary models significantly outperform open-source chat models across all 9 capabilities (p<0.001p < 0.001). The performance gap is narrowest in Safety (g=0.36g=0.36), Instruction Following (g=0.38g=0.38), and Refinement (g=0.46g=0.46), whereas pronounced disparities remain in Multilingual tasks (g=0.84g=0.84), Reasoning (g=0.65g=0.65), Theory of Mind (g=0.59g=0.59), Planning (g=0.58g=0.58), and Tool Usage (g=0.58g=0.58).

  8. Knowl 8 — Distillation of Proprietary Evaluation Feedback into Open-Source Prometheus-2-BGB

    model/method

    To build an accessible open-source evaluator that rivals proprietary evaluator models without incurring recurring API evaluation costs, Prometheus-2-BGB 8x7B is trained by continual feedback distillation:

    • Training Data: 50,490 evaluation judgments (verbal rationales and 1–5 integer scores) generated by GPT-4-1106 across 66 response language models on the 765 BiGGen Bench prompts.
    • Fine-Tuning Configuration: LoRA applied to Prometheus-2 8x7B targeting Q, K, V, O, and W projection matrices along with the language modeling head, using rank r=256r = 256, scaling factor α=512\alpha = 512, dropout 0.10.1, learning rate 10−510^{-5}, batch size 8, sequence length 4096, and 1 epoch in bfloat16.
    • Generalization to Unseen Models: Tested on 28,305 judgments across 37 unseen response language models, Prometheus-2-BGB increases Pearson correlation with GPT-4-1106 from 0.688 to 0.836 (and 0.865 with N=5N=5 self-consistency) and correlation with human judgments from 0.471 to 0.577 (and 0.607 with N=5N=5 self-consistency), matching Claude-3-Opus (0.578) and GPT-4-1106 (0.597).
    • Transfer Across External Benchmarks: Prometheus-2-BGB improves Pearson correlation with GPT-4-1106 on external benchmarks without catastrophic forgetting: Vicuna Bench (0.685→0.7770.685 \rightarrow 0.777), MT Bench (0.665→0.7730.665 \rightarrow 0.773), and FLASK (0.659→0.7640.659 \rightarrow 0.764), while retaining in-domain Feedback Bench evaluation correlation (0.898→0.8900.898 \rightarrow 0.890).
  9. Knowl 9 — Absence of Verbosity Bias in Direct Fine-Grained LM-as-a-Judge Evaluation

    empirical result

    Analysis of verbosity bias across 78,795 direct assessment judgments made by GPT-4-Turbo-2024-04-09 on 103 language models across BiGGen Bench reveals that assigned scores do not exhibit length preference:

    Length=12.52⋅Score+339.31\text{Length} = 12.52 \cdot \text{Score} + 339.31

    where Length\text{Length} is the token count of the evaluated response and Score∈{1,2,3,4,5}\text{Score} \in \{1, 2, 3, 4, 5\}.

    The regression yields a correlation coefficient r=0.050r = 0.050, a coefficient of determination R2=0.049R^2 = 0.049, and p=2.69×10−42p = 2.69 \times 10^{-42}. Despite statistical significance due to the massive sample size, the extremely low correlation (r=0.050r = 0.050) demonstrates that direct assessment coupled with explicit instance-specific scoring rubrics insulates the evaluation pipeline against verbosity bias.

  10. Knowl 10 — Multi-Stage Human Validation and Evaluator Qualification Pipeline for Generation Benchmarks

    experimental setup

    BiGGen Bench employs a four-stage dataset construction and human evaluation pipeline to ensure benchmark quality and ground-truth reliability:

    1. Expert Hand-Crafting: 18 researchers hand-crafted 385 initial instances (25 instances across 5 tasks per non-multilingual capability), and 10 native speakers created 100 multilingual instances (10 instances across 10 tasks in 10 languages). Review removed 23 low-quality or confusing tasks, leaving 77 finalized tasks.
    2. In-Context Diversity Expansion: GPT-4-0125 expanded the hand-crafted seed instances using few-shot in-context demonstrations. For each target instance, 5 candidates were generated, and the one with the lowest semantic similarity (measured by BertScore) was selected across 5 iterative rounds, producing 770 instances.
    3. Independent Cross-Validation: Two independent annotators validated each instance on (1) task fit, (2) difficulty, (3) reference answer quality, and (4) scoring rubric quality. Instances flagged as poor by both annotators were removed (5 discarded), yielding the final 765 benchmark instances.
    4. Human Evaluation & Qualification:
      • Recruitment: 102 Prolific crowd workers were recruited and assigned to capability-matched domains.
      • Qualification: Evaluators were tested on 4 probe instances per capability; qualification required Pearson r≥0.6r \ge 0.6 (p<0.05p < 0.05) against GPT-4-1106 pseudo-references. 1,000 simulations of central-tendency scores (score of 3 plus Gaussian noise) verified that 0% of unreflective annotators pass this filter. 27 workers qualified.
      • Main Study: 29 crowd workers collected 2,780 evaluations across 695 non-multilingual instances (4 responses per prompt), and 10 native annotators produced 420 evaluations across 70 multilingual instances (6 responses per prompt), yielding 3,236 verified human ratings. Krippendorff's alpha inter-annotator agreement ranged from 0.592 (Grounding) to 0.895 (Instruction Following).

Coverage note — Deliberately omitted individual model leaderboard rankings and granular pre-training data mixture case studies (e.g., CodeLlama vs. Llemma on math reasoning and OLMo-SFT vs. DPO progression), as these serve as specific illustrative instances of the comprehensive scaling, capability gap, and evaluation analyses presented in the knowls.

References

  1. 1.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. 2022. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691.
  2. 2.Anurag Ajay, Seungwook Han, Yilun Du, Shuang Li, Abhi Gupta, Tommi Jaakkola, Josh Tenenbaum, Leslie Kaelbling, Akash Srivastava, and Pulkit Agrawal. 2024. Compositional foundation models for hierarchical planning. Advances in Neural Information Processing Systems, 36.
  3. 3.Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng, Jian-Guang Lou, and Weizhu Chen. 2023. Learning from mistakes makes llm better reasoner. arXiv preprint arXiv:2310.20689.
  4. 4.Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q Jiang, Jia Deng, Stella Biderman, and Sean Welleck. 2023. Llemma: An open language model for mathematics. arXiv preprint arXiv:2310.10631.
  5. 5.Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, et al. 2024. Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues. arXiv preprint arXiv:2402.14762.
  6. 6.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609.
  7. 7.Douglas Bates, Martin Maechler, Ben Bolker, Steven Walker, Rune Haubo Bojesen Christensen, Henrik Singmann, Bin Dai, Gabor Grothendieck, Peter Green, and M Ben Bolker. 2015. Package ‘lme4’. convergence, 12(1):2.
  8. 8.Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf. 2023. Open llm leaderboard. https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard.
  9. 9.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Wen-tau Yih, and Yejin Choi. 2019. Abductive commonsense reasoning. arXiv preprint arXiv:1908.05739.
  10. 10.Faeze Brahman, Chandra Bhagavatula, Valentina Pyatkin, Jena D Hwang, Xiang Lorraine Li, Hirona J Arai, Soumya Sanyal, Keisuke Sakaguchi, Xiang Ren, and Yejin Choi. 2023. Plasma: Making small language models better procedural knowledge models for (counterfactual) planning. arXiv preprint arXiv:2305.19472.
  11. 11.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems.
  12. 12.Tianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen, and Denny Zhou. 2023. Large language models as tool makers. arXiv preprint arXiv:2305.17126.
  13. 13.Andrew Caines, Luca Benedetto, Shiva Taslimipoor, Christopher Davis, Yuan Gao, Oeistein Andersen, Zheng Yuan, Mark Elliott, Russell Moore, Christopher Bryant, et al. 2023. On the application of large language models for language teaching and assessment technology. arXiv preprint arXiv:2307.08393.
  14. 14.Charles F Cannell, Peter V Miller, and Lois Oksenberg. 1981. Research on interviewing techniques. Sociological methodology, 12:389–437.
  15. 15.Hyungjoo Chae, Yongho Song, Kai Tzu-iunn Ong, Taeyoon Kwon, Minjin Kim, Youngjae Yu, Dongha Lee, Dongyeop Kang, and Jinyoung Yeo. 2023. Dialogue chain-of-thought distillation for commonsenseaware conversational agents. arXiv preprint arXiv:2310.09343.
  16. 16.Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. Chateval: Towards better llm-based evaluators through multi-agent debate. arXiv preprint arXiv:2308.07201.
  17. 17.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2022a. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588.
  18. 18.Yanran Chen, Jonas Belouadi, and Steffen Eger. 2022b. Reproducibility issues for bert-based evaluation metrics. arXiv preprint arXiv:2204.00004.
  19. 19.Pengyu Cheng, Jiawen Xie, Ke Bai, Yong Dai, and Nan Du. 2023. Everyone deserves a reward: Learning customized human preferences. arXiv preprint arXiv:2309.03126.
  20. 20.Yew Ken Chia, Pengfei Hong, Lidong Bing, and Soujanya Poria. 2023. Instructeval: Towards holistic evaluation of instruction-tuned large language models. arXiv preprint arXiv:2306.04757.
  21. 21.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An opensource chatbot impressing gpt-4 with 90%* chatgpt quality.
  22. 22.Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024. Chatbot arena: An open platform for evaluating llms by human preference.
  23. 23.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  24. 24.Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. 2023. Ultrafeedback: Boosting language models with high-quality feedback. arXiv preprint arXiv:2310.01377.
  25. 25.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2021. Knowledge neurons in pretrained transformers. arXiv preprint arXiv:2104.08696.
  26. 26.Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B Hashimoto. 2024. Length-controlled alpacaeval: A simple way to debias automatic evaluators. arXiv preprint arXiv:2404.04475.
  27. 27.Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al. 2023. Towards measuring the representation of subjective global opinions in language models. arXiv preprint arXiv:2306.16388.
  28. 28.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. Eli5: Long form question answering. arXiv preprint arXiv:1907.09190.
  29. 29.Markus Freitag, David Grangier, and Isaac Caswell. 2020. Bleu might be guilty but references are not innocent. arXiv preprint arXiv:2004.06063.
  30. 30.Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. 2022. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858.
  31. 31.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021. A framework for few-shot language model evaluation.
  32. 32.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Pal: Program-aided language models. In International Conference on Machine Learning, pages 10764–10799. PMLR.
  33. 33.Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna Clinciu, Dipanjan Das, Kaustubh D Dhole, et al. 2021. The gem benchmark: Natural language generation, its evaluation and metrics. arXiv preprint arXiv:2102.01672.
  34. 34.Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran, Alex Wang, Alexandros Papangelis, Aman Madaan, Angelina McMillan-Major, Anna Shvets, Ashish Upadhyay, Bingsheng Yao, et al. 2022. Gemv2: Multilingual nlg benchmarking in a single line of code. arXiv preprint arXiv:2206.11249.
  35. 35.Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2023. Critic: Large language models can self-correct with tool-interactive critiquing. arXiv preprint arXiv:2305.11738.
  36. 36.Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. 2024. Olmo: Accelerating the science of language models. arXiv preprint arXiv:2402.00838.
  37. 37.Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song. 2023. The false promise of imitating proprietary llms. arXiv preprint arXiv:2305.15717.
  38. 38.Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, Diego Zambrano, et al. 2024. Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models. Advances in Neural Information Processing Systems, 36.
  39. 39.Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Luke Benson, Lucy Sun, Ekaterina Zubova, Yujie Qiao, Matthew Burtell, David Peng, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Shafiq Joty, Alexander R. Fabbri, Wojciech Kryscinski, Xi Victoria Lin, Caiming Xiong, and Dragomir Radev. 2022. Folio: Natural language reasoning with first-order logic.
  40. 40.Michael Hanna and Ondˇrej Bojar. 2021. A fine-grained analysis of bertscore. In Proceedings of the Sixth Conference on Machine Translation, pages 507–517.
  41. 41.Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. 2023. Reasoning with language model is planning with world model. arXiv preprint arXiv:2305.14992.
  42. 42.Shreya Havaldar, Sunny Rai, Bhumika Singhal, Langchen Liu Sharath Chandra Guntuku, and Lyle Ungar. 2023. Multilingual language models are not multicultural: A case study in emotion. arXiv preprint arXiv:2307.01370.
  43. 43.Zhiwei He, Tian Liang, Wenxiang Jiao, Zhuosheng Zhang, Yujiu Yang, Rui Wang, Zhaopeng Tu, Shuming Shi, and Xing Wang. 2024. Exploring humanlike translation strategy with large language models. Transactions of the Association for Computational Linguistics, 12:229–246.
  44. 44.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
  45. 45.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874.
  46. 46.Giwon Hong, Jeonghwan Kim, Junmo Kang, SungHyon Myaeng, and Joyce Jiyoung Whang. 2023. Discern and answer: Mitigating the impact of misinformation in retrieval-augmented models with discriminators. arXiv preprint arXiv:2305.01579.
  47. 47.Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2022. Large language models can self-improve. arXiv preprint arXiv:2210.11610.
  48. 48.Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2023a. Large language models cannot self-correct reasoning yet. arXiv preprint arXiv:2310.01798.
  49. 49.Xiaowei Huang, Wenjie Ruan, Wei Huang, Gaojie Jin, Yi Dong, Changshun Wu, Saddek Bensalem, Ronghui Mu, Yi Qi, Xingyu Zhao, et al. 2023b. A survey of safety and trustworthiness of large language models through the lens of verification and validation. arXiv preprint arXiv:2305.11391.
  50. 50.Joel Jang, Seungone Kim, Bill Yuchen Lin, Yizhong Wang, Jack Hessel, Luke Zettlemoyer, Hannaneh Hajishirzi, Yejin Choi, and Prithviraj Ammanabrolu. 2023a. Personalized soups: Personalized large language model alignment via post-hoc parameter merging. arXiv preprint arXiv:2310.11564.
  51. 51.Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2023b. Exploring the benefits of training expert language models over instruction tuning. arXiv preprint arXiv:2302.03202.
  52. 52.Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, and Minjoon Seo. 2021. Towards continual knowledge learning of language models. arXiv preprint arXiv:2110.03215.
  53. 53.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023a. Mistral 7b. arXiv preprint arXiv:2310.06825.
  54. 54.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088.
  55. 55.Dongfu Jiang, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, and Wenhu Chen. 2023b. Tigerscore: Towards building explainable metric for all text generation tasks. arXiv preprint arXiv:2310.00752.
  56. 56.Yuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, and Wei Wang. 2023c. Followbench: A multi-level fine-grained constraints following benchmark for large language models. arXiv preprint arXiv:2310.20410.
  57. 57.Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? MIT Press.
  58. 58.Yimin Jing, Renren Jin, Jiahao Hu, Huishi Qiu, Xiaohua Wang, Peng Wang, and Deyi Xiong. 2023. Followeval: A multi-dimensional benchmark for assessing the instruction-following capability of large language models. arXiv preprint arXiv:2311.09829.
  59. 59.Pei Ke, Bosi Wen, Zhuoer Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, et al. 2023. Critiquellm: Scaling llm-as-critic for effective and explainable evaluation of large language model generation. arXiv preprint arXiv:2311.18702.
  60. 60.Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras, Gunhee Kim, Yejin Choi, and Maarten Sap. 2023a. Fantom: A benchmark for stress-testing machine theory of mind in interactions. arXiv preprint arXiv:2310.15421.
  61. 61.Seungone Kim, Se June Joo, Yul Jang, Hyungjoo Chae, and Jinyoung Yeo. 2023b. Cotever: Chain of thought prompting annotation toolkit for explanation verification. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, pages 195–208.
  62. 62.Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al. 2023c. Prometheus: Inducing fine-grained evaluation capability in language models. arXiv preprint arXiv:2310.08491.
  63. 63.Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024. Prometheus 2: An open source language model specialized in evaluating other language models. arXiv preprint arXiv:2405.01535.
  64. 64.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526.
  65. 65.Miyoung Ko, Ingyu Seong, Hwaran Lee, Joonsuk Park, Minsuk Chang, and Minjoon Seo. 2022. Claimdiff: Comparing and contrasting claims on contentious issues. arXiv preprint arXiv:2205.12221.
  66. 66.Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. 2024. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks. arXiv preprint arXiv:2401.13649.
  67. 67.Seongyun Lee, Seungone Kim, Sue Hyun Park, Geewook Kim, and Minjoon Seo. 2024a. Prometheus-vision: Vision-language model as a judge for fine-grained evaluation. arXiv preprint arXiv:2401.06591.
  68. 68.Seongyun Lee, Sue Hyun Park, Seungone Kim, and Minjoon Seo. 2024b. Aligning to thousands of preferences via system message generalization.
  69. 69.Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. 2023a. Cmmlu: Measuring massive multitask language understanding in chinese. arXiv preprint arXiv:2306.09212.
  70. 70.Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. 2023b. Generative judge for evaluating alignment. arXiv preprint arXiv:2310.05470.
  71. 71.Tianjian Li and Kenton Murray. 2023. Why does zero-shot cross-lingual generation fail? an explanation and a solution. arXiv preprint arXiv:2305.17325.
  72. 72.Xinyu Li, Zachary C Lipton, and Liu Leqi. 2024. Personalized language modeling from personalized human feedback. arXiv preprint arXiv:2402.05133.
  73. 73.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023c. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval.
  74. 74.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. Let’s verify step by step. arXiv preprint arXiv:2305.20050.
  75. 75.Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. 2023. The unlocking spell on base llms: Rethinking alignment via incontext learning. arXiv preprint arXiv:2312.01552.
  76. 76.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  77. 77.Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. 2023a. Agentbench: Evaluating llms as agents. arXiv preprint arXiv:2308.03688.
  78. 78.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b. G-eval: Nlg evaluation using gpt-4 with better human alignment.
  79. 79.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023c. Gpteval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634.
  80. 80.Shayne Longpre, Sayash Kapoor, Kevin Klyman, Ashwin Ramaswami, Rishi Bommasani, Borhane Blili-Hamelin, Yangsibo Huang, Aviya Skowron, Zheng-Xin Yong, Suhas Kotha, et al. 2024. A safe harbor for ai evaluation and red teaming. arXiv preprint arXiv:2403.04893.
  81. 81.Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023a. Eureka: Human-level reward design via coding large language models. arXiv preprint arXiv:2310.12931.
  82. 82.Yingwei Ma, Yue Liu, Yue Yu, Yuanliang Zhang, Yu Jiang, Changjian Wang, and Shanshan Li. 2023b. At which training stage does code data help llms reasoning? arXiv preprint arXiv:2309.16298.
  83. 83.Michael McCloskey and Neal J Cohen. 1989. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of learning and motivation, volume 24, pages 109–165. Elsevier.
  84. 84.Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. Ambigqa: Answering ambiguous open-domain questions. arXiv preprint arXiv:2004.10645.
  85. 85.Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi. 2023. Can llms keep a secret? testing privacy implications of language models via contextual integrity theory. arXiv preprint arXiv:2310.17884.
  86. 86.Seungjun Moon, Yongho Song, Hyungjoo Chae, Dongjin Kang, Taeyoon Kwon, Kai Tzu-iunn Ong, Seung-won Hwang, and Jinyoung Yeo. 2023. Coffee: Boost your code llms by fixing bugs with feedback. arXiv preprint arXiv:2311.07215.
  87. 87.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted questionanswering with human feedback. arXiv preprint arXiv:2112.09332.
  88. 88.Aida Nematzadeh, Kaylee Burns, Erin Grant, Alison Gopnik, and Thomas L Griffiths. 2018. Evaluating theory of mind in question answering. arXiv preprint arXiv:1808.09352.
  89. 89.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  90. 90.Bhargavi Paranjape, Scott Lundberg, Sameer Singh, Hannaneh Hajishirzi, Luke Zettlemoyer, and Marco Tulio Ribeiro. 2023. Art: Automatic multi-step reasoning and tool-use for large language models. arXiv preprint arXiv:2303.09014.
  91. 91.Michael Quinn Patton. 2002. Qualitative research & evaluation methods. sage.
  92. 92.Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju. 2023. In-context unlearning: Language models as few shot unlearners. arXiv preprint arXiv:2310.07579.
  93. 93.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics.
  94. 94.Biqing Qi, Kaiyan Zhang, Haoxiang Li, Kai Tian, Sihang Zeng, Zhang-Ren Chen, and Bowen Zhou. 2023. Large language models are zero shot hypothesis proposers. arXiv preprint arXiv:2311.05965.
  95. 95.Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao, Sangwoo Cho, Xiaoyang Wang, Xuansheng Wu, Fei Liu, Pengfei Liu, and Dong Yu. 2024. Infobench: Evaluating instruction following ability in large language models. arXiv preprint arXiv:2401.03601.
  96. 96.Yiwei Qin, Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2022. T5score: Discriminative fine-tuning of generative evaluation metrics. arXiv preprint arXiv:2212.05726.
  97. 97.Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al. 2023. Toolllm: Facilitating large language models to master 16000+ real-world apis. arXiv preprint arXiv:2307.16789.
  98. 98.Yifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen, Edoardo M Ponti, and Shay B Cohen. 2023. Are large language models temporally grounded? arXiv preprint arXiv:2311.08398.
  99. 99.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36.
  100. 100.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426.
  101. 101.Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950.
  102. 102.Abulhair Saparov and He He. 2022. Language models are greedy reasoners: A systematic formal analysis of chain-of-thought. arXiv preprint arXiv:2210.01240.
  103. 103.William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike. 2022. Self-critiquing models for assisting human evaluators. arXiv preprint arXiv:2206.05802.
  104. 104.Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. 2024. Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems, 36.
  105. 105.Timo Schick, Jane Dwivedi-Yu, Zhengbao Jiang, Fabio Petroni, Patrick Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, and Sebastian Riedel. 2022. Peer: A collaborative language model. arXiv preprint arXiv:2208.11663.
  106. 106.Natalie Schluter. 2017. The limits of automatic summarisation according to rouge. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, pages 41–45. Association for Computational Linguistics.
  107. 107.Thomas Scialom, Tuhin Chakrabarty, and Smaranda Muresan. 2022. Fine-tuned language models are continual learners.
  108. 108.Melanie Sclar, Sachin Kumar, Peter West, Alane Suhr, Yejin Choi, and Yulia Tsvetkov. 2023. Minding language models’(lack of) theory of mind: A plug-andplay multi-character belief tracker. arXiv preprint arXiv:2306.00924.
  109. 109.Natalie Shapira, Guy Zwirn, and Yoav Goldberg. 2023. How well do large language models perform on faux pas tests? In Findings of the Association for Computational Linguistics: ACL 2023, pages 10438–10451.
  110. 110.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et al. 2022. Language models are multilingual chain-of-thought reasoners. arXiv preprint arXiv:2210.03057.
  111. 111.Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, et al. 2024. Aya dataset: An openaccess collection for multilingual instruction tuning. arXiv preprint arXiv:2402.06619.
  112. 112.Guijin Son, Sangwon Baek, Sangdae Nam, Ilgyun Jeong, and Seungone Kim. 2024a. Multi-task inference: Can large language models follow multiple instructions at once? arXiv preprint arXiv:2402.11597.
  113. 113.Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, and Stella Biderman. 2024b. Kmmlu: Measuring massive multitask language understanding in korean. arXiv preprint arXiv:2402.11548.
  114. 114.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615.
  115. 115.Haotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai, and Chao Zhang. 2024. Adaplanner: Adaptive planning from feedback with language models. Advances in Neural Information Processing Systems, 36.
  116. 116.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and finetuned chat models.
  117. 117.Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. 2024. Replacing judges with juries: Evaluating llm generations with a panel of diverse models. arXiv preprint arXiv:2404.18796.
  118. 118.Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. 2024. The instruction hierarchy: Training llms to prioritize privileged instructions. arXiv preprint arXiv:2404.13208.
  119. 119.Bin Wang, Zhengyuan Liu, Xin Huang, Fangkai Jiao, Yang Ding, Ai Ti Aw, and Nancy F Chen. 2023a. Seaeval for multilingual foundation models: From cross-lingual alignment to cultural reasoning. arXiv preprint arXiv:2309.04766.
  120. 120.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022a. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  121. 121.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022b. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560.
  122. 122.Zekun Moore Wang, Zhongyuan Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Man Zhang, et al. 2023b. Rolellm: Benchmarking, eliciting, and enhancing role-playing abilities of large language models. arXiv preprint arXiv:2310.00746.
  123. 123.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.
  124. 124.Sean Welleck, Jiacheng Liu, Ronan Le Bras, Hannaneh Hajishirzi, Yejin Choi, and Kyunghyun Cho. 2021. Naturalproofs: Mathematical theorem proving in natural language. arXiv preprint arXiv:2104.01112.
  125. 125.Sean Welleck, Jiacheng Liu, Ximing Lu, Hannaneh Hajishirzi, and Yejin Choi. 2022a. Naturalprover: Grounded mathematical proof generation with language models. Advances in Neural Information Processing Systems, 35:4913–4927.
  126. 126.Sean Welleck, Ximing Lu, Peter West, Faeze Brahman, Tianxiao Shen, Daniel Khashabi, and Yejin Choi. 2022b. Generating sequences by learning to selfcorrect. arXiv preprint arXiv:2211.00053.
  127. 127.Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, et al. 2023. A survey on large language models for recommendation. arXiv preprint arXiv:2305.19860.
  128. 128.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. 2024. Osworld: Benchmarking multimodal agents for openended tasks in real computer environments. arXiv preprint arXiv:2404.07972.
  129. 129.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023a. Wizardlm: Empowering large language models to follow complex instructions.
  130. 130.Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Yang Wang, and Lei Li. 2023b. Instructscore: Towards explainable text generation evaluation with automatic feedback. arXiv preprint arXiv:2305.14282.
  131. 131.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2020. mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934.
  132. 132.Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin. 2023a. Shadow alignment: The ease of subverting safely-aligned language models. arXiv preprint arXiv:2310.02949.
  133. 133.Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2023b. Alignment for honesty. arXiv preprint arXiv:2312.07000.
  134. 134.Zonglin Yang, Li Dong, Xinya Du, Hao Cheng, Erik Cambria, Xiaodong Liu, Jianfeng Gao, and Furu Wei. 2022. Language models as inductive reasoners. arXiv preprint arXiv:2212.10923.
  135. 135.Shunyu Yao, Howard Chen, Austin W Hanjie, Runzhe Yang, and Karthik Narasimhan. 2023. Collie: Systematic construction of constrained text generation tasks. arXiv preprint arXiv:2307.08689.
  136. 136.Junjie Ye, Guanyu Li, Songyang Gao, Caishuang Huang, Yilong Wu, Sixian Li, Xiaoran Fan, Shihan Dou, Qi Zhang, Tao Gui, et al. 2024. Tooleyes: Finegrained evaluation for tool learning capabilities of large language models in real-world scenarios. arXiv preprint arXiv:2401.00741.
  137. 137.Seonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, Seungone Kim, Yongrae Jo, James Thorne, Juho Kim, and Minjoon Seo. 2023. Flask: Fine-grained language model evaluation based on alignment skill sets. arXiv preprint arXiv:2307.10928.
  138. 138.Dongkeun Yoon, Joel Jang, Sungdong Kim, Seungone Kim, Sheikh Shafayat, and Minjoon Seo. 2024. Langbridge: Multilingual reasoning without multilingual supervision. arXiv preprint arXiv:2401.10695.
  139. 139.Xinyan Velocity Yu, Sewon Min, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022. Crepe: Open-domain question answering with false presuppositions.
  140. 140.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. BARTScore: Evaluating generated text as text generation. In Advances in Neural Information Processing Systems.
  141. 141.Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, et al. 2023. Instruction tuning for large language models: A survey. arXiv preprint arXiv:2308.10792.
  142. 142.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
  143. 143.Ruochen Zhao, Xingxuan Li, Shafiq Joty, Chengwei Qin, and Lidong Bing. 2023. Verify-and-edit: A knowledge-enhanced chain-of-thought framework. arXiv preprint arXiv:2305.03268.
  144. 144.Zirui Zhao, Wee Sun Lee, and David Hsu. 2024. Large language models as commonsense knowledge for large-scale task planning. Advances in Neural Information Processing Systems, 36.
  145. 145.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023a. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685.
  146. 146.Mingqian Zheng, Jiaxin Pei, and David Jurgens. 2023b. Is" a helpful assistant" the best role for large language models? a systematic evaluation of social roles in system prompts. arXiv preprint arXiv:2311.10054.
  147. 147.Shanshan Zhong, Zhongzhan Huang, Shanghua Gao, Wushao Wen, Liang Lin, Marinka Zitnik, and Pan Zhou. 2023. Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation. arXiv preprint arXiv:2312.02439.
  148. 148.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023a. Lima: Less is more for alignment.
  149. 149.Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023b. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911.
  150. 150.Pei Zhou, Aman Madaan, Srividya Pranavi Potharaju, Aditya Gupta, Kevin R McKee, Ari Holtzman, Jay Pujara, Xiang Ren, Swaroop Mishra, Aida Nematzadeh, et al. 2023c. How far are large language models from agents with theory-of-mind? arXiv preprint arXiv:2310.03051.
  151. 151.Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, et al. 2023d. Webarena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854.
  152. 152.Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, et al. 2023e. Sotopia: Interactive evaluation for social intelligence in language agents. arXiv preprint arXiv:2310.11667.
  153. 153.Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021. Tat-qa: A question answering benchmark on a hybrid of tabular and textual content in finance. arXiv preprint arXiv:2105.07624.
  154. 154.Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2023. Judgelm: Fine-tuned large language models are scalable judges. arXiv preprint arXiv:2310.17631.

Citation

MLA
Kim, S., et al. “The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models”. arXiv, 2024, http://arxiv.org/abs/2406.05761v2.
APA
Kim, S., Suk, J., Cho, J. Y., Longpre, S., Kim, C., Yoon, D., Son, G., Cho, Y., Shafayat, S., Baek, J., Park, S. H., Hwang, H., Jo, J., Cho, H., Shin, H., Lee, S., Oh, H., Lee, N., Ho, N., … Seo, M. (2024). The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models. arXiv. http://arxiv.org/abs/2406.05761v2
Chicago
Kim, S., J. Suk, J. Y. Cho, et al. 2024. “The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models”. arXiv. http://arxiv.org/abs/2406.05761v2.
Harvard
Kim, S. et al. (2024) “The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2406.05761v2.
Vancouver
1. Kim S, Suk J, Cho JY, et al (2024) The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models. arXiv

BibTeX

@article{kim2024the,
  title = {The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models},
  author = {Kim, Seungone and Suk, Juyoung and Cho, Ji Yong and Longpre, Shayne and Kim, Chaeeun and Yoon, Dongkeun and Son, Guijin and Cho, Yejin and Shafayat, Sheikh and Baek, Jinheon and Park, Sue Hyun and Hwang, Hyeonbin and Jo, Jinkyung and Cho, Hyowon and Shin, Haebin and Lee, Seongyun and Oh, Hanseok and Lee, Noah and Ho, Namgyu and Joo, Se June and Ko, Miyoung and Lee, Yoonjoo and Chae, Hyungjoo and Shin, Jamin and Jang, Joel and Ye, Seonghyeon and Lin, Bill Yuchen and Welleck, Sean and Neubig, Graham and Lee, Moontae and Lee, Kyungjae and Seo, Minjoon},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2406.05761v2},
  eprint = {2406.05761}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/