QuRating: Selecting High-Quality Data for Training Language Models

Alexander WettigAatmik GuptaSaumya MalikDanqi Chen

article2024ICML183 citations

Presents a scalable data selection method that trains compact rating models on LLM pairwise judgments across qualitative criteria like educational value, enabling language models to match the performance of baselines trained on 50% more data.

Listen

Training highly capable large language models requires vast amounts of text, yet standard pre-training pipelines still depend on blunt heuristics or coarse domain filters to curate data. These conventional methods often fail to capture subtle qualitative dimensions of text and can inadvertently discard valuable knowledge. As high-quality web text becomes scarcer and computing costs escalate, developing scalable techniques that can identify genuinely useful training material has become a critical operational challenge.

The article introduces and evaluates QuRating, a framework designed to capture human intuitions about text quality to guide language model pre-training. Specifically, the article analyzes four criteria: educational value, writing style, factual density, and required expertise. To build this system, an automated judge evaluated 250,000 text pairs to establish relative quality preferences. These comparisons were used to train a compact 1.3-billion-parameter rating model that scored a 260-billion-token corpus. Using probabilistic temperature sampling to balance quality and diversity, the authors selected 30-billion-token subsets and trained new language models from scratch to assess downstream capabilities across ten diverse benchmark tasks.

The core findings demonstrate that educational value is the most powerful selection criterion, delivering an average gain of 1.8 percentage points in in-context learning across all ten benchmarks compared to uniform data selection. Notably, a model trained on educationally curated data achieved performance comparable to a standard model trained on 50 percent more data and compute. The analysis also revealed that probabilistic sampling consistently outperforms rigid top-tier filtering, which excessively strips away sample diversity and hurts general performance. Furthermore, optimizing for writing style achieved the lowest perplexity but delivered negligible gains on downstream tasks, proving that standard perplexity metrics do not reliably predict reasoning and knowledge capabilities. Finally, sequencing data from low to high required expertise created an effective curriculum that enhanced performance using the same underlying data.

These insights show that data selection grounded in educational and explanatory qualities can substantially reduce training timelines, energy usage, and compute costs while producing stronger models. Decision-makers should prioritize educational value over stylistic polish when designing pre-training datasets. Technical teams should adopt soft probabilistic sampling rather than hard quality thresholds to avoid catastrophic loss of data diversity, and they should leverage quality scores to structure progressive training curricula.

Leaders should nevertheless weigh several limitations before large-scale adoption. The experimental findings were established using 1.3-billion-parameter models, meaning validation at larger model scales is recommended before making major infrastructure commitments. Additionally, because the rating model inherits biases from automated judgments, the selection process can exhibit subtle geographic, social, and linguistic skews. Organizations deploying this approach should combine automated selection with deliberate bias evaluations and manual domain balancing.

arXiv: 2402.09739
Cover for QuRating: Selecting High-Quality Data for Training Language Models

Abstract

Selecting high-quality pre-training data is important for creating capable language models, but existing methods rely on simple heuristics. We introduce QuRating, a method for selecting pre-training data that can capture human intuitions about data quality. In this paper, we investigate four qualities—writing style, required expertise, facts & trivia, and educational value—and find that LLMs are able to discern these qualities, especially when making pairwise judgments of texts. We train a QuRater model to learn scalar ratings from pairwise judgments, and use it to annotate a 260B training corpus with quality ratings for each of the four criteria. In our experiments, we select 30B tokens according to the different quality ratings and train 1.3B-parameter language models on the selected data. We find that it is important to balance quality and diversity. When we sample using quality ratings as logits over documents, our models obtain lower perplexity and stronger in-context learning performance than baselines. Our best model is based on educational value and performs similarly to a model trained with uniform sampling for 50% more steps. Beyond data selection, we use the quality ratings to construct a training curriculum which improves performance without changing the training dataset. We extensively analyze the quality ratings and discuss their characteristics, biases, and wider implications.2

Table of Contents

  • 1. Introduction
  • 2. Background
  • 3. Quantifying Qualitative Aspects of Text
  • 3.1. Overview of the Method
  • 3.2. Choice of Criteria and Prompts
  • 3.3. Why Use Pairwise Comparisons?
  • 3.4. Training the QuRater Model
  • 4. Selecting Data By Quality Rating
  • 5. Experiments
  • 5.1. Setup
  • 5.2. Data Selection Methods
  • 5.3. Results
  • 5.4. Curriculum Learning
  • 6. Analysis of Quality Ratings
  • 6.1. Distribution of Quality Ratings
  • 6.2. Data Inspection
  • 6.3. Documenting Social Bias
  • 7. Conclusion
  • Limitations
  • Impact Statement
  • Acknowledgements
  • References
  • A. Full Prompts
  • B. QuRater Model
  • B.1. Judgment Dataset
  • B.2. QuRater Training
  • C. Connection of Exponential Sampling to RLHF
  • D. Experimental Details
  • E. Further Analysis of Quality Ratings
  • F. Inspecting Raw Documents and Ratings

Knowls

  1. Knowl 1 — QuRating Pre-Training Data Selection via Temperature-Scaled Softmax Sampling

    model/method

    QuRating selects pre-training data subsets from a large corpus D\mathcal{D} by scoring each document di∈Dd_i \in \mathcal{D} with a scalar quality rating sis_i and sampling without replacement according to softmax probabilities:

    p(di)∝exp⁡(siτ)p(d_i) \propto \exp\left(\frac{s_i}{\tau}\right)

    where τ>0\tau > 0 is a temperature hyperparameter that balances data quality and data diversity. In the limit τ→0\tau \to 0, the procedure reduces to deterministic top-kk selection (maximizing quality score density at the expense of diversity). In the limit τ→∞\tau \to \infty, the procedure converges to uniform random sampling. Sampling without replacement is implemented at web scale via the Gumbel top-kk trick, which avoids duplicate documents and broadens token coverage.

  2. Knowl 2 — Bradley-Terry Formulation for Learning Scalar Quality Ratings from Pairwise LLM Preferences

    equation

    To quantify subjective text qualities without relying on hand-engineered proxy domains, pairwise comparative judgments between text snippets tAt_A and tBt_B are mapped to scalar ratings using the Bradley-Terry model. The probability that snippet tBt_B is preferred to tAt_A is modeled as:

    pB≻A=σ(sθ(tB)−sθ(tA))p_{B \succ A} = \sigma(s_\theta(t_B) - s_\theta(t_A))

    where σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} is the standard logistic sigmoid function, and sθ(t)∈Rs_\theta(t) \in \mathbb{R} is a scalar scoring function parameterized by a neural network with parameters θ\theta. Given a dataset of pairwise comparisons J={(tA,k,tB,k,pB≻A,k)}k=1N\mathcal{J} = \{(t_{A,k}, t_{B,k}, p_{B \succ A,k})\}_{k=1}^N with soft confidence targets pB≻A∈[0,1]p_{B \succ A} \in [0, 1], the model is trained by minimizing the binary cross-entropy loss:

    Lθ=E(tA,tB,pB≻A)∈J[−pB≻Alog⁡σ(sθ(tB)−sθ(tA))−(1−pB≻A)log⁡σ(sθ(tA)−sθ(tB))]\mathcal{L}_\theta = \mathbb{E}_{(t_A, t_B, p_{B \succ A}) \in \mathcal{J}} \left[ - p_{B \succ A} \log \sigma(s_\theta(t_B) - s_\theta(t_A)) - (1 - p_{B \succ A}) \log \sigma(s_\theta(t_A) - s_\theta(t_B)) \right]

  3. Knowl 3 — Theoretical Equivalence Between Exponential Data Resampling and RLHF

    theoretical result

    In Reinforcement Learning from Human Feedback (RLHF), a language model policy p(y)p(y) is optimized against an unconditional scalar reward s(y)s(y) subject to a Kullback-Leibler (KL) divergence constraint relative to a reference model pD(y)p_\mathcal{D}(y) pre-trained on a corpus D\mathcal{D}:

    p∗(y)=arg⁡max⁡pEy∼p[s(y)−τlog⁡p(y)pD(y)]p^*(y) = \arg\max_p \mathbb{E}_{y \sim p} \left[ s(y) - \tau \log \frac{p(y)}{p_\mathcal{D}(y)} \right]

    This optimization yields the closed-form optimal policy:

    p∗(y)=1ZpD(y)exp⁡(s(y)τ)p^*(y) = \frac{1}{Z} p_\mathcal{D}(y) \exp\left(\frac{s(y)}{\tau}\right)

    where Z=∑ypD(y)exp⁡(s(y)/τ)Z = \sum_y p_\mathcal{D}(y) \exp(s(y)/\tau) is the partition function. Training a language model via maximum likelihood on a dataset resampled with probability proportional to exp⁡(s(y)/τ)\exp(s(y)/\tau) optimizes the weighted cross-entropy objective:

    arg⁡max⁡p∑yexp⁡(s(y)/τ)p^D(y)log⁡p(y)\arg\max_p \sum_y \exp(s(y)/\tau) \hat{p}_\mathcal{D}(y) \log p(y)

    where p^D\hat{p}_\mathcal{D} is the empirical data distribution. Under the condition that D\mathcal{D} is sufficiently large such that pD(y)≈p^D(y)p_\mathcal{D}(y) \approx \hat{p}_\mathcal{D}(y), training on exponentially resampled data directly approximates pre-training on the whole corpus followed by RLHF steering, with the sampling temperature τ\tau acting precisely as the KL regularization coefficient.

  4. Knowl 4 — Comparative Impact of Abstract Quality Criteria on In-Context Learning and Instruction Following

    empirical result

    When selecting 30B pre-training tokens from a 260B corpus (SlimPajama) to train 1.3B-parameter transformer models across four criteria—Writing Style, Facts & Trivia, Educational Value, and Required Expertise:

    1. Educational Value (tau=2.0\\tau = 2.0) produces the strongest performance, achieving an average in-context learning (ICL) accuracy of 46.7% across 10 diverse tasks, outperforming uniform sampling (44.9%) by +1.8% and matching a uniform model trained on 45B tokens (+50% compute and data, 46.8%). In downstream instruction following (AlpacaFarm evaluation against GPT-4 after 10K ShareGPT SFT), it achieves a 57.3% win rate against the uniform baseline.
    2. Facts & Trivia (tau=2.0\\tau = 2.0) achieves 46.2% average ICL (+1.3% over uniform), excelling in world knowledge and reading comprehension tasks while showing minor trade-offs on commonsense reasoning.
    3. Required Expertise (tau=2.0\\tau = 2.0) achieves 46.0% average ICL (+1.1% over uniform), boosting reading comprehension and technical queries.
    4. Writing Style (tau=2.0\\tau = 2.0) yields the lowest held-out validation perplexity (8.90 vs. 8.96 for uniform) but minimal average ICL gain (45.0% vs. 44.9%).
  5. Knowl 5 — Failure of Top-k Selection and Superiority of Temperature Sampling for Corpus Coverage

    empirical result

    Selecting training data via top-kk filtering (corresponding to τ→0\tau \to 0) degrades model performance due to severe loss of diversity and support collapse over the broader language distribution. Across all four quality criteria (Writing Style, Facts & Trivia, Educational Value, Required Expertise), top-kk selection on 30B tokens yields validation perplexities between 10.53 and 11.54, substantially worse than the 8.96 validation perplexity of uniform sampling.

    Introducing temperature sampling with τ=2.0\tau = 2.0 resolves this distribution shift, improving validation perplexity over uniform sampling (8.90 to 8.93 vs. 8.96) while consistently yielding higher and more balanced downstream in-context learning accuracy across individual evaluation tasks compared to top-kk selection.

  6. Knowl 6 — Dissociation Between Held-Out Perplexity and In-Context Task Performance Under Distributional Selection

    empirical result

    Validation perplexity does not serve as a reliable surrogate for downstream model capabilities when evaluating non-uniform data selection methods. While selecting pre-training data by Writing Style (tau=2.0\\tau = 2.0) produces the best validation perplexity (8.90 vs. 8.96 for uniform baseline), it delivers virtually no gain on downstream tasks (+0.1% average ICL). Conversely, data selected by Educational Value (tau=2.0\\tau = 2.0) yields a higher perplexity (8.91) while driving substantial downstream improvements (+1.8% average ICL across 10 tasks).

    Furthermore, sequence-level negative log-likelihood scores from a standard pre-trained LLM (LLaMA-2-7B) do not correlate strongly with quality ratings: Spearman rank correlation is 0.50 for writing style, 0.18 for facts & trivia, 0.14 for educational value, and -0.07 for required expertise.

  7. Knowl 7 — Curriculum Learning via Quality-Rating Scheduling Without Dataset Alteration

    empirical result

    Quality ratings enable effective pre-training curricula without altering the set of training examples. When training a 1.3B language model on a fixed 30B-token dataset sampled uniformly from SlimPajama, altering the presentation order based on Required Expertise ratings improves performance relative to standard random shuffling:

    1. Increasing Expertise Curriculum (low to\\to high expertise): Improves average 10-task ICL accuracy from 44.9% to 45.5% (+0.6%) and reduces held-out validation perplexity from 8.96 to 8.92.
    2. Decreasing Expertise Curriculum (high to\\to low expertise): Improves average 10-task ICL accuracy from 44.9% to 45.4% (+0.5%), though validation perplexity remains unchanged at 8.96.
  8. Knowl 8 — Superiority of Pairwise LLM Comparisons Over Pointwise Scoring for Quality Judgments

    empirical result

    LLMs evaluate qualitative textual attributes with significantly higher fidelity and consistency when prompted with pairwise comparisons rather than absolute scalar scoring. In a benchmark comparing GPT-3.5-turbo assessments against human consensus rankings on subtle gradations of writing style across 45 document pairs:

    1. Pairwise comparisons achieve a Kendall rank correlation coefficient of τ=0.79±0.01\tau = 0.79 \pm 0.01.
    2. Pointwise scoring (rating individual documents on a 1–10 scale) achieves τ=0.61±0.06\tau = 0.61 \pm 0.06.

    Evaluating pairs in both forward (tA,tB)(t_A, t_B) and reversed (tB,tA)(t_B, t_A) orders and averaging the sampled output distributions suppresses positional bias inherent in LLM evaluation.

  9. Knowl 9 — Multi-Task QuRater Model Architecture and Training Protocol

    experimental setup

    The QuRater model is constructed by fine-tuning Sheared-LLaMA-1.3B (a structured pruned version of LLaMA-2-7B) with four independent linear regression heads appended to the final token representation, enabling simultaneous prediction of all four quality criteria in a single forward pass.

    1. Training Data: 245,657 pairwise comparisons generated by GPT-3.5-turbo from 500K unique SlimPajama text snippets (sequence lengths 256–512 tokens). Training is restricted to pairs exhibiting a confidence margin ∣2pB≻A−1∣≥0.50|2p_{B \succ A} - 1| \ge 0.50 to eliminate noisy and position-biased samples.
    2. Optimization: Batch size 512, Adam optimizer with β=(0.9,0.95)\beta = (0.9, 0.95), learning rate 5×10−55 \times 10^{-5}, cosine decay, trained for 2 epochs. The model achieves over 93% accuracy on held-out confident pairwise judgments across all four criteria.
  10. Knowl 10 — Social, Topical, and Demographic Skew Induced by Quality-Based Web Data Selection

    empirical result

    Applying QuRating to the AboutMe dataset (10M web pages with metadata annotations) reveals systematic shifts in topic, social role, and geographic representation:

    1. Amplified Attributes: Sampling by Required Expertise and Educational Value preferentially selects research, law, technology, and university domains (e.g., postdoctoral fellow retained at 19–22% vs. 10% expected), while Writing Style amplifies literary and arts content (e.g., laureates and celebrants retained at 16–17%).
    2. Suppressed Attributes: Commercial/shopping websites (online stores, apparel) and conventionally female social roles (e.g., mommy, manicurist, seamstress retained at 6–7%) are suppressed across criteria.
    3. Mitigation via Temperature: Softmax sampling with τ=2.0\tau = 2.0 substantially attenuates demographic and occupational skew relative to top-kk selection, where retention rates for amplified roles reach 66% and suppressed roles drop to 0%.

Coverage note — Omitted qualitative qualitative inspections of individual cluster texts (Tables 11-40) as they serve as raw qualitative examples rather than generalizable findings.

References

  1. 1.Abbas, A., Tirumala, K., Simig, D., Ganguli, S., and Morcos, A. S. SemDeDup: Data-efficient learning at webscale through semantic deduplication. arXiv preprint arXiv:2303.09540, 2023.
  2. 2.Abid, A., Farooqi, M., and Zou, J. Persistent anti-muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’21, pp. 298–306, New York, NY, USA, 2021. Association for Computing Machinery. ISBN 9781450384735. doi: 10.1145/3461702.3462624. URL https://doi.org/10.1145/3461702.3462624.
  3. 3.Agarwal, R., Vieillard, N., Zhou, Y., Stanczyk, P., Garea, S. R., Geist, M., and Bachem, O. On-policy distillation of language models: Learning from self-generated mistakes. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=3zKtaqxLhW.
  4. 4.Almeida, T. and Hidalgo, J. SMS Spam Collection. UCI Machine Learning Repository, 2012. DOI: https://doi.org/10.24432/C5CC84.
  5. 5.Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. PaLM 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  6. 6.Arora, S. and Goyal, A. A theory for emergence of complex skills in language models. arXiv preprint arXiv:2307.15936, 2023.
  7. 7.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
  8. 8.Barocas, S., Crawford, K., Shapiro, A., and Wallach, H. The problem with bias: Allocative versus representational harms in machine learning. In 9th Annual conference of the special interest group for computing, information and society, 2017.
  9. 9.Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, pp. 610–623, New York, NY, USA, 2021. Association for Computing Machinery. ISBN 9781450383097. doi: 10.1145/3442188.3445922. URL https://doi.org/10.1145/3442188.3445922.
  10. 10.Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pp. 7432–7439, 2020.
  11. 11.Blodgett, S. L., Barocas, S., Daume III, H., and Wallach, H. Language (technology) is power: A critical survey of “bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 5454–5476, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.aclmain.485. URL https://aclanthology.org/2020.acl-main.485.
  12. 12.Bradley, R. A. and Terry, M. E. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952.
  13. 13.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  14. 14.Caliskan, A., Bryson, J. J., and Narayanan, A. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334): 183–186, 2017. doi: 10.1126/science.aal4230. URL https://www.science.org/doi/abs/10.1126/science.aal4230.
  15. 15.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023.
  16. 16.Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 2924–2936, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1300. URL https://aclanthology.org/N19-1300.
  17. 17.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? Try ARC, the AI2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  18. 18.Dev, S., Monajatipoor, M., Ovalle, A., Subramonian, A., Phillips, J., and Chang, K.-W. Harms of gender exclusivity and challenges in non-binary representation in language technologies. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 1968–1994, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlpmain.150. URL https://aclanthology.org/2021.emnlp-main.150.
  19. 19.Dodge, J., Sap, M., Marasovic, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M., and Gardner, M. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 1286–1305, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.98. URL https://aclanthology.org/2021.emnlp-main.98.
  20. 20.Doudna, J. A. and Charpentier, E. Genome editing. the new frontier of genome engineering with crispr-cas9. Science (New York, N.Y.), 346(6213): 1258096, November 2014. ISSN 0036-8075. doi: 10.1126/science.1258096. URL https://doi.org/10.1126/science.1258096.
  21. 21.Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., Zoph, B., Fedus, L., Bosma, M. P., Zhou, Z., Wang, T., Wang, E., Webster, K., Pellat, M., Robinson, K., Meier-Hellstern, K., Duke, T., Dixon, L., Zhang, K., Le, Q., Wu, Y., Chen, Z., and Cui, C. GLaM: Efficient scaling of language models with mixture-of-experts. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 5547–5569. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/du22c.html.
  22. 22.Du, Z., Zeng, A., Dong, Y., and Tang, J. Understanding emergent abilities of language models from the loss perspective. arXiv preprint arXiv:2403.15796, 2024.
  23. 23.Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P., and Hashimoto, T. Alpacafarm: A simulation framework for methods that learn from human feedback. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=4hturzLcKX.
  24. 24.Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. The Pile: An 800GB dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  25. 25.Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for fewshot language model evaluation, September 2021. URL https://doi.org/10.5281/zenodo.5371628.
  26. 26.Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 3356–3369, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findingsemnlp.301. URL https://aclanthology.org/2020.findings-emnlp.301.
  27. 27.Gienapp, L., Frobe, M., Hagen, M., and Potthast, M. Sparse pairwise re-ranking with pre-trained transformers. In Proceedings of the 2022 ACM SIGIR International Conference on Theory of Information Retrieval, ICTIR ’22, pp. 72–80, New York, NY, USA, 2022. Association for Computing Machinery. ISBN 9781450394123. doi: 10.1145/3539813.3545140. URL https://doi.org/10.1145/3539813.3545140.
  28. 28.Go, D., Korbak, T., Kruszewski, G., Rozen, J., Ryu, N., and Dymetman, M. Aligning language models with preferences through f-divergence minimization. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 11546–11583. PMLR, 23–29 Jul 2023. URL https://proceedings.mlr.press/v202/go23a.html.
  29. 29.Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Giorno, A. D., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., Salim, A., Shah, S., Behl, H. S., Wang, X., Bubeck, S., Eldan, R., Kalai, A. T., Lee, Y. T., and Li, Y. Textbooks are all you need, 2023.
  30. 30.Gururangan, S., Li, M., Lewis, M., Shi, W., Althoff, T., Smith, N. A., and Zettlemoyer, L. Scaling expert language models with unsupervised domain discovery. arXiv preprint arXiv:2303.14177, 2023.
  31. 31.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021.
  32. 32.Hinton, G. E., Vinyals, O., and Dean, J. Distilling the knowledge in a neural network. ArXiv, abs/1503.02531, 2015. URL https://api.semanticscholar.org/CorpusID:7200347.
  33. 33.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J. W., and Sifre, L. Training compute-optimal large language models. In Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=iBBcRUlOAPR.
  34. 34.Jiang, T., Yuan, X., Chen, Y., Cheng, K., Wang, L., Chen, X., and Ma, J. Fuzzydedup: Secure fuzzy deduplication for cloud storage. IEEE Transactions on Dependable and Secure Computing, 20(3):2466–2483, 2023. doi: 10.1109/TDSC.2022.3185313.
  35. 35.Kim, C., Sabharwal, A., and Ermon, S. Exact sampling with integer linear programs and random perturbations. Proceedings of the AAAI Conference on Artificial Intelligence, 30(1), Mar. 2016. doi: 10.1609/aaai.v30i1.10421. URL https://ojs.aaai.org/index.php/AAAI/article/view/10421.
  36. 36.Kim, Y. and Rush, A. M. Sequence-level knowledge distillation. In Su, J., Duh, K., and Carreras, X. (eds.), Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 1317–1327, Austin, Texas, November 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1139. URL https://aclanthology.org/D16-1139.
  37. 37.Kingma, D. and Ba, J. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), San Diega, CA, USA, 2015.
  38. 38.Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems, volume 35, pp. 22199–22213. Curran Associates, Inc., 2022. URL https://proceedings.neurips.cc/paper_files/paper/2022/file/8bb0d291acd4acf06ef112099c16f326-Paper-Conference.pdf.
  39. 39.Kool, W., Van Hoof, H., and Welling, M. Stochastic beams and where to find them: The Gumbel-top-k trick for sampling sequences without replacement. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pp. 3499–3508. PMLR, 09–15 Jun 2019. URL https://proceedings.mlr.press/v97/kool19a.html.
  40. 40.Korbak, T., Elsahar, H., Kruszewski, G., and Dymetman, M. On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting. In Advances in Neural Information Processing Systems, 2022a. URL https://openreview.net/forum?id=XvI6h-s4un.
  41. 41.Korbak, T., Perez, E., and Buckley, C. RL with KL penalties is better viewed as Bayesian inference. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 1083–1091, Abu Dhabi, United Arab Emirates, December 2022b. Association for Computational Linguistics. doi: 10.18653/v1/2022.findingsemnlp.77. URL https://aclanthology.org/2022.findings-emnlp.77.
  42. 42.Korbak, T., Shi, K., Chen, A., Bhalerao, R. V., Buckley, C., Phang, J., Bowman, S. R., and Perez, E. Pretraining language models with human preferences. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 17506–17533. PMLR, 23–29 Jul 2023. URL https://proceedings.mlr.press/v202/korbak23a.html.
  43. 43.Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466, 2019.
  44. 44.Lacoste, A., Luccioni, A., Schmidt, V., and Dandres, T. Quantifying the carbon emissions of machine learning. arXiv preprint arXiv:1910.09700, 2019.
  45. 45.Laurencon, H., Saulnier, L., Wang, T., Akiki, C., del Moral, A. V., Scao, T. L., Werra, L. V., Mou, C., Ponferrada, E. G., Nguyen, H., Frohberg, J., Saɘ sko, M., Lhoest, Q., McMillan-Major, A., Dupont, G., Biderman, S., Rogers, A., allal, L. B., Toni, F. D., Pistilli, G., Nguyen, O., Nikpoor, S., Masoud, M., Colombo, P., de la Rosa, J., Villegas, P., Thrush, T., Longpre, S., Nagel, S., Weber, L., Munoz, M. R., Zhu, J., Strien, D. V., Alyafeai, Z., Almubarak, K., Chien, V. M., Gonzalez-Dios, I., Soroa, A., Lo, K., Dey, M., Suarez, P. O., Gokaslan, A., Bose, S., Adelani, D. I., Phan, L., Tran, H., Yu, I., Pai, S., Chim, J., Lepercq, V., Ilic, S., Mitchell, M., Luccioni, S., and Jernite, Y. The bigscience ROOTS corpus: A 1.6TB composite multilingual dataset. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022. URL https://openreview.net/forum?id=UoEw6KigkUn.
  46. 46.Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N. Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 8424–8445, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acllong.577. URL https://aclanthology.org/2022.acl-long.577.
  47. 47.Lesterhuis, M., Bouwer, R., Van Daal, T., Donche, V., and De Maeyer, S. Validity of comparative judgment scores: How assessors evaluate aspects of text quality when comparing argumentative texts. In Frontiers in Education, volume 7, pp. 216. Frontiers, 2022.
  48. 48.Li, Y., Bubeck, S., Eldan, R., Giorno, A. D., Gunasekar, S., and Lee, Y. T. Textbooks are all you need ii: phi-1.5 technical report, 2023.
  49. 49.Liu, J., Cui, L., Liu, H., Huang, D., Wang, Y., and Zhang, Y. Logiqa: A challenge dataset for machine reading comprehension with logical reasoning. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20, pp. 3622–3628, 2020.
  50. 50.Lucy, L., Gururangan, S., Soldaini, L., Strubell, E., Bamman, D., Klein, L., and Dodge, J. AboutMe: Using self-descriptions in webpages to document the effects of english pretraining data filters. arXiv preprint arXiv:2401.06408, 2024.
  51. 51.Manzini, T., Yao Chong, L., Black, A. W., and Tsvetkov, Y. Black is to criminal as Caucasian is to police: Detecting and removing multiclass bias in word embeddings. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 615–621, Minneapolis, Minnesota, 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1062. URL https://aclanthology.org/N19-1062.
  52. 52.Marion, M., Ust un, A., Pozzobon, L., Wang, A., Fadaee, M., and Hooker, S. When less is more: Investigating data pruning for pretraining LLMs at scale, 2023.
  53. 53.Metsis, V., Androutsopoulos, I., and Paliouras, G. Spam filtering with naive bayes-which naive bayes? In CEAS, volume 17, pp. 28–69. Mountain View, CA, 2006.
  54. 54.Muennighoff, N., Rush, A. M., Barak, B., Scao, T. L., Tazi, N., Piktus, A., Pyysalo, S., Wolf, T., and Raffel, C. Scaling data-constrained language models. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=j5BuTrEj35.
  55. 55.Nadeem, M., Bethke, A., and Reddy, S. StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 5356–5371, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.416. URL https://aclanthology.org/2021.acl-long.416.
  56. 56.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback, 2022.
  57. 57.Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.M., Rothchild, D., So, D., Texier, M., and Dean, J. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350, 2021.
  58. 58.Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Alobeidli, H., Cappelli, A., Pannier, B., Almazrouei, E., and Launay, J. The refinedweb dataset for falcon LLM: Outperforming curated corpora with web data only. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. URL https://openreview.net/forum?id=kM5eGcdCzq.
  59. 59.Peters, J. and Schaal, S. Reinforcement learning by reward-weighted regression for operational space control. In Proceedings of the 24th international conference on Machine learning, pp. 745–750, 2007.
  60. 60.Petroni, F., Rocktaschel, T., Riedel, S., Lewis, P., Bakhtin, A., Wu, Y., and Miller, A. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 2463–2473, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1250. URL https://aclanthology.org/D19-1250.
  61. 61.Pollitt, A. The method of adaptive comparative judgement. Assessment in Education: principles, policy & practice, 19(3):281–300, 2012.
  62. 62.Pollitt, A. and Crisp, V. Could comparative judgements of script quality replace traditional marking and improve the validity of exam questions. In BERA annual conference, UMIST Manchester, England, 2004.
  63. 63.Qin, Z., Jagerman, R., Hui, K., Zhuang, H., Wu, J., Shen, J., Liu, T., Liu, J., Metzler, D., Wang, X., and Bendersky, M. Large language models are effective text rankers with pairwise ranking prompting, 2023.
  64. 64.Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021.
  65. 65.Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=HPuSIXJaa9.
  66. 66.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  67. 67.Ramalho, M. A war of flags between guyana and venezuela. https://www.bellingcat.com/news/2023/12/13/a-war-of-flags-between-guyana-and-venezuela/, 2023. Accessed: 2024-01-04.
  68. 68.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  69. 69.Sanh, V., Debut, L., Chaumond, J., and Wolf, T. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. ArXiv, abs/1910.01108, 2019. URL https://api.semanticscholar.org/CorpusID:203626972.
  70. 70.Shazeer, N. M. GLU variants improve transformer. ArXiv, abs/2002.05202, 2020.
  71. 71.Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., and Dey, N. SlimPajama: A 627B token cleaned and deduplicated version of RedPajama. https://www.cerebras.net/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama, 2023. URL https://huggingface.co/datasets/cerebras/SlimPajama-627B.
  72. 72.Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., Hofmann, V., Jha, A. H., Kumar, S., Lucy, L., Lyu, X., Magnusson, I., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Ravichander, A., Richardson, K., Shen, Z., Strubell, E., Subramani, N., Tafjord, O., Walsh, E. P., Hajishirzi, H., Smith, N. A., Zettlemoyer, L., Beltagy, I., Groeneveld, D., Dodge, J., and Lo, K. Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research. arXiv preprint, 2023.
  73. 73.Strubell, E., Ganesh, A., and McCallum, A. Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 3645–3650, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1355. URL https://aclanthology.org/P19-1355.
  74. 74.Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024. ISSN 0925-2312. doi: https://doi.org/10.1016/j.neucom.2023.127063. URL https://www.sciencedirect.com/science/article/pii/S0925231223011864.
  75. 75.Sun, W., Yan, L., Ma, X., Wang, S., Ren, P., Chen, Z., Yin, D., and Ren, Z. Is ChatGPT good at search? investigating large language models as re-ranking agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 14918–14937, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlpmain.923. URL https://aclanthology.org/2023.emnlp-main.923.
  76. 76.Suresh, H. and Guttag, J. A framework for understanding sources of harm throughout the machine learning life cycle. In EAAMO ’21, New York, NY, USA, 2021. Association for Computing Machinery. ISBN 9781450385534. doi: 10.1145/3465416.3483305. URL https://doi.org/10.1145/3465416.3483305.
  77. 77.Tan, Y. C. and Celis, L. E. Assessing social and intersectional biases in contextualized word representations. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Red Hook, NY, USA, 2019. Curran Associates Inc.
  78. 78.Thurstone, L. L. A law of comparative judgment. Psychological Review, 34(4):273–286, 1927. doi: 10.1037/h0070288.
  79. 79.Tirumala, K., Simig, D., Aghajanyan, A., and Morcos, A. D4: Improving LLM pretraining via document de-duplication and diversification. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems, volume 36, pp. 53983–53995. Curran Associates, Inc., 2023. URL https://proceedings.neurips.cc/paper_files/paper/2023/file/a8f8cbd7f7a5fb2c837e578c75e5b615-Paper-Datasets_and_Benchmarks.pdf.
  80. 80.TogetherAI. RedPajama: An open source recipe to reproduce llama training dataset, 2023. URL https://github.com/togethercomputer/RedPajama-Data.
  81. 81.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a. URL https://arxiv.org/pdf/2302.13971.pdf.
  82. 82.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and finetuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  83. 83.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
  84. 84.Vieira, T. Gumbel-max trick and weighted reservoir sampling, 2014. URL http://timvieira.github.io/blog/post/2014/08/01/gumbel-max-trick-and-weighted-reservoir-sampling/.
  85. 85.Wang, P., Li, L., Chen, L., Zhu, D., Lin, B., Cao, Y., Liu, Q., Liu, T., and Sui, Z. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926, 2023.
  86. 86.Wei, J., Wang, X., Schuurmans, D., Bosma, M., brian ichter, Xia, F., Chi, E. H., Le, Q. V., and Zhou, D. Chain of thought prompting elicits reasoning in large language models. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=_VjQlMeSB_J.
  87. 87.Welbl, J., Liu, N. F., and Gardner, M. Crowdsourcing multiple choice science questions. In Proceedings of the 3rd Workshop on Noisy User-generated Text, pp. 94–106, 2017.
  88. 88.Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzman, F., Joulin, A., and Grave, E. CCNet: Extracting high quality monolingual datasets from web crawl data. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pp. 4003–4012, Marseille, France, May 2020. European Language Resources Association. ISBN 979-10-95546-34-4. URL https://aclanthology.org/2020.lrec-1.494.
  89. 89.Xia, M., Artetxe, M., Zhou, C., Lin, X. V., Pasunuru, R., Chen, D., Zettlemoyer, L., and Stoyanov, V. Training trajectories of language models across scales. In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 13711–13738, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acllong.767. URL https://aclanthology.org/2023.acl-long.767.
  90. 90.Xia, M., Gao, T., Zeng, Z., and Chen, D. Sheared LLaMA: Accelerating language model pre-training via structured pruning. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=09iOdaeOzp.
  91. 91.Xie, S. M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P., Le, Q. V., Ma, T., and Yu, A. W. DoReMi: Optimizing data mixtures speeds up language model pretraining. In Thirty-seventh Conference on Neural Information Processing Systems, 2023a. URL https://openreview.net/forum?id=lXuByUeHhd.
  92. 92.Xie, S. M., Santurkar, S., Ma, T., and Liang, P. Data selection for language models via importance resampling. Advances in Neural Information Processing Systems (NeurIPS), 2023b.
  93. 93.Yu, D., Kaur, S., Gupta, A., Brown-Cohen, J., Goyal, A., and Arora, S. Skill-mix: a flexible and expandable family of evaluations for AI models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Jf5gplvglq.
  94. 94.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 4791–4800, 2019.
  95. 95.Zeng, Z., Yu, J., Gao, T., Meng, Y., Goyal, T., and Chen, D. Evaluating large language models at evaluating instruction following. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=tr0KidwPLc.
  96. 96.Zhang, X., Zhao, J., and LeCun, Y. Character-level convolutional networks for text classification. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1, NIPS’15, pp. 649–657, Cambridge, MA, USA, 2015. MIT Press.

Citation

MLA
Wettig, A., et al. “QuRating: Selecting High-Quality Data for Training Language Models”. arXiv, 2024, http://arxiv.org/abs/2402.09739v3.
APA
Wettig, A., Gupta, A., Malik, S., & Chen, D. (2024). QuRating: Selecting High-Quality Data for Training Language Models. arXiv. http://arxiv.org/abs/2402.09739v3
Chicago
Wettig, A., A. Gupta, S. Malik, and D. Chen. 2024. “QuRating: Selecting High-Quality Data for Training Language Models”. arXiv. http://arxiv.org/abs/2402.09739v3.
Harvard
Wettig, A. et al. (2024) “QuRating: Selecting High-Quality Data for Training Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.09739v3.
Vancouver
1. Wettig A, Gupta A, Malik S, Chen D (2024) QuRating: Selecting High-Quality Data for Training Language Models. arXiv

BibTeX

@article{wettig2024qurating,
  title = {QuRating: Selecting High-Quality Data for Training Language Models},
  author = {Wettig, Alexander and Gupta, Aatmik and Malik, Saumya and Chen, Danqi},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.09739v3},
  eprint = {2402.09739}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/