Learning or Self-aligning? Rethinking Instruction Fine-tuning

Mengjie RenBoxi CaoHongyu LinCao LiuXianpei HanKe ZengGuanglu WanXunliang CaiLe Sun

article2024ACL53 citations

Demonstrates that instruction fine-tuning succeeds primarily by aligning outputs with a language model's existing internal parameter knowledge rather than teaching new facts, warning that introducing inconsistent world knowledge during tuning degrades model performance.

Listen

Instruction fine-tuning is a vital step in transforming pre-trained large language models from basic text generators into helpful, task-oriented assistants. In industry and academia, fine-tuning is often treated as a standard supervised learning method to inject new domain-specific knowledge into models. However, this approach frequently underperforms or causes unintended performance degradations, creating substantial uncertainty around how to best construct fine-tuning datasets and adapt models for specialized domains.

The article evaluates whether instruction fine-tuning actually teaches models new factual knowledge or primarily aligns their output behavior with knowledge they already possess. To demonstrate this, the authors establish a knowledge intervention framework across four specialized domains—medicine, history, engineering, and jurisprudence—using four leading open-source models ranging from 7 billion to 70 billion parameters (LLaMA-2 and Mistral variants). The researchers probed each model's internal baseline knowledge using few-shot in-context learning and constructed experimental training sets that systematically manipulated the consistency between the training data and the model's pre-existing knowledge.

The findings reveal that attempting to force models to learn new factual knowledge during instruction fine-tuning produces poor results and often harms performance across related and unrelated tasks. Remarkably, models trained on factually incorrect responses that matched their internal knowledge outperformed models trained on factually correct responses that conflicted with their internal knowledge, showing average accuracy gains of roughly 5% to 10%. Furthermore, providing external factual context directly within the prompt during training prevented these negative effects, yielding performance improvements of approximately 4% to 9.5%. Most importantly, statistical analysis confirmed that the primary driver of fine-tuning success is maintaining consistency between the model's internal knowledge before and after tuning, rather than attempting to alter its internal factual state.

These insights demonstrate that instruction fine-tuning operates as a self-alignment mechanism rather than a knowledge injection phase. For strategic decision-makers, this shifts how resources and risk should be allocated in artificial intelligence initiatives. Attempting to update a model's factual worldview through simple instruction tuning introduces significant performance risks, increases hallucinations, and wastes engineering time. Instead, fine-tuning should be restricted to teaching style, behavioral norms, and response formatting, while core domain facts should be embedded during pre-training or retrieved dynamically at runtime.

Organizations developing or deploying large language models should immediately audit their fine-tuning workflows to ensure datasets align with base model internal capabilities rather than attempting to inject raw facts. When conflicting domain knowledge must be processed during training, teams should provide explicit contextual reference data within the prompt to decouple behavioral learning from factual learning. Decision-makers should also evaluate alignment pipelines that use the base model’s own internal representations to guide fine-tuning.

While the article's core statistical conclusions are supported by strong confidence levels across multiple architectures, certain boundaries apply. The empirical validation relies primarily on multiple-choice formats and focuses mostly on models around the 7-billion to 13-billion parameter range, with limited evaluations at the 70-billion scale and no free-form text generation benchmarks. Additional pilots and evaluations in open-ended generative settings are recommended before overhauling enterprise-scale generation systems.

No sufficiently relevant recommendations were found.

Cover for Learning or Self-aligning? Rethinking Instruction Fine-tuning

Abstract

Instruction Fine-tuning (IFT) is a crucial phase in building large language models (LLMs). Previous works mainly focus on the IFT’s role in the transfer of behavioral norms and the learning of additional world knowledge. However, the understanding of the underlying mechanisms of IFT remains significantly limited. In this paper, we design a knowledge intervention framework to decouple the potential underlying factors of IFT, thereby enabling individual analysis of different factors. Surprisingly, our experiments reveal that attempting to learn additional world knowledge through IFT often struggles to yield positive impacts and can even lead to markedly negative effects. Further, we discover that maintaining internal knowledge consistency before and after IFT is a critical factor for achieving successful IFT. Our findings reveal the underlying mechanisms of IFT and provide robust support for some very recent and potential future works. We release our experimental dataset and codes to facilitate future work¹.

Knowls

  1. Knowl 1 — Knowledge intervention separates learning from behavioral alignment

    model/method

    The knowledge intervention framework studies whether instruction fine-tuning (IFT) changes a model by teaching it additional world knowledge or by changing how it expresses knowledge already encoded in its parameters. For each base model and multiple-choice question, the researchers first probe the model’s existing answer using five in-context examples from the question’s domain. They then construct IFT data with different relationships between the target answer and the probed answer, fine-tune the model, and evaluate it on held-out questions from the same dataset (homogeneous evaluation), related MMLU subjects (in-domain), and other MMLU subjects (out-of-domain). Comparing these interventions is intended to isolate the consequences of supplying knowledge that agrees or conflicts with the model’s existing knowledge.

  2. Knowl 2 — Three IFT conditions manipulate agreement with the base model

    definition

    The paper defines three IFT data conditions by comparing each target answer with the base model’s answer from in-context probing. In the harmonious condition, the target answer is correct and agrees with the model’s answer, so the example supplies no knowledge the model lacks. In the incompatible condition, the target answer is correct but disagrees with the model, requiring the model to acquire knowledge that conflicts with its existing answer. In the self-aligning condition, the questions are the same as in the incompatible condition, but the targets are changed to the model’s probed answers; these targets are therefore incorrect relative to the dataset’s golden answers but agree with the model’s existing knowledge. The researchers keep the three groups the same size for a given model and domain.

  3. Knowl 3 — Evaluation and fine-tuning cover four domains and four base models

    experimental setup

    The experiments use multiple-choice data in medicine, history, engineering, and jurisprudence. The medical training corpus is a 20,000-example subset of MedMCQA filtered for explanations and a single correct answer; the other domain corpora come from Xiezhi Benchmark. Evaluation comprises a held-out split from the domain’s source dataset, in-domain MMLU subjects, and the remaining MMLU subjects as out-of-domain tests. Base models are LLaMA-2-7B, LLaMA-2-13B, LLaMA-2-70B, and Mistral-7B. Fine-tuning runs for three epochs with batch size 256, a learning rate of 2×10−52\times10^{-5} (or 1×10−51\times10^{-5} for Mistral-7B, because the higher rate did not converge), and loss computed only on output tokens. An equal proportion of general instruction data is added to stabilize training.

  4. Knowl 4 — Correct IFT answers perform better when they agree with existing model knowledge

    empirical result

    Across the four base models, four domains, and homogeneous, in-domain, and out-of-domain evaluations, fine-tuning on harmonious data generally outperforms fine-tuning on incompatible data. Both conditions contain correct target answers; the difference is whether those answers agree with the model’s probed knowledge. The paper reports average harmonious-condition gains over the incompatible condition of 11.27% on homogeneous tests, 14.58% on in-domain tests, and 14.57% on out-of-domain tests. Thus, in these experiments, correct targets that require learning conflicting knowledge yield worse results than correct targets already known to the base model.

  5. Knowl 5 — Incorrect but self-aligned targets outperform correct conflicting targets

    empirical result

    For each model and domain, the self-aligning and incompatible datasets use the same questions. The incompatible targets are the dataset’s correct answers, whereas the self-aligning targets reproduce the base model’s probed answers and are incorrect when that model answered incorrectly. Despite using these incorrect targets, fine-tuning on self-aligning data outperforms fine-tuning on incompatible data on average by 5.25% on homogeneous evaluations, 9.78% on in-domain evaluations, and 6.97% on out-of-domain evaluations. The comparison holds across the tested model sizes and architectures and indicates that introducing correct knowledge that conflicts with the model’s existing knowledge can be more damaging than preserving its existing answers.

  6. Knowl 6 — Providing conflicting knowledge in the context mitigates its fine-tuning cost

    model/method

    To separate knowledge use from knowledge acquisition, the researchers augment incompatible examples with an evidence passage generated by GPT-3.5 that contains the information needed to answer the question. The question, passage, and correct answer are used together for training, so the model can produce the expected answer by using the supplied context rather than having to internalize conflicting knowledge. At test time, the passage is not supplied. Compared with vanilla incompatible IFT, contextualized IFT improves average accuracy across homogeneous, in-domain, and out-of-domain tests by 8.16 percentage points for LLaMA-2-7B, 9.48 for LLaMA-2-13B, and 3.98 for Mistral-7B. For example, the reported overall scores for vanilla versus contextualized IFT are 32.48 versus 40.65 for LLaMA-2-7B, 36.37 versus 45.84 for LLaMA-2-13B, and 37.91 versus 41.89 for Mistral-7B. Mistral-7B’s homogeneous score is a reported exception, falling from 38.77 to 37.22, while its in-domain and out-of-domain scores increase.

  7. Knowl 7 — A mixture of aligned and conflicting examples often beats either extreme

    empirical result

    The researchers vary the consistency ratio by mixing self-aligning examples with incompatible examples; a ratio of 0 means all examples are incompatible, and a ratio of 1 means all are self-aligning. Across the tested domains and models, the all-incompatible endpoint performs poorly, while the all-self-aligning endpoint does not reliably give the best result. The highest performance is most often observed at an intermediate mixture of the two conditions, and the best ratio varies by model and domain. Consequently, agreement between IFT targets and the base model’s answers is influential but is not, by itself, a monotonic rule for selecting IFT data.

  8. Knowl 8 — Prediction agreement before and after IFT is positively associated with fine-tuned accuracy

    empirical result

    The paper measures internal-knowledge consistency as the average, over test questions, of the Pearson correlation between the rankings of candidate answers produced by the base model under five-shot in-context learning and by its fine-tuned model under zero-shot prompting. It then computes a Spearman partial correlation between this consistency score and fine-tuned accuracy while controlling for the base model’s accuracy on the same test set. The partial correlations are positive for every reported model and evaluation type. For the pooled models, the coefficients are 0.43 on homogeneous tests, 0.57 on in-domain tests, and 0.48 on out-of-domain tests, each with a reported p-value of 0.00. By model, the coefficients for homogeneous, in-domain, and out-of-domain tests are respectively 0.78, 0.81, and 0.82 for Mistral-7B; 0.27, 0.19, and 0.21 for LLaMA-2-7B; and 0.56, 0.78, and 0.87 for LLaMA-2-13B. The LLaMA-2-7B correlations are not statistically significant at the 0.05 level (reported p-values 0.14, 0.30, and 0.24); the other listed model-specific correlations have reported p-values of 0.00. The association therefore appears across evaluation types and remains after controlling for base accuracy, though its strength and significance vary by model.

  9. Knowl 9 — Extreme data conditions can separate fine-tuned output distributions from the base model

    empirical result

    As an additional measure of internal change, the researchers calculate KL divergence between each fine-tuned model’s zero-shot prediction distribution and its base model’s in-context prediction distribution. For the best-performing fine-tuned model, the self-aligning model, and the incompatible model, respectively, the reported divergences are 0.24, 0.34, and 0.37 for Mistral-7B; 0.16, 0.93, and 0.51 for LLaMA-2-7B; and 0.18, 0.68, and 0.34 for LLaMA-2-13B. “Best” denotes the model with the highest average performance across the three evaluation types. These results show that either extreme condition can be associated with a distribution that departs from the base model’s outputs; the pattern is not identical across model families. The paper interprets the stronger performance of intermediate mixtures as consistent with preserving the base model’s knowledge distribution while avoiding an overly sharp distribution induced by training only on self-aligning data.

  10. Knowl 10 — The study’s evidence is limited to multiple-choice tasks and a narrow range of model scales

    limitation

    The knowledge-probing and evaluation framework in this study uses multiple-choice questions, so the findings have not been established for free-form generation. Hardware constraints also mean that most experiments use models around 7B–13B parameters; the 70B LLaMA-2 model is included only in some experiments. The paper identifies extending the framework to free-form tasks and repeating the study with larger models and more domains as necessary to assess how broadly its conclusions apply.

Coverage note — Appendix prompt templates and exhaustive MMLU subject lists are omitted because they are implementation details rather than standalone contributions.

References

  1. 1.Christopher M Bishop. 2006. Pattern recognition and machine learning, volume 4. Springer.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, and Jared D. Kaplan. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901.
  3. 3.Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, et al. 2023. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. arXiv preprint arXiv:2312.09390.
  4. 4.Hao Chen, Yiming Zhang, Qi Zhang, Hantao Yang, Xiaomeng Hu, Xuetao Ma, Yifan Yanggong, and Junbo Zhao. 2023a. Maybe only 0.5% data is needed: A preliminary exploration of low training data instruction tuning. arXiv preprint arXiv:2305.09246.
  5. 5.Yirong Chen, Xiaofen Xing, Jingkai Lin, Huimin Zheng, and Zhenyu Wang. 2023b. Soulchat: Improving llms’ empathy, listening, and comfort abilities through fine-tuning with multi-turn empathy conversations. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 1170–1183, Singapore. Association for Computational Linguistics.
  6. 6.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research, 25(70):1–53.
  7. 7.Jiaxi Cui, Zongjian Li, Yang Yan, Bohua Chen, and Li Yuan. 2023. Chatlaw: Open-source legal large language model with integrated external knowledge bases. arXiv preprint arXiv:2306.16092.
  8. 8.Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022. A survey on in-context learning. arXiv preprint arXiv:2301.00234.
  9. 9.Lingyu Gao, Aditi Chaudhary, Krishna Srinivasan, Kazuma Hashimoto, Karthik Raman, and Michael Bendersky. 2023. Ambiguity-aware in-context learning with large language models. arXiv preprint arXiv:2309.07900.
  10. 10.Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. MIT Press.
  11. 11.Zhouhong Gu, Xiaoxuan Zhu, Haoning Ye, Lin Zhang, Jianchen Wang, Yixin Zhu, Sihang Jiang, Zhuozhi Xiong, Zihan Li, Weijie Wu, et al. 2024. Xiezhi: An ever-updating benchmark for holistic domain knowledge evaluation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 18099–18107.
  12. 12.Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song. 2023. The false promise of imitating proprietary llms. arXiv preprint arXiv:2305.15717.
  13. 13.Hongyi Guo, Yuanshun Yao, Wei Shen, Jiaheng Wei, Xiaoying Zhang, Zhaoran Wang, and Yang Liu. 2024. Human-instruction-free llm self-alignment with limited samples. arXiv preprint arXiv:2401.06785.
  14. 14.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In International Conference on Learning Representations.
  15. 15.Srinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, Todor Mihaylov, Daniel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu, Punit Singh Koura, et al. 2022. Opt-iml: Scaling language model instruction meta learning through the lens of generalization. arXiv preprint arXiv:2212.12017.
  16. 16.Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2023. Exploring the benefits of training expert language models over instruction tuning. In Proceedings of the 40th International Conference on Machine Learning, pages 14702–14729.
  17. 17.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  18. 18.Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, Steve Jiang, and You Zhang. 2023. Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge. Cureus, 15(6).
  19. 19.Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. 2024. The unlocking spell on base LLMs: Rethinking alignment via in-context learning. In The Twelfth International Conference on Learning Representations.
  20. 20.Alisa Liu, Xiaochuang Han, Yizhong Wang, Yulia Tsvetkov, Yejin Choi, and Noah A Smith. 2024. Tuning language models by proxy. arXiv preprint arXiv:2401.08565.
  21. 21.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11048–11064.
  22. 22.Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Conference on Health, Inference, and Learning, pages 248–260. PMLR.
  23. 23.Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023. Instruction tuning with gpt-4. arXiv preprint arXiv:2304.03277.
  24. 24.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’20, pages 3505–3506, New York, NY, USA. Association for Computing Machinery.
  25. 25.Ming Shen. 2024. Rethinking data selection for supervised fine-tuning. arXiv preprint arXiv:2402.06094.
  26. 26.C. Spearman. 1961. The Proof and Measurement of Association Between Two Things. Studies in Individual Differences: The Search for Intelligence. Appleton-Century-Crofts, East Norwalk, CT, US.
  27. 27.Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan. 2024. Principle-driven self-alignment of language models from scratch with minimal human supervision. Advances in Neural Information Processing Systems, 36.
  28. 28.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  29. 29.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  30. 30.Fanqi Wan, Xinting Huang, Leyang Cui, Xiaojun Quan, Wei Bi, and Shuming Shi. 2024. Mitigating hallucinations of large language models via knowledge consistent alignment. arXiv preprint arXiv:2401.10768.
  31. 31.Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, et al. 2023. Baichuan 2: Open large-scale language models. arXiv preprint arXiv:2309.10305.
  32. 32.Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. 2024. Self-rewarding language models. arXiv preprint arXiv:2401.10020.
  33. 33.Hanning Zhang, Shizhe Diao, Yong Lin, Yi R Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, and Tong Zhang. 2023. R-tuning: Teaching large language models to refuse unknown questions. arXiv preprint arXiv:2311.09677.
  34. 34.Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al. 2023. PyTorch fsdp: experiences on scaling fully sharded data parallel. arXiv preprint arXiv:2304.11277.
  35. 35.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36.
  36. 36.Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2024. Lima: Less is more for alignment. Advances in Neural Information Processing Systems, 36.

Citation

MLA
Ren, M., et al. “Learning or Self-aligning? Rethinking Instruction Fine-tuning”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 6090–105, https://doi.org/10.18653/v1/2024.acl-long.330.
APA
Ren, M., Cao, B., Lin, H., Liu, C., Han, X., Zeng, K., Guanglu, W., Cai, X., & Sun, L. (2024). Learning or Self-aligning? Rethinking Instruction Fine-tuning. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6090–6105. https://doi.org/10.18653/v1/2024.acl-long.330
Chicago
Ren, M., B. Cao, H. Lin, et al. 2024. “Learning or Self-aligning? Rethinking Instruction Fine-tuning”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6090–6105. https://doi.org/10.18653/v1/2024.acl-long.330.
Harvard
Ren, M. et al. (2024) “Learning or Self-aligning? Rethinking Instruction Fine-tuning”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6090–6105. Available at: https://doi.org/10.18653/v1/2024.acl-long.330.
Vancouver
1. Ren M, Cao B, Lin H, Liu C, Han X, Zeng K, Guanglu W, Cai X, Sun L (2024) Learning or Self-aligning? Rethinking Instruction Fine-tuning. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6090–6105

BibTeX

@inproceedings{ren-etal-2024-learning,
    title = "Learning or Self-aligning? Rethinking Instruction Fine-tuning",
    author = "Ren, Mengjie  and
      Cao, Boxi  and
      Lin, Hongyu  and
      Liu, Cao  and
      Han, Xianpei  and
      Zeng, Ke  and
      Guanglu, Wan  and
      Cai, Xunliang  and
      Sun, Le",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.330/",
    doi = "10.18653/v1/2024.acl-long.330",
    pages = "6090--6105"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/