RACE: Retrieval-augmented Commit Message Generation

Ensheng ShiYanlin WangWei TaoLun DuHongyu ZhangShi HanDongmei ZhangHongbin Sun

article2022EMNLP63 citations

Proposes a retrieval-augmented generation framework that uses an exemplar guider to control the influence of retrieved historical commits, significantly improving the quality and accuracy of automated commit messages across multiple programming languages.

Listen

Clear commit messages are vital for understanding software evolution, but developers frequently lack the time or motivation to document code changes manually, resulting in missing or low-quality descriptions. Existing automated approaches—such as pure rule-based, retrieval-based, or neural sequence-to-sequence models—often produce repetitive, vague, or inaccurate messages, while prior hybrid techniques fail to train the generation model to effectively leverage retrieved examples.

The article develops and evaluates a new retrieval-augmented framework, named RACE, designed to generate accurate, readable, and informative commit messages from source code changes. The model pairs a given code change with a semantically similar historical exemplar and uses an adaptive mechanism to guide message generation based on the degree of similarity.

The authors implemented a two-stage approach consisting of a retrieval module and a generation module. The retrieval module embeds fine-grained, token-level code changes into a high-dimensional space to find the most relevant commit from a historical repository. The generation module uses three encoders (for current code, retrieved code, and retrieved message), a learned weighting mechanism to dynamically scale how much influence the retrieved message has, and a decoder to generate the final text. Experiments were conducted across five programming languages (Java, C#, C++, Python, and JavaScript) using a curated public dataset containing roughly 900,000 code-message pairs, evaluating performance against eleven baseline systems using automated text similarity metrics and human assessments.

The experimental findings show that the proposed method consistently outperformed all eleven state-of-the-art baselines across all five programming languages, achieving average relative improvements of up to 46% on evaluation metrics compared to top-performing models. Furthermore, the retrieval-augmented framework generalized effectively across existing sequence-to-sequence architectures, boosting their baseline performance between 7% and 73% (with an average metric improvement of 11% to 43% across different models). An ablation study confirmed that removing the similarity-based weighting component caused measurable performance degradation across all languages. Finally, a blind human evaluation of developers rated the generated messages significantly higher in informativeness, conciseness, and expressiveness.

These results demonstrate that combining dense semantic retrieval with an adaptive guiding mechanism solves key limitations of purely generative or purely retrieval-based methods. For organizations, adopting this framework can streamline developer workflows, improve software documentation quality, and reduce the maintenance overhead and risk associated with poorly documented code repositories.

Engineering leaders should consider incorporating retrieval-augmented generation frameworks into their developer tooling and continuous integration pipelines to assist software teams. Because the framework requires access to an indexed code base, organizations should maintain high-quality historical commit repositories to serve as effective retrieval pools. Future development should explore expanding beyond the five studied programming languages, optimizing retrieval search across very large code bases, and developing strategies to process long code changes exceeding standard input limits without truncation.

Shi et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for RACE: Retrieval-augmented Commit Message Generation

Abstract

Commit messages are important for developers to understand changes in code repositories. However, writing commit messages is a time-consuming and tedious task for developers. To alleviate this burden, many approaches have been proposed to automatically generate commit messages. Among them, the state-of-the-art approaches are based on neural machine translation (NMT) models. However, these approaches only utilize the difference between the pre- and post-commit versions of the code (i.e., code diff) to generate commit messages, and ignore the rich information in the commit history. In this paper, we propose a novel approach named RACE (Retrieval-augmented Commit message gEneration) that retrieves similar code diffs and their corresponding commit messages from the commit history, and leverages the retrieved commit messages to guide the generation of commit messages. We evaluate RACE on a large-scale dataset collected from GitHub. Experimental results show that RACE outperforms the state-of-the-art approaches by a significant margin.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Work
  • 3 Proposed Approach
  • 3.1 Retrieval module
  • 3.2 Generation module
  • 4 Experimental Setup
  • 4.1 Dataset
  • 4.2 Data pre-processing
  • 4.3 Hyperparameters
  • 4.4 Evaluation metrics
  • 4.5 Baselines
  • 5 Experimental Results
  • 5.1 How does RACE perform compared with baseline approaches?
  • 5.2 What is the effectiveness of exemplar guider?
  • 5.3 What is the performance when we retrieve k relevant commits?
  • 5.4 Can our framework boost the performance of existing models?
  • 5.5 Human evaluation
  • 6 Conclusion
  • Limitations
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — RACE uses retrieved commits as guided generation exemplars

    model/method

    RACE is a retrieval-augmented model for generating a commit message from a code diff. It retrieves a training example containing a semantically similar code diff and its paired commit message, then uses the retrieved code diff to estimate how relevant that example is to the current diff. In generation, Transformer encoders represent the current diff, retrieved diff, and retrieved message; an exemplar guider uses the two diff representations to control how the current diff and retrieved message contribute to a Transformer decoder’s output. This lets the generator use relevant information from an exemplar without treating every retrieved message as fully reliable.

  2. Knowl 2 — The exemplar guider gates the retrieved message by diff similarity

    model/method

    For the current code diff, let HqH_q be the sequence of contextual token representations from its encoder; let HrH_r be the corresponding sequence for the retrieved code diff. RACE calculates a scalar guidance weight as λ=σ ⁣(Ws[mean⁡(Hq);mean⁡(Hr)])\lambda=\sigma\!\left(W_s[\operatorname{mean}(H_q);\operatorname{mean}(H_r)]\right), where mean is dimension-wise averaging, the semicolon denotes concatenation, WsW_s is a learnable projection, and σ\sigma is the sigmoid function. The resulting λ∈(0,1)\lambda\in(0,1) weights the retrieved commit-message representations, while 1−λ1-\lambda weights the current code-diff representations; the weighted sequences are concatenated as the decoder input. Thus, the guider learns from the current and retrieved diffs how much to rely on the retrieved message.

  3. Knowl 3 — Retrieval uses mean-pooled representations from a trained diff encoder

    model/method

    RACE trains a Transformer-based code-diff encoder as part of an encoder–decoder commit-message generator on approximately 0.9 million code-diff/message pairs, using cross-entropy training. For a code diff dd, if its encoder produces contextual token vectors h1,…,hLh_1,\ldots,h_L for its LL tokens, RACE forms the semantic vector v(d)=1L∑i=1Lhiv(d)=\frac{1}{L}\sum_{i=1}^{L}h_i. It compares these vectors by cosine similarity and retrieves the paired code diff and commit message with the highest similarity from the parallel training corpus. When retrieving for a training example, RACE uses the second-ranked result so that it does not retrieve the example itself; for other queries, it uses the first-ranked result.

  4. Knowl 4 — RACE outperforms baselines across five programming languages

    data/table

    RACE was evaluated against 11 baseline approaches on Java, C#, C++, Python, and JavaScript using B-Norm BLEU, METEOR, ROUGE-L, and CIDEr. The table reports RACE’s scores and the percentage improvements printed with its results in the paper. RACE scored higher than every baseline on all four metrics for all five languages; the paper reports all results as statistically significant at p<0.01p<0.01.

    Could not parse LaTeX table

    The gains are the paper’s reported percentages, not scores. They accompany RACE’s results in the comparison and indicate its advantage over the strongest baseline for each language and metric.

  5. Knowl 5 — Adding retrieval improves four existing generation models

    empirical result

    The authors adapted RACE’s retrieval-and-guidance framework to NMTGen, CommitBERT, CodeT5-small, and CodeT5-base, using each model’s encoder for retrieval representations and generation inputs and its decoder to generate messages. Averaged across the five evaluated programming languages, the augmented models improved over their original versions on BLEU, METEOR, ROUGE-L, and CIDEr, respectively, by: NMTGen, 43%, 49%, 33%, and 61%; CommitBERT, 11%, 9%, 11%, and 12%; CodeT5-small, 15%, 14%, 11%, and 26%; and CodeT5-base, 16%, 10%, 8%, and 32%. The paper also reports that BLEU improved for each model in every language, with language-specific improvements ranging from 7% to 73%.

  6. Knowl 6 — Removing the exemplar guider lowers generation scores

    empirical result

    The ablation compares full RACE with a variant that directly concatenates retrieved-result representations for the decoder and omits the exemplar guider. The values below are BLEU, METEOR, ROUGE-L, and CIDEr, respectively. The no-guider variant scores lower on every metric in every language, supporting the value of learning how strongly to use the retrieved exemplar.

    Could not parse LaTeX table
  7. Knowl 7 — Human raters prefer RACE on information, concision, and fluency

    empirical result

    Four experienced software developers rated generated messages from RACE and four baselines—CommitBERT, NNGen, NMTGen, and CoRec—on 50 test-set code diffs. The resulting 250 code-diff/message pairs were each scored by all four raters from 0 to 4 on informativeness, concision (freedom from extraneous content), and expressiveness (grammaticality and fluency); reported values are means with standard deviations. RACE had the highest mean on all three dimensions. Rater agreement was Krippendorff’s α=0.90\alpha=0.90, with pairwise Kendall’s τ\tau from 0.73 to 0.95. Wilcoxon signed-rank tests found RACE’s human-rating improvements over the baselines significant at p<0.05p<0.05.

    Could not parse LaTeX table
  8. Knowl 8 — Code diffs are represented as token-level change actions

    definition

    RACE represents each code diff as a sequence of spans of tokens marked by change actions, extracted with Python’s difflib. The action types are <keep> for unchanged spans, <insert> for added spans, and <delete> for removed spans. A replacement is represented with separate <replace_old> and <replace_new> markers for the old and new token spans. The experimental setup adds start/end special tokens for these action spans, allowing the encoder to distinguish changed code from unchanged context.

  9. Knowl 9 — Dataset and training configuration for the five-language evaluation

    experimental setup

    The evaluation uses MCMD commits from five languages, collected from popular GitHub repositories and filtered to remove redundant or noisy messages. The authors further exclude commits with multiple files or files that cannot be parsed. The final train/validation/test counts are: Java, 160,018/19,825/20,159; C#, 149,907/18,688/18,702; C++, 160,948/20,000/20,141; Python, 206,777/25,912/25,837; and JavaScript, 197,529/24,899/24,773. Maximum input lengths are 200 tokens for code diffs and 50 for commit messages. CodeT5-base weights initialize the code-diff encoders and decoder; the vocabulary is expanded from 32,100 to 32,109 tokens with nine change-action tokens. Training uses AdamW with learning rate 2×10−52\times10^{-5}, batch size 32, and at most 20 epochs. Results are averaged over random seeds 0, 1, and 2. Evaluation uses B-Norm BLEU, METEOR, ROUGE-L, and CIDEr.

  10. Knowl 10 — RACE depends on a codebase and has language and length limits

    limitation

    The reported evaluation covers only Java, C#, C++, Python, and JavaScript, so the authors state that the framework’s generality to other languages remains to be tested. Unlike a purely neural generator, RACE requires a codebase from which to retrieve examples. The authors report a training time of about 35 hours. Code diffs longer than 512 tokens are truncated, which can remove information; better handling of long diffs is left for future work.

Coverage note — The retrieved-neighbor-count sensitivity experiment (k = 1, 3, 5, 7, 9) is omitted as a separate knowl: the paper reports generally stable performance but does not provide exact per-k scores, making it less informative than the included results.

References

  1. 1.Apache. 2011. Apache lucene.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In IEEvaluation@ACL.
  3. 3.Mike Barnett, Christian Bird, João Brunet, and Shuvendu K. Lahiri. 2015. Helping developers help themselves: Automatic decomposition of code review changesets. In ICSE (1), pages 134–144. IEEE Computer Society.
  4. 4.Raymond P. L. Buse and Westley Weimer. 2010. Automatically documenting program changes. In ASE, pages 33–42. ACM.
  5. 5.Luis Fernando Cortes-Coy, Mario Linares Vásquez, Jairo Aponte, and Denys Poshyvanyk. 2014. On automatically generating commit messages via summarization of source code changes. In SCAM, pages 275–284. IEEE Computer Society.
  6. 6.Martin Dias, Alberto Bacchelli, Georgios Gousios, Damien Cassou, and Stéphane Ducasse. 2015. Untangling fine-grained code changes. In SANER, pages 341–350. IEEE Computer Society.
  7. 7.Jinhao Dong, Yiling Lou, Qihao Zhu, Zeyu Sun, Zhilin Li, Wenjie Zhang, and Dan Hao. 2022. Fira: Fine-grained graph-based code change representation for automated commit message generation.
  8. 8.Lun Du, Xiaozhou Shi, Yanlin Wang, Ensheng Shi, Shi Han, and Dongmei Zhang. 2021. Is a single model enough? mucos: A multi-model ensemble learning approach for semantic code search. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, pages 2994–2998.
  9. 9.Robert Dyer, Hoan Anh Nguyen, Hridesh Rajan, and Tien N. Nguyen. 2013. Boa: a language and infrastructure for analyzing ultra-large-scale software repositories. In ICSE, pages 422–431. IEEE Computer Society.
  10. 10.Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. Codebert: A pre-trained model for programming and natural languages. In EMNLP (Findings), volume EMNLP 2020 of Findings of ACL, pages 1536–1547. Association for Computational Linguistics.
  11. 11.Xiaodong Gu, Hongyu Zhang, and Sunghun Kim. 2018. Deep code search. In ICSE, pages 933–944. ACM.
  12. 12.Andrew F Hayes and Klaus Krippendorff. 2007. Answering the call for a standard reliability measure for coding data. Communication methods and measures, 1(1):77–89.
  13. 13.Yuan Huang, Nan Jia, Hao-Jie Zhou, Xiangping Chen, Zibin Zheng, and Mingdong Tang. 2020. Learning human-written commit messages to document code changes. J. Comput. Sci. Technol., 35(6):1258–1277.
  14. 14.Yuan Huang, Qiaoyang Zheng, Xiangping Chen, Yingfei Xiong, Zhiyong Liu, and Xiaonan Luo. 2017. Mining version control system for automatically generating commit comment. In ESEM, pages 414–423. IEEE Computer Society.
  15. 15.Siyuan Jiang, Ameer Armaly, and Collin McMillan. 2017. Automatically generating commit messages from diffs using neural machine translation. In ASE.
  16. 16.Tae Hwan Jung. 2021. Commitbert: Commit message generation using pre-trained programming language model. In Proceedings of the 1st Workshop on Natural Language Processing for Programming (NLP4Prog 2021), pages 26–33.
  17. 17.Maurice G Kendall. 1945. The treatment of ties in ranking problems. Biometrika, 33(3):239–251.
  18. 18.Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. In NeurIPS.
  19. 19.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out.
  20. 20.Qin Liu, Zihe Liu, Hongming Zhu, Hongfei Fan, Bowen Du, and Yu Qian. 2019. Generating commit messages from diffs using pointer-generator network. In MSR, pages 299–309. IEEE / ACM.
  21. 21.Shangqing Liu, Cuiyun Gao, Sen Chen, Lun Yiu Nie, and Yang Liu. 2020. ATOM: commit message generation based on abstract syntax tree and hybrid ranking. TSE, PP:1–1.
  22. 22.Zhongxin Liu, Xin Xia, Ahmed E. Hassan, David Lo, Zhenchang Xing, and Xinyu Wang. 2018. Neural-machine-translation-based commit message generation: how far are we? In ASE, pages 373–384. ACM.
  23. 23.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In ICLR.
  24. 24.Pablo Loyola, Edison Marrese-Taylor, and Yutaka Matsuo. 2017. A neural architecture for generating natural language descriptions from source code changes. In ACL (2), pages 287–292. Association for Computational Linguistics.
  25. 25.Laura Moreno, Jairo Aponte, Giriprasad Sridhara, Andrian Marcus, Lori L. Pollock, and K. Vijay-Shanker. 2013. Automatic generation of natural language summaries for java classes. In ICPC, pages 23–32. IEEE Computer Society.
  26. 26.Lun Yiu Nie, Cuiyun Gao, Zhicong Zhong, Wai Lam, Yang Liu, and Zenglin Xu. 2021. Coregen: Contextualized code representation learning for commit message generation. Neurocomputing, 459:97–107.
  27. 27.Sebastiano Panichella, Annibale Panichella, Moritz Beller, Andy Zaidman, and Harald C. Gall. 2016. The impact of test case summaries on bug fixing performance: an empirical investigation. In ICSE, pages 547–558. ACM.
  28. 28.Sheena Panthaplackel, Pengyu Nie, Milos Gligoric, Junyi Jessy Li, and Raymond J. Mooney. 2020. Learning to update natural language comments based on code changes. In ACL, pages 1853–1868. Association for Computational Linguistics.
  29. 29.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In ACL.
  30. 30.Jinfeng Shen, Xiaobing Sun, Bin Li, Hui Yang, and Jiajun Hu. 2016. On automatic summarization of what and why information in source code changes. In COMPSAC, pages 103–112. IEEE Computer Society.
  31. 31.Ensheng Shi, Wenchao Gub, Yanlin Wang, Lun Du, Hongyu Zhang, Shi Han, Dongmei Zhang, and Hongbin Sun. 2022a. Enhancing semantic code search with multimodal contrastive learning and soft data augmentation. arXiv preprint arXiv:2204.03293.
  32. 32.Ensheng Shi, Yanlin Wang, Lun Du, Junjie Chen, Shi Han, Hongyu Zhang, Dongmei Zhang, and Hongbin Sun. 2022b. On the evaluation of neural code summarization. In ICSE.
  33. 33.Ensheng Shi, Yanlin Wang, Lun Du, Hongyu Zhang, Shi Han, Dongmei Zhang, and Hongbin Sun. 2021a. Cast: Enhancing code summarization with hierarchical splitting and reconstruction of abstract syntax trees. In EMNLP.
  34. 34.Ensheng Shi, Yanlin Wang, Lun Du, Hongyu Zhang, Shi Han, Dongmei Zhang, and Hongbin Sun. 2021b. CAST: enhancing code summarization with hierarchical splitting and reconstruction of abstract syntax trees. In EMNLP (1), pages 4053–4062. Association for Computational Linguistics.
  35. 35.Wei Tao, Yanlin Wang, Ensheng Shi, Lun Du, Shi Han, Hongyu Zhang, Dongmei Zhang, and Wenqiang Zhang. 2021. On the evaluation of commit message generation models: An experimental study. In ICSME.
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NIPS, pages 5998–6008.
  37. 37.Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In CVPR.
  38. 38.Haoye Wang, Xin Xia, David Lo, Qiang He, Xinyu Wang, and John Grundy. 2021a. Context-aware retrieval-based deep commit message generation. ACM Trans. Softw. Eng. Methodol., 30(4):56:1–56:30.
  39. 39.Yanlin Wang, Lun Du, Ensheng Shi, Yuxuan Hu, Shi Han, and Dongmei Zhang. 2020. Cocogum: Contextual code summarization with multi-relational gnn on umls. Technical report, Microsoft, MSR-TR-2020-16. [Online].
  40. 40.Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021b. Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. In EMNLP (1), pages 8696–8708. Association for Computational Linguistics.
  41. 41.Bolin Wei, Yongmin Li, Ge Li, Xin Xia, and Zhi Jin. 2020. Retrieve and refine: exemplar-based neural comment generation. In 2020 35th IEEE/ACM International Conference on Automated Software Engineering (ASE), pages 349–360. IEEE.
  42. 42.Frank Wilcoxon, SK Katti, and Roberta A Wilcox. 1970. Critical values and probability levels for the wilcoxon rank sum test and the wilcoxon signed rank test. Selected tables in mathematical statistics, 1:171–259.
  43. 43.Shengbin Xu, Yuan Yao, Feng Xu, Tianxiao Gu, Hanghang Tong, and Jian Lu. 2019. Commit message generation for source code changes. In IJCAI, pages 3975–3981. ijcai.org.
  44. 44.HongChien Yu, Chenyan Xiong, and Jamie Callan. 2021. Improving query representations for dense retrieval with pseudo relevance feedback. In CIKM, pages 3592–3596. ACM.
  45. 45.Jian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun, and Xudong Liu. 2020. Retrieval-based neural source code summarization. In ICSE.

Citation

MLA
Shi, E., et al. “RACE: Retrieval-augmented Commit Message Generation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 5520–30, https://doi.org/10.18653/v1/2022.emnlp-main.372.
APA
Shi, E., Wang, Y., Tao, W., Du, L., Zhang, H., Han, S., Zhang, D., & Sun, H. (2022). RACE: Retrieval-augmented Commit Message Generation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 5520–5530. https://doi.org/10.18653/v1/2022.emnlp-main.372
Chicago
Shi, E., Y. Wang, W. Tao, et al. 2022. “RACE: Retrieval-augmented Commit Message Generation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 5520–30. https://doi.org/10.18653/v1/2022.emnlp-main.372.
Harvard
Shi, E. et al. (2022) “RACE: Retrieval-augmented Commit Message Generation”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 5520–5530. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.372.
Vancouver
1. Shi E, Wang Y, Tao W, Du L, Zhang H, Han S, Zhang D, Sun H (2022) RACE: Retrieval-augmented Commit Message Generation. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 5520–5530

BibTeX

@inproceedings{shi-etal-2022-race,
    title = "{RACE}: Retrieval-augmented Commit Message Generation",
    author = "Shi, Ensheng  and
      Wang, Yanlin  and
      Tao, Wei  and
      Du, Lun  and
      Zhang, Hongyu  and
      Han, Shi  and
      Zhang, Dongmei  and
      Sun, Hongbin",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.372/",
    doi = "10.18653/v1/2022.emnlp-main.372",
    pages = "5520--5530"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/