RACE: Retrieval-augmented Commit Message Generation
Ensheng ShiYanlin WangWei TaoLun DuHongyu ZhangShi HanDongmei ZhangHongbin Sun
Proposes a retrieval-augmented generation framework that uses an exemplar guider to control the influence of retrieved historical commits, significantly improving the quality and accuracy of automated commit messages across multiple programming languages.
Clear commit messages are vital for understanding software evolution, but developers frequently lack the time or motivation to document code changes manually, resulting in missing or low-quality descriptions. Existing automated approaches—such as pure rule-based, retrieval-based, or neural sequence-to-sequence models—often produce repetitive, vague, or inaccurate messages, while prior hybrid techniques fail to train the generation model to effectively leverage retrieved examples.
The article develops and evaluates a new retrieval-augmented framework, named RACE, designed to generate accurate, readable, and informative commit messages from source code changes. The model pairs a given code change with a semantically similar historical exemplar and uses an adaptive mechanism to guide message generation based on the degree of similarity.
The authors implemented a two-stage approach consisting of a retrieval module and a generation module. The retrieval module embeds fine-grained, token-level code changes into a high-dimensional space to find the most relevant commit from a historical repository. The generation module uses three encoders (for current code, retrieved code, and retrieved message), a learned weighting mechanism to dynamically scale how much influence the retrieved message has, and a decoder to generate the final text. Experiments were conducted across five programming languages (Java, C#, C++, Python, and JavaScript) using a curated public dataset containing roughly 900,000 code-message pairs, evaluating performance against eleven baseline systems using automated text similarity metrics and human assessments.
The experimental findings show that the proposed method consistently outperformed all eleven state-of-the-art baselines across all five programming languages, achieving average relative improvements of up to 46% on evaluation metrics compared to top-performing models. Furthermore, the retrieval-augmented framework generalized effectively across existing sequence-to-sequence architectures, boosting their baseline performance between 7% and 73% (with an average metric improvement of 11% to 43% across different models). An ablation study confirmed that removing the similarity-based weighting component caused measurable performance degradation across all languages. Finally, a blind human evaluation of developers rated the generated messages significantly higher in informativeness, conciseness, and expressiveness.
These results demonstrate that combining dense semantic retrieval with an adaptive guiding mechanism solves key limitations of purely generative or purely retrieval-based methods. For organizations, adopting this framework can streamline developer workflows, improve software documentation quality, and reduce the maintenance overhead and risk associated with poorly documented code repositories.
Engineering leaders should consider incorporating retrieval-augmented generation frameworks into their developer tooling and continuous integration pipelines to assist software teams. Because the framework requires access to an indexed code base, organizations should maintain high-quality historical commit repositories to serve as effective retrieval pools. Future development should explore expanding beyond the five studied programming languages, optimizing retrieval search across very large code bases, and developing strategies to process long code changes exceeding standard input limits without truncation.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Read this foundational account of retrieval-augmented generation first to understand the retrieve-and-generate design principle that RACE adapts to commit messages.
No sufficiently relevant recommendations were found.
