Fault-Aware Neural Code Rankers

Jeevana Priya InalaChenglong WangMei YangAndrés CodasMark EncarnaciónShuvendu K. LahiriMadanlal MusuvathiJianfeng Gao

article2022NeurIPS59 citations

Presents CodeRanker, an execution-free neural ranking framework that predicts specific compile and runtime error types to select correct model-generated programs, substantially boosting pass@1 accuracy across standard benchmarks without requiring unit test execution.

Listen

Large language models show substantial promise in automated software development, but generating accurate code on the first attempt remains a major challenge. While generating multiple candidate solutions increases the likelihood of producing a correct program, identifying the best candidate traditionally requires executing the code against unit tests. This execution-based filtering creates significant operational and security hurdles, including the risk of running unsafe code, the requirement for pre-written test suites, and missing environment dependencies during active development.

The article evaluates CodeRanker, an execution-free neural ranking framework designed to predict the correctness of generated Python programs without running them. Rather than relying on simple binary classification (correct versus wrong), the system is trained to be fault-aware by learning specific failure modes, such as compiler exceptions, runtime error types, erroneous line locations, and output mismatches.

To evaluate this approach, researchers fine-tuned an encoder-based language model on execution error metadata gathered from candidate programs across standard coding benchmarks. The ranking system was tested across four code-generation models of varying sizes (Codex, GPT-J, and two GPT-Neo variants) on the APPS, HumanEval, and MBPP datasets.

The evaluation yielded several key findings. First, the ranking system substantially improved single-attempt code generation accuracy across all evaluated models; for example, Codex's top-one accuracy rose from 26% to 39.6% on the primary validation benchmark. Second, the rankers demonstrated strong zero-shot transferability across entirely different problem sets, boosting Codex's top-one accuracy by roughly 6 percentage points on both HumanEval and MBPP without benchmark-specific retraining. Third, fault-aware multi-class classifiers consistently outperformed standard binary classifiers, proving particularly effective at identifying execution errors over logic errors. Finally, combining training datasets generated by different models yielded performance improvements, showing that data from smaller, inexpensive models can enhance ranker quality.

These findings suggest that execution-free neural ranking offers a practical, secure path to improving code generation in developer tools, such as integrated development environments, without the infrastructure overhead or security risks of executing untrusted code. A smaller 125-million parameter model paired with the ranker even outperformed a baseline model fifty times its size, presenting significant cost-saving opportunities for model deployment.

Organizations developing automated coding tools should consider integrating execution-free neural ranking into their generation pipelines. Recommended next steps include testing early-stage ranking on incomplete code snippets to reduce generation latency and exploring richer structural code representations.

Decision-makers should note certain limitations: the ranking process is probabilistic and can still misclassify programs, and generating multiple candidate solutions increases total inference time and compute cost. However, the consistent cross-benchmark performance supports strong confidence in fault-aware ranking as an effective method for enhancing AI-assisted software development.

Cover for Fault-Aware Neural Code Rankers

Abstract

Large language models (LLMs) have demonstrated an impressive ability to generate code for various programming tasks. In many instances, LLMs can generate a correct program for a task when given numerous trials. Consequently, a recent trend is to do large scale sampling of programs using a model and then filtering/ranking the programs based on the program execution on a small number of known unit tests to select one candidate solution. However, these approaches assume that the unit tests are given and assume the ability to safely execute the generated programs (which can do arbitrary dangerous operations such as file manipulations). Both of the above assumptions are impractical in real-world software development. In this paper, we propose CodeRanker, a neural ranker that can predict the correctness of a sampled program without executing it. Our CodeRanker is fault-aware i.e., it is trained to predict different kinds of execution information such as predicting the exact compile/runtime error type (e.g., an IndexError or a TypeError). We show that CodeRanker can significantly increase the pass@1 accuracy of various code generation models (including Codex [11], GPT-Neo, GPT-J) on APPS [25], HumanEval [11] and MBPP [3] datasets.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Code generation
  • 2.2 Code ranking
  • 3 Fault-Aware Neural Code Ranker
  • 3.1 Code Ranker Dataset
  • 3.2 Code Ranker Tasks
  • 3.3 Code Ranker Models
  • 4 Evaluation
  • 4.1 Experiment Setup
  • 4.2 Main Results: CODERANKER improves code generation models
  • 4.3 Ablations
  • 5 Related works
  • 6 Conclusions and Future Directions
  • References
  • Checklist

Knowls

  1. Knowl 1 — CODERANKER reranks sampled programs without execution

    model/method

    CODERANKER selects better programs from a set of programs sampled from a code-generation model without executing those programs during inference. Let GG be a code-generation task, let FF be a generator with sampling distribution PF(S∣G)P_F(S\mid G), and let S1,…,SnS_1,\ldots,S_n be sampled programs. A ranker RR receives (G,Si)(G,S_i) and assigns the score

    si=PR(CORRECT∣G,Si),s_i=P_R(\mathrm{CORRECT}\mid G,S_i),

    where PRP_R is the ranker's predicted probability that program SiS_i satisfies the task's unit tests. The programs are ordered as So1,…,SonS_{o_1},\ldots,S_{o_n} so that so1≥⋯≥sons_{o_1}\geq\cdots\geq s_{o_n}. To produce a final answer, the system returns the first kk programs in this ordering. Thus, execution and unit tests are required to create training labels, but not to rank programs at inference time. The method is intended to improve the probability that at least one of the top-ranked programs is correct, while also allowing ranking by whether a program executes without a compile/runtime error.

  2. Knowl 2 — Fault-aware classification tasks expose why generated programs fail

    definition

    CODERANKER uses execution-derived labels at several levels of detail. A program is CORRECT if it satisfies the task's unit tests. A WRONG program is classified as having an intent error when it executes but produces an output that differs from the expected output, or an execution error when it fails with a compile/runtime error such as a syntax, type, indexing, or timeout failure.

    The binary task BB predicts {CORRECT, WRONG}. The ternary task TT predicts {CORRECT, intent error, execution error}. The intent-aware task II retains execution error as one aggregate class and splits intent errors into nine classes: NoneError, EmptyError, OutputTypeError, LengthError, IntSmallError, IntLargeError, StringSmallError, StringLargeError, and Misc. These classes respectively describe missing or empty outputs, output-type mismatches, length mismatches, small or large integer deviations, small or large string-length deviations, and other output mismatches.

    The execution-aware task EE retains intent error as one aggregate class and splits execution errors into ten classes: NameError, ValueError, EOFError, TypeError, IndexError, KeyError, TimeoutException, SyntaxError, Function not found, and Misc. The combined task E+LE+L predicts both the execution-error class and the erroneous source-code line. Its line label includes −1-1 when there is no execution error.

  3. Knowl 3 — Execution of sampled programs creates the fault-aware ranker dataset

    model/method

    For every APPS training task, each code-generation model samples n=100n=100 complete programs. The sampled programs are executed on the task's unit tests, and the resulting outputs or compiler/runtime diagnostics are converted into CORRECT, intent-error, execution-error, error-type, and error-line labels. Codex is used in a few-shot manner; GPT-J and GPT-Neo models are fine-tuned for at most two epochs, with the checkpoint having the lowest validation loss selected, to preserve diversity among sampled programs.

    The resulting data are strongly imbalanced toward incorrect programs. The following counts give the number of correct, intent-error, and execution-error programs in the ranker datasets generated by each model:

    Could not parse LaTeX table

    Across these datasets, the number of wrong programs is approximately five to forty times the number of correct programs, with the imbalance generally increasing for smaller generation models.

  4. Knowl 4 — CodeBERT provides classification and erroneous-line prediction heads

    model/method

    Each CODERANKER is implemented by fine-tuning pretrained CodeBERT on the concatenation of a natural-language task description and a generated program. The input uses CodeBERT's special sequence format with a [CLS] token, task tokens, a separator, code tokens, and an end token; the maximum encoded length is 512 tokens. The final hidden representation of [CLS] is C∈RHC\in\mathbb{R}^{H}, where HH is CodeBERT's hidden dimension.

    For a classification task with KK output classes, a trainable matrix W∈RK×HW\in\mathbb{R}^{K\times H} produces class probabilities

    p=softmax⁡(CWT),p=\operatorname{softmax}(CW^{\mathsf T}),

    where pp is a KK-dimensional probability vector. All CodeBERT and classification-head parameters are fine-tuned with cross-entropy loss.

    For the E+LE+L task, a second trainable vector S∈RHS\in\mathbb{R}^{H} predicts the erroneous line. If Ti∈RHT_i\in\mathbb{R}^{H} is the hidden representation of the iith newline token in the encoded program, the probability that newline ii marks the error is

    Pi=exp⁡(STTi)∑jexp⁡(STTj).P_i=\frac{\exp(S^{\mathsf T}T_i)}{\sum_j\exp(S^{\mathsf T}T_j)}.

    A newline token is added before the code to represent no erroneous line, and another is added after the encoded code to represent an error beyond the 512-token input limit.

  5. Knowl 5 — Evaluation uses 100-sample ranking on APPS, HumanEval, and MBPP

    experimental setup

    The experiments use APPS, HumanEval, and MBPP. APPS contains 5,000 training and 5,000 test tasks; 600 tasks are held out from the original APPS training problems for validation. The validation tasks are selected so that Codex produces at least one correct program among 100 samples. HumanEval contains 164 test tasks, and MBPP contains 974 tasks, including 474 training and 500 test problems. Code-generation and ranker models are trained only with APPS data, then evaluated on all three datasets; task descriptions for HumanEval and MBPP are programmatically transformed to the APPS-style format for GPT-J and GPT-Neo transfer.

    The evaluated generators are Codex, GPT-J 6B, GPT-Neo 1.3B, and GPT-Neo 125M. GPT-J and GPT-Neo are fine-tuned on APPS for two epochs with batch size 256 and learning rate 10−510^{-5}. Generation uses temperature 0.80.8 for Codex and 0.90.9 for GPT-J and GPT-Neo unless an ablation changes the temperature; each sample contains up to 512 new tokens and is truncated at the prompt's stop sequence. Rankers are fine-tuned for 30 epochs with batch size 512 and learning rate 10−410^{-4}, using class weights to compensate for label imbalance and selecting the checkpoint with the best validation ranked pass@1. Experiments use V100 GPUs with 32 GB of memory.

    For kk sampled programs, pass@k is the fraction of tasks with at least one program that passes all unit tests, whereas exec@k is the fraction with at least one program that executes without a compile/runtime error, even if its output is wrong. Ranked versions compute the same quantities after sorting the samples with CODERANKER. The reported metrics use an unbiased estimator from 100 samples.

  6. Knowl 6 — Fault-aware ranking substantially improves APPS validation performance

    empirical result

    On the 600-task APPS validation set, selecting programs with the best fault-aware CODERANKER improves pass@1, pass@5, and exec@1 for every evaluated generation model. The values are percentages; pass@100 is the upper bound supplied by the available 100 samples.

    Could not parse LaTeX table

    The absolute pass@1 improvements range from 5.1 percentage points for GPT-Neo 125M to 13.6 points for Codex. The ranker also makes execution-aware selection especially effective: for example, GPT-Neo 1.3B's exec@1 rises from 52.1% to 85.6%. A 125M generator paired with a 125M ranker therefore outperforms the unranked GPT-J generator on several metrics despite GPT-J having roughly fifty times more parameters.

  7. Knowl 7 — CODERANKER improves APPS test-set selection despite harder tasks

    empirical result

    On the 5,000-task APPS test set, the tasks are harder than the validation tasks, as shown by their lower pass@100 values. Nevertheless, fault-aware ranking improves pass@1, pass@5, and exec@1 for all four generators.

    Could not parse LaTeX table

    For Codex, ranked pass@1 increases from 3.8% to 4.5%, while ranked exec@1 increases from 59.6% to 73.4%. The gains in execution-based metrics are consistently larger than the gains in correctness-based metrics, indicating that the ranker is better at identifying programs with execution failures than programs that execute but implement the wrong behavior.

  8. Knowl 8 — APPS-trained rankers transfer to HumanEval and MBPP without retraining

    empirical result

    CODERANKER models trained from APPS-generated data improve nearly all metrics when applied zero-shot to HumanEval and MBPP. Values are percentages.

    Could not parse LaTeX table

    The only listed degradation is GPT-J's MBPP ranked pass@5, which falls from 28.9% to 28.2%. Codex pass@1 rises from 26.3% to 32.3% on HumanEval and from 36.4% to 41.8% on MBPP. The transfer results support the paper's claim that many generated-code failure patterns recur across programming datasets.

  9. Knowl 9 — Ranker granularity, sampling temperature, and mixed training data affect performance

    empirical result

    Ablations show that the best fault-aware classification task depends on the generator. On the APPS validation set, the best ranked pass@1 model for Codex and GPT-J uses the ternary task TT, the best model for GPT-Neo 1.3B uses execution-error-plus-line prediction E+LE+L, and the best model for GPT-Neo 125M uses the intent-aware task II. Binary ranking is substantially worse, especially for the smaller generators. For ranked exec@1 specifically, execution-aware tasks EE and E+LE+L perform best, while binary BB and intent-aware II perform worst.

    Lowering Codex's sampling temperature from 0.80.8 to 0.20.2 increases unranked pass@1 but lowers pass@100. CODERANKER still improves HumanEval pass@1, but slightly hurts MBPP pass@1 when the temperature is low:

    Could not parse LaTeX table

    When ranker training data are generated by a different model than the target generator, the same-model dataset is usually strongest and datasets from models with large size differences are weakest. A mixed-small dataset, formed by taking 25% of each model-specific dataset, is slightly worse than the same-model dataset except for GPT-J. A mixed-large dataset containing all four datasets gives the best overall performance, suggesting that additional data from smaller and less expensive generators can augment ranker training.

  10. Knowl 10 — CODERANKER is not sound and increases inference cost

    limitation

    CODERANKER is a learned classifier rather than a sound verifier: it can label a correct program as wrong and a wrong program as correct. Its use also increases inference time because the system must generate many programs, typically n≫kn\gg k, and rank them before presenting only the best kk candidates. The current method requires complete programs before ranking, so it cannot prune an incorrect partial program during generation.

    The paper identifies ranking partial programs, exploiting richer code-structure-aware ranker architectures, predicting complete error messages, and evaluating transfer to generic software-development tasks as future directions. Such generic-task transfer is constrained by the lack of a suitable large-scale dataset.

Coverage note — Detailed illustrative error-program examples, pretraining-source descriptions of the generator models, and supplementary appendix analyses were omitted because they support interpretation or reproducibility but are not additional load-bearing contributions beyond the methods, ablations, results, and limitations captured here.

References

  1. 1.Github copilot: Your ai pair programmer. https://copilot.github.com/, 2021.
  2. 2.Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. Unified pre-training for program understanding and generation. arXiv preprint arXiv:2103.06333, 2021.
  3. 3.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  4. 4.Antonio Valerio Miceli Barone and Rico Sennrich. A parallel corpus of python functions and documentation strings for automated code documentation and code generation. arXiv preprint arXiv:1707.02275, 2017.
  5. 5.Zeki Bilgin, Mehmet Akif Ersoy, Elif Ustundag Soykan, Emrah Tomur, Pinar Çomak, and Leyli Karaçay. Vulnerability prediction from source code using machine learning. IEEE Access, 8:150672–150684, 2020.
  6. 6.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. Gpt-neox-20b: An open-source autoregressive language model. arXiv preprint arXiv:2204.06745, 2022.
  7. 7.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  8. 8.Jose Cambronero, Hongyu Li, Seohyun Kim, Koushik Sen, and Satish Chandra. When deep learning met code search. In Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pages 964–974, 2019.
  9. 9.Saikat Chakraborty, Rahul Krishna, Yangruibo Ding, and Baishakhi Ray. Deep learning based vulnerability detection: Are we there yet. IEEE Transactions on Software Engineering, 2021.
  10. 10.Shubham Chandel, Colin B Clement, Guillermo Serrato, and Neel Sundaresan. Training and evaluating a jupyter notebook data science assistant. arXiv preprint arXiv:2201.12901, 2022.
  11. 11.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  12. 12.Xinyun Chen, Chang Liu, and Dawn Song. Execution-guided neural program synthesis. In International Conference on Learning Representations, 2018.
  13. 13.Xinyun Chen, Chang Liu, and Dawn Song. Tree-to-tree neural networks for program translation. Advances in neural information processing systems, 31, 2018.
  14. 14.Xinyun Chen, Dawn Song, and Yuandong Tian. Latent execution for neural program synthesis beyond domain-specific languages. Advances in Neural Information Processing Systems, 34, 2021.
  15. 15.Xiao Cheng, Haoyu Wang, Jiayi Hua, Guoai Xu, and Yulei Sui. Deepwukong: Statically detecting software vulnerabilities using deep graph neural network. ACM Transactions on Software Engineering and Methodology (TOSEM), 30(3):1–33, 2021.
  16. 16.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  17. 17.Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555, 2020.
  18. 18.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  19. 19.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  20. 20.Kevin Ellis, Maxwell Nye, Yewen Pu, Felix Sosa, Josh Tenenbaum, and Armando Solar-Lezama. Write, execute, assess: Program synthesis with a repl. Advances in Neural Information Processing Systems, 32, 2019.
  21. 21.Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al. Codebert: A pre-trained model for programming and natural languages. arXiv preprint arXiv:2002.08155, 2020.
  22. 22.Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Wen-tau Yih, Luke Zettlemoyer, and Mike Lewis. Incoder: A generative model for code infilling and synthesis. arXiv preprint arXiv:2204.05999, 2022.
  23. 23.Xiaodong Gu, Hongyu Zhang, and Sunghun Kim. Deep code search. In 2018 IEEE/ACM 40th International Conference on Software Engineering (ICSE), pages 933–944. IEEE, 2018.
  24. 24.Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al. Graphcodebert: Pre-training code representations with data flow. arXiv preprint arXiv:2009.08366, 2020.
  25. 25.Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, et al. Measuring coding challenge competence with apps. arXiv preprint arXiv:2105.09938, 2021.
  26. 26.Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. Codesearchnet challenge: Evaluating the state of semantic code search. arXiv preprint arXiv:1909.09436, 2019.
  27. 27.Sumith Kulal, Panupong Pasupat, Kartik Chandra, Mina Lee, Oded Padon, Alex Aiken, and Percy S Liang. Spoc: Search-based pseudocode to code. Advances in Neural Information Processing Systems, 32, 2019.
  28. 28.Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. Competition-level code generation with alphacode. arXiv preprint arXiv:2203.07814, 2022.
  29. 29.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  30. 30.Lili Mou, Ge Li, Lu Zhang, Tao Wang, and Zhi Jin. Convolutional neural networks over tree structures for programming language processing. In Thirtieth AAAI conference on artificial intelligence, 2016.
  31. 31.Rohan Mukherjee, Yeming Wen, Dipak Chaudhari, Thomas Reps, Swarat Chaudhuri, and Christopher Jermaine. Neural program generation modulo static analysis. Advances in Neural Information Processing Systems, 34, 2021.
  32. 32.Anh Tuan Nguyen, Tung Thanh Nguyen, and Tien N Nguyen. Divide-and-conquer approach for multi-phase statistical migration for source code (t). In 2015 30th IEEE/ACM International Conference on Automated Software Engineering (ASE), pages 585–596. IEEE, 2015.
  33. 33.Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. A conversational paradigm for program synthesis. arXiv preprint arXiv:2203.13474, 2022.
  34. 34.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114, 2021.
  35. 35.Tal Schuster, Ashwin Kalyan, Oleksandr Polozov, and Adam Tauman Kalai. Programming puzzles. arXiv preprint arXiv:2106.05784, 2021.
  36. 36.Jianhao Shen, Yichun Yin, Lin Li, Lifeng Shang, Xin Jiang, Ming Zhang, and Qun Liu. Generate & rank: A multi-task framework for math word problems. arXiv preprint arXiv:2109.03034, 2021.
  37. 37.Jeffrey Svajlenko, Judith F Islam, Iman Keivanloo, Chanchal K Roy, and Mohammad Mamun Mia. Towards a big data curated benchmark of inter-project code clones. In 2014 IEEE International Conference on Software Maintenance and Evolution, pages 476–480. IEEE, 2014.
  38. 38.Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. arXiv preprint arXiv:2109.00859, 2021.
  39. 39.Frank F Xu, Uri Alon, Graham Neubig, and Vincent J Hellendoorn. A systematic evaluation of large language models of code. arXiv preprint arXiv:2202.13169, 2022.
  40. 40.Yaqin Zhou, Shangqing Liu, Jingkai Siow, Xiaoning Du, and Yang Liu. Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks. Advances in neural information processing systems, 32, 2019.

Citation

MLA
Inala, J. P., et al. “Fault-Aware Neural Code Rankers”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 13419–32, https://proceedings.neurips.cc/paper_files/paper/2022/file/5762c579d09811b7639be2389b3d07be-Paper-Conference.pdf.
APA
Inala, J. P., Wang, C., Yang, M., Codas, A., Encarnación, M., Lahiri, S., Musuvathi, M., & Gao, J. (2022). Fault-Aware Neural Code Rankers. Advances in Neural Information Processing Systems, 35, 13419–13432. https://proceedings.neurips.cc/paper_files/paper/2022/file/5762c579d09811b7639be2389b3d07be-Paper-Conference.pdf
Chicago
Inala, J. P., C. Wang, M. Yang, et al. 2022. “Fault-Aware Neural Code Rankers”. Advances in Neural Information Processing Systems 35: 13419–32. https://proceedings.neurips.cc/paper_files/paper/2022/file/5762c579d09811b7639be2389b3d07be-Paper-Conference.pdf.
Harvard
Inala, J.P. et al. (2022) “Fault-Aware Neural Code Rankers”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 13419–13432. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/5762c579d09811b7639be2389b3d07be-Paper-Conference.pdf.
Vancouver
1. Inala JP, Wang C, Yang M, Codas A, Encarnación M, Lahiri S, Musuvathi M, Gao J (2022) Fault-Aware Neural Code Rankers. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 13419–13432

BibTeX

@inproceedings{inala2022fault,
  title = {Fault-Aware Neural Code Rankers},
  author = {Inala, Jeevana Priya and Wang, Chenglong and Yang, Mei and Codas, Andres and Encarnación, Mark and Lahiri, Shuvendu and Musuvathi, Madanlal and Gao, Jianfeng},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {13419-13432},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/5762c579d09811b7639be2389b3d07be-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission