Fault-Aware Neural Code Rankers
Jeevana Priya InalaChenglong WangMei YangAndrés CodasMark EncarnaciónShuvendu K. LahiriMadanlal MusuvathiJianfeng Gao
Presents CodeRanker, an execution-free neural ranking framework that predicts specific compile and runtime error types to select correct model-generated programs, substantially boosting pass@1 accuracy across standard benchmarks without requiring unit test execution.
Large language models show substantial promise in automated software development, but generating accurate code on the first attempt remains a major challenge. While generating multiple candidate solutions increases the likelihood of producing a correct program, identifying the best candidate traditionally requires executing the code against unit tests. This execution-based filtering creates significant operational and security hurdles, including the risk of running unsafe code, the requirement for pre-written test suites, and missing environment dependencies during active development.
The article evaluates CodeRanker, an execution-free neural ranking framework designed to predict the correctness of generated Python programs without running them. Rather than relying on simple binary classification (correct versus wrong), the system is trained to be fault-aware by learning specific failure modes, such as compiler exceptions, runtime error types, erroneous line locations, and output mismatches.
To evaluate this approach, researchers fine-tuned an encoder-based language model on execution error metadata gathered from candidate programs across standard coding benchmarks. The ranking system was tested across four code-generation models of varying sizes (Codex, GPT-J, and two GPT-Neo variants) on the APPS, HumanEval, and MBPP datasets.
The evaluation yielded several key findings. First, the ranking system substantially improved single-attempt code generation accuracy across all evaluated models; for example, Codex's top-one accuracy rose from 26% to 39.6% on the primary validation benchmark. Second, the rankers demonstrated strong zero-shot transferability across entirely different problem sets, boosting Codex's top-one accuracy by roughly 6 percentage points on both HumanEval and MBPP without benchmark-specific retraining. Third, fault-aware multi-class classifiers consistently outperformed standard binary classifiers, proving particularly effective at identifying execution errors over logic errors. Finally, combining training datasets generated by different models yielded performance improvements, showing that data from smaller, inexpensive models can enhance ranker quality.
These findings suggest that execution-free neural ranking offers a practical, secure path to improving code generation in developer tools, such as integrated development environments, without the infrastructure overhead or security risks of executing untrusted code. A smaller 125-million parameter model paired with the ranker even outperformed a baseline model fifty times its size, presenting significant cost-saving opportunities for model deployment.
Organizations developing automated coding tools should consider integrating execution-free neural ranking into their generation pipelines. Recommended next steps include testing early-stage ranking on incomplete code snippets to reduce generation latency and exploring richer structural code representations.
Decision-makers should note certain limitations: the ranking process is probabilistic and can still misclassify programs, and generating multiple candidate solutions increases total inference time and compute cost. However, the consistent cross-benchmark performance supports strong confidence in fault-aware ranking as an effective method for enhancing AI-assisted software development.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). This seminal work introduced Codex and the HumanEval benchmark, establishing the standard sampling-and-filtering paradigms evaluated by CodeRanker.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). This paper introduced the Mostly Basic Programming Problems (MBPP) benchmark and analyzed scaling properties of large language models for program synthesis, serving as a primary dataset and foundation for the source.
- Paper: Natural Language to Code Translation with Execution, Freda Shi et al. (2022). It provides the foundational execution-based consensus ranking method (MBR-EXEC) that the source directly addresses and improves upon by ranking candidate code without runtime execution.
- Paper: Competition-level code generation with AlphaCode, Yujia Li et al. (2022). This paper pioneered large-scale sampling followed by execution-based test filtering and clustering for code generation, establishing the execution-dependent baseline framework contrasted by the source.
- Paper: CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation, Yue Wang et al. (2021). It establishes unified pre-trained encoder-decoder representations for code understanding and defect detection that inform neural code modeling.
- Paper: CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation, Shuai Lu et al. (2021). This work establishes the standard benchmark tasks and baseline models for code intelligence and program defect prediction.
- Paper: Teaching Large Language Models to Self-Debug, Xinyun Chen et al. (2023). It builds upon the concept of fault identification in generated code by enabling LLMs to iteratively debug and repair their own runtime errors.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). It advances the evaluation of LLM code correctness beyond standard unit tests by introducing rigorous input-mutation framework EvalPlus to catch subtle runtime faults.
- Paper: Code Llama: Open Foundation Models for Code, Baptiste Rozière et al. (2023). It scales open foundation models for code generation across the same core benchmarks (HumanEval, MBPP, APPS) where neural ranking methods are applied.
- Paper: Repair Is Nearly Generation: Multilingual Program Repair with LLMs, Harshit Joshi et al. (2023). It extends the treatment of minor syntax and runtime defects in LLM code by formulating automated program repair as a generation task.
- Paper: Generative Verifiers: Reward Modeling as Next-Token Prediction, Lunjun Zhang et al. (2025). It extends non-execution verifiers and best-of-N selection techniques by framing solution verification as generative next-token prediction.
- Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Naman Jain et al. (2025). It develops a comprehensive, contamination-free evaluation suite assessing LLM capabilities on code generation, repair, and execution prediction without running programs.
