Towards Functional Correctness of Large Code Models with Selective Generation
Jaewoo JeongTaesoo KimSangdon Park
Proposes a selective code generation framework that uses automatically generated unit tests to assess functional correctness, enabling large language models to abstain from unreliable outputs and theoretically control false discovery rates.
Large language models frequently generate code that contains subtle errors, known as code hallucination. In high-stakes production environments, deploying functionally incorrect software introduces severe security vulnerabilities, compliance risks, and operational failures. While techniques exist to measure uncertainty in natural language, verifying the functional correctness of generated programming code remains difficult because static code text cannot easily reveal whether a program will execute as intended under all possible conditions.
The article develops a certified framework that enables code models to selectively generate code by abstaining with an "I don't know" response when uncertainty is high. The primary objective is to mathematically bound the error rate—specifically the false discovery rate among non-abstained solutions—to a user-specified risk tolerance level while maximizing the proportion of accepted, useful code outputs.
To achieve this, the authors exploit the executable nature of software by re-purposing dynamic analysis tools, specifically fuzz testing, to automatically generate large suites of execution-based unit tests. Using these tests, the framework establishes a statistical standard of functional correctness and learns a confidence-calibrated selection threshold across calibration data. The authors evaluate this approach across five open and closed code models (including GPT-4o, Gemini 1.5 Pro, and DeepSeek-R1), four programming languages (C++, Java, JavaScript, and Perl), and four standard benchmark datasets.
The findings demonstrate that the selective generation framework reliably enforces the pre-set error bounds where conventional baseline methods fail. On difficult coding benchmarks, applying the method to advanced models reduced error rates from over 43% down to approximately 22%, strictly complying with a 30% error ceiling. In contrast, heuristic alternatives either routinely breached error thresholds or collapsed in utility by rejecting almost all generations. Furthermore, expanding automated test suites via fuzzing significantly sharpened calibration accuracy and improved evaluation reliability compared to small, manually curated test sets.
These results show that organizations do not need to rely solely on expensive model fine-tuning to achieve dependable automated coding. By implementing dynamic test-driven selective generation as a post-processing filter, engineering teams can set strict reliability targets and drastically lower the risk of deploying broken code into production systems. This transforms code generation from an unconstrained liability into a controllable, risk-managed workflow.
Organizations deploying automated coding systems should consider implementing dynamic test execution and confidence-calibrated abstention mechanisms on top of existing code models. However, users should note that the mathematical guarantees depend on standard statistical assumptions that the incoming distribution of coding prompts remains stable. Furthermore, when underlying model quality is low or scoring functions are poorly calibrated, the system preserves safety primarily by increasing abstention rates, reducing overall output efficiency.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). This paper establishes the necessity of rigorous, mutation-expanded unit tests to evaluate functional correctness and expose subtle flaws in code models, forming the direct motivation for dynamic-test evaluation.
- Paper: Natural Language to Code Translation with Execution, Freda Shi et al. (2022). This work introduces execution-based candidate selection and consensus filtering on sample inputs, providing a core foundation for utilizing runtime verification to select correct code.
- Paper: Fault-Aware Neural Code Rankers, Jeevana Priya Inala et al. (2022). This study analyzes how to predict correctness and rank generated code candidates based on execution error awareness, motivating selective generation and abstention paradigms.
- Paper: TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark, Kush Jain et al. (2025). This paper formalizes benchmarks and execution metrics for automated unit test generation, providing prerequisite context for synthesizing functional test suites.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). This foundational work establishes the pass@k benchmark methodology and unit-test execution framework for measuring functional correctness in large code models.
- Paper: Teaching Large Language Models to Self-Debug, Xinyun Chen et al. (2023). This paper demonstrates using dynamic execution feedback and runtime traces to detect bugs and guide model refinement.
- Paper: Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness, Jiuhai Chen et al. (2024). This work explores quantifying uncertainty and calibrating confidence to abstain from erroneous language model generations, laying theoretical groundwork for selective generation.
- Paper: The Verification Horizon: No Silver Bullet for Coding Agent Rewards, Binghai Wang et al. (2026). This work extends functional verification by examining the limits of test- and judge-based reward models against reward hacking in autonomous coding agents.
- Paper: BlueCodeAgent: A Blue Teaming Agent Enabled by Automated Red Teaming for CodeGen AI, Chengquan Guo et al. (2026). This paper applies dynamic unit testing and runtime execution within an agentic pipeline to verify security vulnerabilities and defend code generation models.
- Paper: Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?, Wang Bill Zhu et al. (2026). This paper investigates whether models pass unit tests through precise minimal debugging versus unconstrained code regeneration, extending the evaluation of functional test-passing behavior.
- Paper: SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation, Yuru Feng et al. (2026). This work builds on autonomous test generation by synthesizing executable unit tests to verify failure hypotheses during tree-search skill refinement.
- Paper: CornerCase: Automated Extremal Testing of Protocol Implementations using LLMs, Rathin Singha et al. (2026). This research extends automated specification-based test generation and execution to uncover extremal boundary defects across network protocol implementations.
- Paper: Self Improvement via Fast Tree-search, Xinghong Fu et al. (2026). This work investigates tree-search optimization for self-improving coding agents by decoupling expensive full-benchmark test execution with fast evaluators.
- Paper: Code as Agent Harness, Xuying Ning et al. (2026). This survey generalizes the role of executable code and verification test harnesses into the core operating substrate for autonomous agent planning and control.
