Towards Functional Correctness of Large Code Models with Selective Generation

Jaewoo JeongTaesoo KimSangdon Park

article2026arXiv1 citations

Proposes a selective code generation framework that uses automatically generated unit tests to assess functional correctness, enabling large language models to abstain from unreliable outputs and theoretically control false discovery rates.

Listen

Large language models frequently generate code that contains subtle errors, known as code hallucination. In high-stakes production environments, deploying functionally incorrect software introduces severe security vulnerabilities, compliance risks, and operational failures. While techniques exist to measure uncertainty in natural language, verifying the functional correctness of generated programming code remains difficult because static code text cannot easily reveal whether a program will execute as intended under all possible conditions.

The article develops a certified framework that enables code models to selectively generate code by abstaining with an "I don't know" response when uncertainty is high. The primary objective is to mathematically bound the error rate—specifically the false discovery rate among non-abstained solutions—to a user-specified risk tolerance level while maximizing the proportion of accepted, useful code outputs.

To achieve this, the authors exploit the executable nature of software by re-purposing dynamic analysis tools, specifically fuzz testing, to automatically generate large suites of execution-based unit tests. Using these tests, the framework establishes a statistical standard of functional correctness and learns a confidence-calibrated selection threshold across calibration data. The authors evaluate this approach across five open and closed code models (including GPT-4o, Gemini 1.5 Pro, and DeepSeek-R1), four programming languages (C++, Java, JavaScript, and Perl), and four standard benchmark datasets.

The findings demonstrate that the selective generation framework reliably enforces the pre-set error bounds where conventional baseline methods fail. On difficult coding benchmarks, applying the method to advanced models reduced error rates from over 43% down to approximately 22%, strictly complying with a 30% error ceiling. In contrast, heuristic alternatives either routinely breached error thresholds or collapsed in utility by rejecting almost all generations. Furthermore, expanding automated test suites via fuzzing significantly sharpened calibration accuracy and improved evaluation reliability compared to small, manually curated test sets.

These results show that organizations do not need to rely solely on expensive model fine-tuning to achieve dependable automated coding. By implementing dynamic test-driven selective generation as a post-processing filter, engineering teams can set strict reliability targets and drastically lower the risk of deploying broken code into production systems. This transforms code generation from an unconstrained liability into a controllable, risk-managed workflow.

Organizations deploying automated coding systems should consider implementing dynamic test execution and confidence-calibrated abstention mechanisms on top of existing code models. However, users should note that the mathematical guarantees depend on standard statistical assumptions that the incoming distribution of coding prompts remains stable. Furthermore, when underlying model quality is low or scoring functions are poorly calibrated, the system preserves safety primarily by increasing abstention rates, reducing overall output efficiency.

Cover for Towards Functional Correctness of Large Code Models with Selective Generation

Abstract

The hallucination of code generation models hinders their applicability to systems requiring higher safety standards. One critical bottleneck in addressing code hallucination is the difficulty of identifying the functional correctness of generated code, due to its unnatural form. We address this core bottleneck by automatically generating unit tests using dynamic code analysis tools, leveraging the \emph{executable nature} of code. Accordingly, we propose a \emph{selective code generator} that abstains from uncertain generations -- based on the functional correctness evaluated by generated unit tests -- to theoretically control the correctness among non-abstained answers, \ie the false discovery rate. Finally, we propose to use generated unit tests in evaluation as well as in learning for precise code evaluation, calling this paradigm \emph{FuzzEval}. We demonstrate the efficacy of our method along with the controllability of code hallucination and reasonable selection efficiency.

Table of Contents

  • 1 Introduction
  • 1.1 Related Work
  • 2 Preliminary
  • 3 Problem
  • 4 Method: Selective Code Generation
  • 4.1 Code Entailment
  • 4.2 False Discovery Rate via Code Entailment
  • 4.3 Code Entailment Estimation
  • 4.4 False Discovery Rate Bound
  • 4.5 FDR-CE Control Algorithm
  • 4.6 Controllability Guarantee
  • 4.7 FuzzEval: Evaluation via Fuzzing
  • 5 Experiments
  • 5.1 Setup
  • 5.2 Results
  • 5.2.1 Controllability and Selection Efficiency
  • 5.2.2 Benefit of Fuzzing
  • 5.2.3 Extension over Prior Work
  • 5.2.4 Effect of Scoring Function
  • 6 Conclusion
  • References
  • A Preliminary
  • B Qualitative Result of our Methods

Knowls

  1. Knowl 1 — alpha-Code Entailment

    definition

    Let U\mathcal{U} and V\mathcal{V} denote the sets of input and output execution states for code snippets, respectively. Let X\mathcal{X} be the set of problem descriptions (prompts), and Y\mathcal{Y} be the set of code snippets. A unit test generator F:X→Δ(U×V)\mathcal{F}: \mathcal{X} \to \Delta(\mathcal{U} \times \mathcal{V}) produces input-output state pairs (u,v)∼F(x)(u, v) \sim \mathcal{F}(x) for a given problem description x∈Xx \in \mathcal{X}.

    Given an error tolerance parameter α∈[0,1]\alpha \in [0, 1], a unit test generator F(x)\mathcal{F}(x) is said to α\alpha-entail a generated code snippet y^∈Y\hat{y} \in \mathcal{Y} if the expected functional correctness of y^\hat{y} on test cases sampled from F(x)\mathcal{F}(x) satisfies:

    P(u,v)∼F(x){y^(u)=v}≥1−α\mathbb{P}_{(u, v) \sim \mathcal{F}(x)} \{\hat{y}(u) = v\} \ge 1 - \alpha

    The set of all α\alpha-entailed code snippets for problem xx is denoted by Eα(x)E_\alpha(x):

    Eα(x):={yˉ∈Y  |  P(u,v)∼F(x){yˉ(u)=v}≥1−α}E_\alpha(x) := \left\{\bar{y} \in \mathcal{Y} \;\middle|\; \mathbb{P}_{(u, v) \sim \mathcal{F}(x)} \{\bar{y}(u) = v\} \ge 1 - \alpha\right\}

    This formulation provides a relaxed, one-directional probabilistic standard of semantic correctness between a target specification and generated code.

  2. Knowl 2 — False Discovery Rate with Code Entailment

    definition

    In selective code generation, a selective generator S^:X→Y∪{IDK}\hat{S}: \mathcal{X} \to \mathcal{Y} \cup \{\text{IDK}\} either outputs generated code G(x)∈YG(x) \in \mathcal{Y} produced by a base code generator G:X→YG: \mathcal{X} \to \mathcal{Y} when a binary selection function s^(x)=1\hat{s}(x) = 1, or abstains from answering by returning IDK\text{IDK} ("I don't know") when s^(x)=0\hat{s}(x) = 0:

    S^(x):={G(x)if s^(x)=1IDKotherwise\hat{S}(x) := \begin{cases} G(x) & \text{if } \hat{s}(x) = 1 \\ \text{IDK} & \text{otherwise} \end{cases}

    Let D\mathcal{D} be an underlying distribution over problem descriptions X\mathcal{X} and reference code Y\mathcal{Y}, and let Eα(x)E_\alpha(x) be the set of α\alpha-entailed code snippets with respect to a unit test generator F(x)\mathcal{F}(x). The False Discovery Rate with Code Entailment (FDR-CE), denoted by Rα(S^)R_\alpha(\hat{S}), measures the probability of non-entailed code among the non-abstained predictions:

    Rα(S^):=Px∼D,F,S^(S^(x)∉Eα(x)  |  S^(x)≠IDK)R_\alpha(\hat{S}) := \mathbb{P}_{x \sim \mathcal{D}, \mathcal{F}, \hat{S}} \left( \hat{S}(x) \notin E_\alpha(x) \;\middle|\; \hat{S}(x) \ne \text{IDK} \right)

    This metric serves as the primary risk objective for controlling functional hallucination in code generation.

  3. Knowl 3 — FDR-CE Upper Bound via Estimated Code Entailment

    theoretical result

    Because the exact entailment set Eα(x)E_\alpha(x) cannot be computed analytically, it is approximated by an estimated entailment set E^α,εE(x)\hat{E}_{\alpha, \varepsilon_E}(x) using empirical test execution.

    Let Ux={(uj,vj)}j=1nx∼F(x)nxU_x = \{(u_j, v_j)\}_{j=1}^{n_x} \sim \mathcal{F}(x)^{n_x} be a sample of nxn_x unit tests from the test generator F(x)\mathcal{F}(x), and let k^=∑(u,v)∈Ux1(yˉ(u)=v)\hat{k} = \sum_{(u, v) \in U_x} \mathbf{1}(\bar{y}(u) = v) be the number of passed unit tests for candidate code yˉ\bar{y}. Using the Clopper-Pearson lower binomial tail bound L^(F(x),yˉ,nx,εE):=L^Binom(k^;nx,εE)\hat{L}(\mathcal{F}(x), \bar{y}, n_x, \varepsilon_E) := \hat{L}_{\text{Binom}}(\hat{k}; n_x, \varepsilon_E), which satisfies P(L^≤P(u,v)∼F(x){yˉ(u)=v})≥1−εE\mathbb{P}(\hat{L} \le \mathbb{P}_{(u,v) \sim \mathcal{F}(x)}\{\bar{y}(u) = v\}) \ge 1 - \varepsilon_E, the estimated entailment set is defined as:

    E^α,εE(x):={yˉ∈Y  |  L^(F(x),yˉ,nx,εE)≥1−α}\hat{E}_{\alpha, \varepsilon_E}(x) := \left\{\bar{y} \in \mathcal{Y} \;\middle|\; \hat{L}(\mathcal{F}(x), \bar{y}, n_x, \varepsilon_E) \ge 1 - \alpha\right\}

    Let the empirical surrogate risk evaluated on the estimated entailment set be Rα,εE(S^):=P(S^(x)∉E^α,εE(x)∣S^(x)≠IDK)R_{\alpha, \varepsilon_E}(\hat{S}) := \mathbb{P}(\hat{S}(x) \notin \hat{E}_{\alpha, \varepsilon_E}(x) \mid \hat{S}(x) \ne \text{IDK}). For any α∈(0,1)\alpha \in (0, 1), estimation confidence parameter εE∈(0,1)\varepsilon_E \in (0, 1), and selective generator S^\hat{S}, the true FDR-CE is upper-bounded by:

    Rα(S^)≤εE+Rα,εE(S^)R_\alpha(\hat{S}) \le \varepsilon_E + R_{\alpha, \varepsilon_E}(\hat{S})

    where εE\varepsilon_E upper-bounds the false entailment rate (FER) incurred by the finite binomial sample estimation.

  4. Knowl 4 — Selective Code Generator Learning Algorithm

    algorithm

    The selective code generator learning algorithm ASCG\mathcal{A}_{\text{SCG}} computes an optimal acceptance threshold τ^\hat{\tau} for the selection function s^(x)=1(f(x,G(x))≥τ)\hat{s}(x) = \mathbf{1}(f(x, G(x)) \ge \tau), where ff is a confidence scoring function and GG is a base code generator. The algorithm searches for the minimum threshold τ\tau (thereby maximizing selection efficiency) subject to controlling the upper bound on the FDR-CE within a user-specified target risk εS\varepsilon_S.

    Given a calibration dataset Z={(xi,yi)}i=1nZ = \{(x_i, y_i)\}_{i=1}^n, the algorithm sorts ZZ in ascending order of confidence scores and executes binary search over the sorted scores. At each step ii, candidate threshold τS(i)\tau_S^{(i)} defines the selected subset Z^(i)={(x,y)∈Z∣f(x,G(x))≥τS(i)}\hat{Z}^{(i)} = \{(x, y) \in Z \mid f(x, G(x)) \ge \tau_S^{(i)}\}. The number of empirical false discoveries is counted as k^(i)=∑(x,y)∈Z^(i)1(G(x)∉E^α,εE(x))\hat{k}^{(i)} = \sum_{(x,y) \in \hat{Z}^{(i)}} \mathbf{1}(G(x) \notin \hat{E}_{\alpha, \varepsilon_E}(x)). The upper binomial tail bound U^Binom(k^(i);∣Z^(i)∣,δS/⌈log⁡2∣Z∣⌉)\hat{U}_{\text{Binom}}(\hat{k}^{(i)}; |\hat{Z}^{(i)}|, \delta_S / \lceil \log_2 |Z| \rceil) is computed with a Bonferroni correction over the ⌈log⁡2∣Z∣⌉\lceil \log_2 |Z| \rceil binary search steps. If εE+U^Binom≤εS\varepsilon_E + \hat{U}_{\text{Binom}} \le \varepsilon_S, the search explores lower thresholds; otherwise, it increases the threshold.

    procedure LEARNSCG(f, Z, G, \alpha, \varepsilon_E, \delta_S, \varepsilon_S)
        Z_prime = SORT_BY_SCORE(Z, f, G)
        low = 1
        high = |Z|
        tau_best = f(Z_prime[high].x, G(Z_prime[high].x))
        U_best = 1.0
        for step = 1 to ceil(log2(|Z|)) do
            mid = ceil((low + high) / 2)
            tau_candidate = f(Z_prime[mid].x, G(Z_prime[mid].x))
            Z_hat = { (x, y) in Z_prime | f(x, G(x)) >= tau_candidate }
            k_hat = sum_{(x, y) in Z_hat} 1(G(x) not in E_hat_{alpha, epsilon_E}(x))
            delta_step = delta_S / ceil(log2(|Z|))
            U_hat = epsilon_E + U_Binom(k_hat, |Z_hat|, delta_step)
            if U_hat <= epsilon_S then
                tau_best = tau_candidate
                U_best = U_hat
                high = mid
            else
                low = mid + 1
            end if
        end for
        return tau_best, U_best
  5. Knowl 5 — PAC Controllability Guarantee for Selective Code Generation

    theoretical result

    Let D\mathcal{D} be a distribution over problem prompts and code snippets, G:X→YG: \mathcal{X} \to \mathcal{Y} be a code generator, f:X×Y→Rf: \mathcal{X} \times \mathcal{Y} \to \mathbb{R} be a scoring function, and F\mathcal{F} be a unit test generator. Given a calibration set Z∼DnZ \sim \mathcal{D}^n, let (S^,U^):=ASCG(Z)(\hat{S}, \hat{U}) := \mathcal{A}_{\text{SCG}}(Z) be the selective generator and empirical upper bound returned by the calibration algorithm.

    For any target risk level εS∈(0,1)\varepsilon_S \in (0, 1), failure tolerance δS∈(0,1)\delta_S \in (0, 1), entailment tolerance α∈(0,1)\alpha \in (0, 1), and estimation parameter εE∈(0,1)\varepsilon_E \in (0, 1), the false discovery rate with code entailment satisfies:

    PZ∼Dn(Rα(S^)≤U^)≥1−δS\mathbb{P}_{Z \sim \mathcal{D}^n} \left( R_\alpha(\hat{S}) \le \hat{U} \right) \ge 1 - \delta_S

    When the optimization problem is feasible such that U^≤εS\hat{U} \le \varepsilon_S, the learned selective code generator satisfies Rα(S^)≤εSR_\alpha(\hat{S}) \le \varepsilon_S with probability at least 1−δS1 - \delta_S over calibration datasets, without requiring human-annotated entailment labels.

  6. Knowl 6 — Adaptive Unit Test Size Computation via Binomial Tail Bound

    algorithm

    To dynamically determine how many generated unit tests nxn_x are necessary to verify α\alpha-code entailment with confidence level 1−εE1 - \varepsilon_E for a problem xx and candidate code yˉ\bar{y}, the procedure COMPUNITTESTSIZE incrementally executes unit tests (u,v)∼F(x)(u, v) \sim \mathcal{F}(x) and evaluates the lower binomial tail bound L^(F(x),yˉ,nx,εE)\hat{L}(\mathcal{F}(x), \bar{y}, n_x, \varepsilon_E). Sampling terminates once the lower bound exceeds the target threshold 1−α1 - \alpha or reaches a predefined computational budget nmaxn_{\text{max}}.

    procedure COMPUNITTESTSIZE(\alpha, \varepsilon_E, \mathcal{F}(x), \bar{y}, n_max)
        n_x = 0
        while L_hat(\mathcal{F}(x), \bar{y}, n_x, \varepsilon_E) < 1 - \alpha do
            (u, v) ~ \mathcal{F}(x)
            n_x = n_x + 1
            if n_x >= n_max then
                break
            end if
        end while
        return n_x
  7. Knowl 7 — Confidence Scoring Functions for Selective Code Generation

    model/method

    The selective code generator relies on a scoring function f(x,G(x))f(x, G(x)) to rank the certainty of generated code snippets. Two primary scoring formulations are utilized:

    1. Length-normalized log-probability (fnormf_{\text{norm}}):

    fnorm(x,G(x))=1∣G(x)∣∑i=1∣G(x)∣ln⁡pif_{\text{norm}}(x, G(x)) = \frac{1}{|G(x)|} \sum_{i=1}^{|G(x)|} \ln p_i

    where pip_i is the probability assigned by the code generator GG to generating the ii-th token conditional on xx and previous tokens.

    1. Mixed Scoring Function (fmixedf_{\text{mixed}}): Combines token-level model likelihood with execution pass rate on a small hold-out set of UU unit tests:

    fmixed(x,G(x))=0.5fnorm(x,G(x))+0.5(1U∑j=1U1(G(x)(uj)=vj))f_{\text{mixed}}(x, G(x)) = 0.5 f_{\text{norm}}(x, G(x)) + 0.5 \left( \frac{1}{U} \sum_{j=1}^U \mathbf{1}(G(x)(u_j) = v_j) \right)

    where (uj,vj)∼F(x)(u_j, v_j) \sim \mathcal{F}(x). Other baseline scoring variants include sequence log-probability fseq(x,G(x))=∑iln⁡pif_{\text{seq}}(x, G(x)) = \sum_i \ln p_i, minimum token log-probability fmin(x,G(x))=min⁡iln⁡pif_{\text{min}}(x, G(x)) = \min_i \ln p_i, and verbalized model confidence fverb(x,G(x))f_{\text{verb}}(x, G(x)) prompted directly from the model.

  8. Knowl 8 — Empirical FDR-CE and Selection Efficiency on Coding Benchmarks with GPT-4o

    data/table

    Experimental comparison of Selective Code Generation (SCG) against baseline selective approaches on four code generation benchmarks (APPS-f, MBPP-f, HumanEval-f, and Mercury-f) using GPT-4o. The parameters are fixed at α=0.35\alpha = 0.35, δS=0.1\delta_S = 0.1, εE=0.05\varepsilon_E = 0.05. Baseline methods compared are:

    • Unselective generation (τ=−∞\tau = -\infty)
    • SCG-EM: Exact string match with solution code (Geifman & El-Yaniv, 2017)
    • SCG-manual: Selects the top k%k\% highest confidence scores
    • SCG-small: SCG calibrated with only 21 unit tests per problem
    • SCG-H: Heuristic ablation of SCG setting εE=0\varepsilon_E = 0 (omitting false entailment estimation error)

    SCG consistently maintains the empirical FDR-CE below the specified εS\varepsilon_S target while achieving the highest valid selection efficiency. SCG-EM yields near-zero efficiency because correct functional code rarely matches canonical solutions character-for-character. SCG-H and SCG-manual violate the desired FDR-CE bounds under several configurations.

    Methods w/o Sel. w/ Selective Generation
    τ=−∞\tau = -\infty SCG-EM SCG-manual SCG-small SCG-H SCG
    APPS-f (εS=0.3\varepsilon_S = 0.3)
    1−PASS@1(↓)1 - \text{PASS@1} (\downarrow) 0.436±0.0120.436 \pm 0.012 0.020±0.0980.020 \pm 0.098 0.293±0.0140.293 \pm 0.014 0.226±0.0150.226 \pm 0.015 0.282±0.0200.282 \pm 0.020 0.227±0.0170.227 \pm 0.017
    FDR-CE(↓)\text{FDR-CE} (\downarrow) 0.431±0.0110.431 \pm 0.011 0.020±0.0990.020 \pm 0.099 0.291±0.0140.291 \pm 0.014 0.226±0.0150.226 \pm 0.015 0.280±0.0200.280 \pm 0.020 0.224±0.018\mathbf{0.224 \pm 0.018}
    EFFICIENCY(↑)\text{EFFICIENCY} (\uparrow) 1.000±0.0001.000 \pm 0.000 0.000±0.0000.000 \pm 0.000 0.497±0.0140.497 \pm 0.014 0.339±0.0120.339 \pm 0.012 0.463±0.0190.463 \pm 0.019 0.337±0.011\mathbf{0.337 \pm 0.011}
    MBPP-f (εS=0.4\varepsilon_S = 0.4)
    1−PASS@11 - \text{PASS@1} 0.299±0.0220.299 \pm 0.022 0.000±0.0000.000 \pm 0.000 0.258±0.0400.258 \pm 0.040 – 0.301±0.0280.301 \pm 0.028 0.304±0.0320.304 \pm 0.032
    FDR-CE\text{FDR-CE} 0.294±0.0270.294 \pm 0.027 0.000±0.0000.000 \pm 0.000 0.254±0.0410.254 \pm 0.041 – 0.296±0.0290.296 \pm 0.029 0.300±0.032\mathbf{0.300 \pm 0.032}
    EFFICIENCY\text{EFFICIENCY} 1.000±0.0001.000 \pm 0.000 0.001±0.0030.001 \pm 0.003 0.499±0.0380.499 \pm 0.038 – 0.997±0.0040.997 \pm 0.004 0.996±0.006\mathbf{0.996 \pm 0.006}
    HUMANEVAL-f (εS=0.3\varepsilon_S = 0.3)
    1−PASS@11 - \text{PASS@1} 0.185±0.0580.185 \pm 0.058 0.000±0.0000.000 \pm 0.000 0.142±0.0800.142 \pm 0.080 – 0.156±0.1650.156 \pm 0.165 0.069±0.1730.069 \pm 0.173
    FDR-CE\text{FDR-CE} 0.207±0.0640.207 \pm 0.064 0.000±0.0000.000 \pm 0.000 0.156±0.0860.156 \pm 0.086 – 0.145±0.1230.145 \pm 0.123 0.049±0.111\mathbf{0.049 \pm 0.111}
    EFFICIENCY\text{EFFICIENCY} 1.000±0.0001.000 \pm 0.000 0.008±0.0170.008 \pm 0.017 0.492±0.0800.492 \pm 0.080 – 0.578±0.4010.578 \pm 0.401 0.164±0.118\mathbf{0.164 \pm 0.118}
    MERCURY-f (εS=0.3\varepsilon_S = 0.3)
    1−PASS@11 - \text{PASS@1} 0.174±0.0180.174 \pm 0.018 0.000±0.0000.000 \pm 0.000 0.138±0.0240.138 \pm 0.024 – 0.169±0.0170.169 \pm 0.017 0.170±0.0200.170 \pm 0.020
    FDR-CE\text{FDR-CE} 0.170±0.0190.170 \pm 0.019 0.000±0.0000.000 \pm 0.000 0.133±0.0220.133 \pm 0.022 – 0.164±0.0190.164 \pm 0.019 0.165±0.020\mathbf{0.165 \pm 0.020}
    EFFICIENCY\text{EFFICIENCY} 1.000±0.0001.000 \pm 0.000 0.001±0.0020.001 \pm 0.002 0.504±0.0330.504 \pm 0.033 – 0.998±0.0030.998 \pm 0.003 0.998±0.002\mathbf{0.998 \pm 0.002}
  9. Knowl 9 — Integration of SCG with Test-Time Search and Debugging Baselines

    data/table

    SCG operates as a modular, post-processing filter that can be applied to outputs generated by external code synthesis, search, and debugging methods. Below are empirical results evaluated on MBPP-f and HumanEval-f using GPT-3.5-Turbo as the base generator, alongside three advanced code generation algorithms: CodeT (dual-execution test ranking), LDB (stepwise execution debugger), and SFS (smarter code space search). Parameters are set to α=0.3\alpha = 0.3, δS=0.1\delta_S = 0.1, εE=0.05\varepsilon_E = 0.05, and target risk εS=0.25\varepsilon_S = 0.25, using the mixed scoring function fmixedf_{\text{mixed}}.

    Applying SCG consistently brings the empirical FDR-CE and 1−PASS@11-\text{PASS@1} error below the target bound εS=0.25\varepsilon_S = 0.25 across all generation paradigms, while preserving selection efficiency between 57.6%57.6\% and 87.8%87.8\%.

    Methods Base Model CodeT LDB SFS
    w/o SCG w/ SCG w/o SCG w/ SCG w/o SCG w/ SCG w/o SCG w/ SCG
    MBPP-F
    1−PASS@1(↓)1 - \text{PASS@1} (\downarrow) 0.496±0.0390.496 \pm 0.039 0.163±0.0580.163 \pm 0.058 0.502±0.0340.502 \pm 0.034 0.155±0.0460.155 \pm 0.046 0.442±0.0440.442 \pm 0.044 0.150±0.0450.150 \pm 0.045 0.486±0.0370.486 \pm 0.037 0.151±0.0430.151 \pm 0.043
    FDR-CE(↓)\text{FDR-CE} (\downarrow) 0.491±0.0420.491 \pm 0.042 0.145±0.052\mathbf{0.145 \pm 0.052} 0.498±0.0340.498 \pm 0.034 0.148±0.045\mathbf{0.148 \pm 0.045} 0.442±0.0440.442 \pm 0.044 0.142±0.043\mathbf{0.142 \pm 0.043} 0.483±0.0390.483 \pm 0.039 0.140±0.042\mathbf{0.140 \pm 0.042}
    EFFICIENCY(↑)\text{EFFICIENCY} (\uparrow) 1.000±0.0001.000 \pm 0.000 0.576±0.0390.576 \pm 0.039 1.000±0.0001.000 \pm 0.000 0.587±0.0400.587 \pm 0.040 1.000±0.0001.000 \pm 0.000 0.655±0.0440.655 \pm 0.044 1.000±0.0001.000 \pm 0.000 0.604±0.0380.604 \pm 0.038
    HUMANEVAL-F
    1−PASS@1(↓)1 - \text{PASS@1} (\downarrow) 0.294±0.0570.294 \pm 0.057 0.100±0.0690.100 \pm 0.069 0.308±0.0580.308 \pm 0.058 0.129±0.0720.129 \pm 0.072 0.253±0.0690.253 \pm 0.069 0.143±0.0710.143 \pm 0.071 0.274±0.0680.274 \pm 0.068 0.135±0.0590.135 \pm 0.059
    FDR-CE(↓)\text{FDR-CE} (\downarrow) 0.305±0.0550.305 \pm 0.055 0.112±0.075\mathbf{0.112 \pm 0.075} 0.312±0.0620.312 \pm 0.062 0.135±0.069\mathbf{0.135 \pm 0.069} 0.230±0.0620.230 \pm 0.062 0.121±0.066\mathbf{0.121 \pm 0.066} 0.288±0.0720.288 \pm 0.072 0.158±0.058\mathbf{0.158 \pm 0.058}
    EFFICIENCY(↑)\text{EFFICIENCY} (\uparrow) 1.000±0.0001.000 \pm 0.000 0.778±0.0650.778 \pm 0.065 1.000±0.0001.000 \pm 0.000 0.808±0.0610.808 \pm 0.061 1.000±0.0001.000 \pm 0.000 0.878±0.0520.878 \pm 0.052 1.000±0.0001.000 \pm 0.000 0.850±0.0560.850 \pm 0.056
  10. Knowl 10 — FuzzEval Dynamic Unit Test Generation Methodology

    model/method

    FuzzEval replaces sparse, manually written benchmark test suites with large-scale test suites generated via coverage-guided fuzzing (e.g., Atheris for Python) on reference canonical code solutions.

    To construct FuzzEval datasets (APPS-f, Mercury-f, HumanEval-f, and MBPP-f), reference solutions are instrumented with fuzzing harnesses that map random binary byte seeds s∼DUs \sim \mathcal{D}_\mathcal{U} into typed program inputs u∈Uu \in \mathcal{U} satisfying the problem's input preconditions. The reference solution yy is executed repeatedly on mutated inputs to explore diverse control flow and branch execution paths, capturing the corresponding output states v=y(u)v = y(u) to yield input-output pairs (u,v)(u, v). For each benchmark problem, at least 600 input-output pairs are extracted and split between calibration and evaluation sets, significantly reducing evaluation error in expected functional correctness compared to conventional sparse test suites.

  11. Knowl 11 — Limitations of Selective Code Generation

    limitation

    The selective code generation approach is subject to two main limitations:

    1. Dependence on Base Generator Accuracy and Score Calibration: While SCG guarantees that the FDR-CE among retained generations does not exceed εS\varepsilon_S, the selection efficiency (the proportion of non-abstained prompts) is strictly limited by the base code model's underlying performance. If the base model produces poorly calibrated confidence scores (e.g., CodeLlama 13B), the calibration algorithm may fail to find a threshold τ\tau satisfying tight risk targets εS\varepsilon_S, returning a trivial selective generator or defaulting to the empirical upper bound U^\hat{U}.

    2. Independent and Identically Distributed (i.i.d.) Assumption: The PAC-style statistical guarantees for controlling the false discovery rate rely on the assumption that test queries follow the same distribution D\mathcal{D} as the calibration dataset Z∼DnZ \sim \mathcal{D}^n. In the presence of distribution shift between calibration and runtime inference, the theoretical FDR-CE bound may not hold.

Coverage note — None was omitted; all primary contributions including theoretical definitions, the FDR-CE bound, calibration and unit test generation algorithms, scoring functions, core empirical tables, and stated limitations were fully extracted.

References

  1. 1.Ahn, J., Verma, R., Lou, R., Liu, D., Zhang, R., and Yin, W. Large language models for mathematical reasoning: Progresses and challenges. In Falk, N., Papi, S., and Zhang, M. (eds.), Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, pp. 225–237, St. Julian’s, Malta, March 2024. Association for Computational Linguistics. URL https://aclanthology.org/2024.eacl-srw.17/.
  2. 2.Alshahwan, N., Chheda, J., Finogenova, A., Gokkaya, B., Harman, M., Harper, I., Marginean, A., Sengupta, S., and Wang, E. Automated unit test improvement using large language models at meta. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, FSE 2024, pp. 185–196, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400706585. doi: 10.1145/3663529.3663839. URL https://doi.org/10.1145/3663529.3663839.
  3. 3.Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., and Sutton, C. Program synthesis with large language models, 2021. URL https://arxiv.org/abs/2108.07732.
  4. 4.Bowman, S., Angeli, G., Potts, C., and Manning, C. D. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp. 632–642, 2015.
  5. 5.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners, 2020.
  6. 6.Cassano, F., Gouwar, J., Nguyen, D., Nguyen, S., Phipps-Costin, L., Pinckney, D., Yee, M.-H., Zi, Y., Anderson, C. J., Feldman, M. Q., et al. Multipl-e: A scalable and extensible approach to benchmarking neural code generation. arXiv preprint arXiv:2208.08227, 2022.
  7. 7.Chen, B., Zhang, F., Nguyen, A., Zan, D., Lin, Z., Lou, J.-G., and Chen, W. Codet: Code generation with generated tests. In The Eleventh International Conference on Learning Representations, 2023a. URL https://openreview.net/forum?id=ktrw68Cmu9c.
  8. 8.Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  9. 9.Chen, W., Ma, X., Wang, X., and Cohen, W. W. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. Transactions on Machine Learning Research, 2023b.
  10. 10.Chen, Y. and Su, Z. Guided differential testing of certificate validation in ssl/tls implementations. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering, ESEC/FSE 2015, pp. 793–804, New York, NY, USA, 2015. Association for Computing Machinery. ISBN 9781450336758. doi: 10.1145/2786805.2786835. URL https://doi.org/10.1145/2786805.2786835.
  11. 11.Claessen, K. and Hughes, J. Quickcheck: a lightweight tool for random testing of haskell programs. SIGPLAN Not., 35(9):268–279, September 2000. ISSN 0362-1340. doi: 10.1145/357766.351266. URL https://doi.org/10.1145/357766.351266.
  12. 12.Clopper, C. J. and Pearson, E. S. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika, 26(4):404–413, 1934.
  13. 13.Ding, Y., Peng, J., Min, M. J., Kaiser, G., Yang, J., and Ray, B. Semcoder: Training code language models with comprehensive semantics reasoning. In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., and Zhang, C. (eds.), Advances in Neural Information Processing Systems, volume 37, pp. 60275–60308. Curran Associates, Inc., 2024. URL https://proceedings.neurips.cc/paper_files/paper/2024/file/6efcc7fd8efeee29a050a79c843c90e0-Paper-Conference.pdf.
  14. 14.Du, M., Tuan, L. A., Ji, B., Liu, Q., and Ng, S.-K. Mercury: A code efficiency benchmark for code large language models. In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., and Zhang, C. (eds.), Advances in Neural Information Processing Systems, volume 37, pp. 16601–16622. Curran Associates, Inc., 2024. URL https://proceedings.neurips.cc/paper_files/paper/2024/file/1df1df43b58845650b8dada00fca9772-Paper-Datasets_and_Benchmarks_Track.pdf.
  15. 15.Fakhoury, S., Naik, A., Sakkas, G., Chakraborty, S., Musuvathi, M., and Lahiri, S. Exploring the effectiveness of llm based test-driven interactive code generation: User study and empirical evaluation. In Proceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings, ICSE-Companion ’24, pp. 390–391, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400705021. doi: 10.1145/3639478.3643525. URL https://doi.org/10.1145/3639478.3643525.
  16. 16.Fakhoury, S., Kuppe, M., Lahiri, S. K., Ramananandro, T., and Swamy, N. 3dgen: Ai-assisted generation of provably correct binary format parsers. In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), pp. 2535–2547, Los Alamitos, CA, USA, May 2025. IEEE Computer Society. doi: 10.1109/ICSE55347.2025.00173. URL https://doi.ieeecomputersociety.org/10.1109/ICSE55347.2025.00173.
  17. 17.Felsing, D., Grebing, S., Klebanov, V., Rümmer, P., and Ulbrich, M. Automating regression verification. In Proceedings of the 29th ACM/IEEE International Conference on Automated Software Engineering, ASE ’14, pp. 349–360, New York, NY, USA, 2014. Association for Computing Machinery. ISBN 9781450330138. doi: 10.1145/2642937.2642987. URL https://doi.org/10.1145/2642937.2642987.
  18. 18.Geifman, Y. and El-Yaniv, R. Selective classification for deep neural networks. Advances in neural information processing systems, 30, 2017.
  19. 19.Google. Atheris: A coverage-guided, native python fuzzer, 2020. URL https://github.com/google/atheris.
  20. 20.Guo, D., Yang, D., Zhang, H., Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, R., Ma, S., Bi, X., et al. Deepseek-r1 incentivizes reasoning in llms through reinforcement learning. Nature, 645(8081):633–638, 2025.
  21. 21.Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., and Steinhardt, J. Measuring coding challenge competence with APPS. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://openreview.net/forum?id=sD93GOzH3i5.
  22. 22.Hossain, S. B., Jiang, N., Zhou, Q., Li, X., Chiang, W.-H., Lyu, Y., Nguyen, H., and Tripp, O. A deep dive into large language models for automated bug localization and repair. Proc. ACM Softw. Eng., 1(FSE), July 2024. doi: 10.1145/3660773. URL https://doi.org/10.1145/3660773.
  23. 23.Huang, B., Lu, S., Wan, X., and Duan, N. Enhancing large language models in coding through multi-perspective self-consistency. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1429–1450, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.78. URL https://aclanthology.org/2024.acl-long.78/.
  24. 24.Jakobs, M.-C. and Wiesner, M. Peqtest: Testing functional equivalence. In Johnsen, E. B. and Wimmer, M. (eds.), Fundamental Approaches to Software Engineering, pp. 184–204, Cham, 2022. Springer International Publishing. ISBN 978-3-030-99429-7.
  25. 25.Kim, T., Han, H., Park, S., Jeong, D. R., Kim, D., Kim, D., Kim, E., Kim, J., Wang, J., Kim, K., Ji, S., Song, W., Zhao, H., Chin, A., Lee, G., Stevens, K., Alharthi, M., Zhai, Y., Zhang, C., Jang, J., Jang, Y., Askar, A., Kim, D., Fleischer, F., Cho, J., Kim, J., Ko, K., Yun, I., Park, S., Baik, D., Lee, H., Heo, H., Gwon, M., Lee, M., Baek, M., Min, S., Kim, W., Jin, Y., Park, Y., Choi, Y., Jung, J., Lee, G., Jang, J., Kim, K., Cha, Y., and Kim, Y. Atlantis: Ai-driven threat localization, analysis, and triage intelligence system, 2025. URL https://arxiv.org/abs/2509.14589.
  26. 26.Kuhn, L., Gal, Y., and Farquhar, S. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=VD-AYtP0dve.
  27. 27.Lahiri, S. K., Hawblitzel, C., Kawaguchi, M., and Rebêlo, H. Symdiff: A language-agnostic semantic diff tool for imperative programs. In Madhusudan, P. and Seshia, S. A. (eds.), Computer Aided Verification, pp. 712–717, Berlin, Heidelberg, 2012. Springer Berlin Heidelberg. ISBN 978-3-642-31424-7.
  28. 28.Le, H., Wang, Y., Gotmare, A. D., Savarese, S., and Hoi, S. CodeRL: Mastering code generation through pretrained models and deep reinforcement learning. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=WaGvb7OzySA.
  29. 29.Lee, M., Kim, K., Kim, T., and Park, S. Selective generation for controllable language models. Advances in Neural Information Processing Systems, 37:50494–50527, 2024.
  30. 30.Li, D., Cao, S., Cao, C., Li, X., Tan, S., Keutzer, K., Xing, J., Gonzalez, J. E., and Stoica, I. S*: Test time scaling for code generation. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 15964–15978, Suzhou, China, November 2025. Association for Computational Linguistics. ISBN 979-8-89176-335-7. doi: 10.18653/v1/2025.findings-emnlp.865. URL https://aclanthology.org/2025.findings-emnlp.865/.
  31. 31.Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Lago, A. D., Hubert, T., Choy, P., de Masson d’Autume, C., Babuschkin, I., Chen, X., Huang, P.-S., Welbl, J., Gowal, S., Cherepanov, A., Molloy, J., Mankowitz, D. J., Robson, E. S., Kohli, P., de Freitas, N., Kavukcuoglu, K., and Vinyals, O. Competition-level code generation with alphacode. Science, 378(6624):1092–1097, 2022. doi: 10.1126/science.abq1158. URL https://www.science.org/doi/abs/10.1126/science.abq1158.
  32. 32.Light, J., Wu, Y., Sun, Y., Yu, W., Liu, Y., Zhao, X., Hu, Z., Chen, H., and Cheng, W. SFS: Smarter code space search improves LLM inference scaling. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=MCHuGOkExF.
  33. 33.Liu, J., Xia, C. S., Wang, Y., and Zhang, L. Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA, 2023. Curran Associates Inc.
  34. 34.Manakul, P., Liusie, A., and Gales, M. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023.
  35. 35.Massalin, H. Superoptimizer: a look at the smallest program. In Proceedings of the Second International Conference on Architectual Support for Programming Languages and Operating Systems, ASPLOS II, pp. 122–126, New York, NY, USA, 1987. Association for Computing Machinery. ISBN 0818608056. doi: 10.1145/36206.36194. URL https://doi.org/10.1145/36206.36194.
  36. 36.McKeeman, W. M. Differential testing for software. Digit. Tech. J., 10:100–107, 1998.
  37. 37.Miller, B. P., Fredriksen, L., and So, B. An empirical study of the reliability of unix utilities. Communications of the ACM, 33(12):32–44, 1990.
  38. 38.Mohri, C. and Hashimoto, T. Language models with conformal factuality guarantees. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024.
  39. 39.OpenAI. Gpt-4 technical report, 2024. URL https://arxiv.org/abs/2303.08774.
  40. 40.OpenAI. Introducing gpt-4.1 in the api, 2025. URL https://openai.com/index/gpt-4-1/.
  41. 41.OpenAI, :, El-Kishky, A., Wei, A., Saraiva, A., Minaiev, B., Selsam, D., Dohan, D., Song, F., Lightman, H., Clavera, I., Pachocki, J., Tworek, J., Kuhn, L., Kaiser, L., Chen, M., Schwarzer, M., Rohaninejad, M., McAleese, N., o3 contributors, Murk, O., Garg, R., Shu, R., Sidor, S., Kosaraju, V., and Zhou, W. Competitive programming with large reasoning models, 2025. URL https://arxiv.org/abs/2502.06807.
  42. 42.Quach, V., Fisch, A., Schuster, T., Yala, A., Sohn, J. H., Jaakkola, T. S., and Barzilay, R. Conformal language modeling. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=pzUhfQ74c5.
  43. 43.Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Sauvestre, R., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C. C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G. Code llama: Open foundation models for code, 2024. URL https://arxiv.org/abs/2308.12950.
  44. 44.Semmle. Codeql, 2019. URL https://codeql.github.com/.
  45. 45.Spiess, C., Gros, D., Pai, K. S., Pradel, M., Rabin, M. R. I., Alipour, A., Jha, S., Devanbu, P., and Ahmed, T. Calibration and correctness of language models for code. In Proceedings of the IEEE/ACM 47th International Conference on Software Engineering, ICSE ’25, pp. 540–552. IEEE Press, 2025. ISBN 9798331505691. doi: 10.1109/ICSE55347.2025.00040. URL https://doi.org/10.1109/ICSE55347.2025.00040.
  46. 46.Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024. URL https://arxiv.org/abs/2403.05530.
  47. 47.Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 5433–5442, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.330. URL https://aclanthology.org/2023.emnlp-main.330/.
  48. 48.Tian, Y., Yan, W., Yang, Q., Zhao, X., Chen, Q., Wang, W., Luo, Z., Ma, L., and Song, D. Codehalu: Investigating code hallucinations in llms via execution-based verification. In Proceedings of the AAAI Conference on Artificial Intelligence, pp. 25300–25308, 2025.
  49. 49.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. Llama: Open and efficient foundation language models, 2023.
  50. 50.Vovk, V., Gammerman, A., and Shafer, G. Algorithmic learning in a random world. Springer Science & Business Media, 2005.
  51. 51.Williams, A., Nangia, N., and Bowman, S. R. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of NAACL-HLT, pp. 1112–1122, 2018.
  52. 52.Zalewski, M. Technical “whitepaper” for afl-fuzz. URl: http://lcamtuf.coredump.cx/afl/technical_details.txt, 2014.
  53. 53.Zhang, Z., Wang, Y., Wang, C., Chen, J., and Zheng, Z. Llm hallucinations in practical code generation: Phenomena, mechanism, and mitigation, 2025. URL https://arxiv.org/abs/2409.20550.
  54. 54.Zhong, L., Wang, Z., and Shang, J. Debug like a human: A large language model debugger via verifying runtime execution step by step. In Findings of the Association for Computational Linguistics ACL 2024, pp. 851–870, 2024.

Citation

MLA
Jeong, J., et al. “Towards Functional Correctness of Large Code Models with Selective Generation”. arXiv, 2025, http://arxiv.org/abs/2505.13553v3.
APA
Jeong, J., Kim, T., & Park, S. (2025). Towards Functional Correctness of Large Code Models with Selective Generation. arXiv. http://arxiv.org/abs/2505.13553v3
Chicago
Jeong, J., T. Kim, and S. Park. 2025. “Towards Functional Correctness of Large Code Models with Selective Generation”. arXiv. http://arxiv.org/abs/2505.13553v3.
Harvard
Jeong, J., Kim, T. and Park, S. (2025) “Towards Functional Correctness of Large Code Models with Selective Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2505.13553v3.
Vancouver
1. Jeong J, Kim T, Park S (2025) Towards Functional Correctness of Large Code Models with Selective Generation. arXiv

BibTeX

@article{jeong2025towards,
  title = {Towards Functional Correctness of Large Code Models with Selective Generation},
  author = {Jeong, Jaewoo and Kim, Taesoo and Park, Sangdon},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2505.13553v3},
  eprint = {2505.13553}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/