Designing Quantum Error Correcting Codes to fit decoders via Reinforcement Learning

Omer S. SellaRobert PinslerThomas Heinis

article2026arXiv1 citations

Develops a reinforcement learning framework using Proximal Policy Optimization to generate Bivariate Bicycle quantum error-correcting codes tailored specifically to maximize the performance of chosen decoder architectures.

Listen

Practical quantum computing requires quantum error correction to suppress physical noise faster and more reliably than errors accumulate. Traditional approaches typically design quantum error-correcting codes and decoding algorithms separately. However, heuristic decoders depend heavily on specific code structures, and poorly matched combinations lead to excessive decoding latency or high logical error rates. The article evaluates whether reinforcement learning can co-design quantum codes to directly optimize the performance of a fixed decoder architecture.

To achieve this, the authors formulate the design of Bivariate Bicycle codes as a sequential decision-making process. The system uses Proximal Policy Optimization to train an agent that modifies code polynomials by flipping individual coefficients. The reward function combines two objectives: penalizing codes that fail to meet a required minimum number of logical qubits and integrating the area under the physical-to-logical error rate curve across simulated depolarizing noise channels. To improve sampling efficiency, the authors incorporate a transformer-based encoder pretrained on a dataset of over 39 million code evaluation records across various code dimensions.

The findings demonstrate that reinforcement learning successfully navigates the code design space to discover high-performing quantum codes tailored to the chosen decoder. Across multiple training runs, the agent progressively shifts from generating invalid or low-capacity codes to consistently discovering codes that match or exceed published benchmark performance, reaching reference reward levels such as 0.546 for specific code sizes. The analysis also reveals an underlying tension in transferability: while the pretrained encoder accurately captures logical qubit counts and error curves on seen dimensions, attempts to extrapolate both metrics simultaneously to larger, unseen code sizes showed negative rank correlation.

These results show that automated, decoder-aware code generation is viable and can replace laborious manual code construction. Tailoring codes to specific decoding algorithms and hardware constraints reduces data corruption risks and operational latencies, which is essential for scaling fault-tolerant architectures. The authors recommend expanding this framework to incorporate circuit-level noise simulations, applying execution budgets to limit decoding costs during training, and exploring simultaneous optimization of both code structures and parameterized decoder settings.

Confidence in the findings is supported by extensive empirical evaluations across 300 encoder models and multi-seed training runs. Nevertheless, the study remains limited by its reliance on a simplified, symmetric depolarizing noise model rather than full hardware-level circuit noise, and the current framework operates on central processing units rather than fully integrated real-time quantum control hardware.

arXiv: 2608.15754
  • Paper: Proximal Policy Optimization Algorithms, John Schulman et al. (2017). This chapter introduces Proximal Policy Optimization (PPO), which serves as the core reinforcement learning algorithm adapted by the source paper to optimize quantum code stabilizer sets.
  • Paper: Quantum Anticodes, Chunjun Cao et al. (2025). This paper establishes foundational algebraic and structural perspectives on quantum error-correcting codes, providing essential domain context for analyzing and designing stabilizer codes.
  • Paper: Trust Region Policy Optimization, John Schulman et al. (2015). This work develops Trust Region Policy Optimization, establishing the theoretical policy-constraint foundations that directly motivated PPO and stable policy-gradient optimization.
  • Paper: Parameterized quantum circuits as machine learning models, Marcello Benedetti et al. (2019). This paper surveys parameterized quantum circuits as trainable machine learning models, offering relevant background on integrating learning algorithms with quantum state and circuit optimization.

No sufficiently relevant recommendations were found.

Cover for Designing Quantum Error Correcting Codes to fit decoders via Reinforcement Learning

Abstract

We present a reinforcement learning (RL) approach to the co-design of stabilizer sets of Quantum Error Correcting Codes (QECCs) and decoders. We show how to produce a generative model that produces Bivariate Bicycle (BB) codes based on the choice of decoder. Specifically, we fix a decoder architecture and use Proximal Policy Optimisation (PPO) to train an agent over BB codes to maximise decoder performance under a depolarising channel noise model.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Qubits and Pauli operators
  • 2.2 Bivariate Bicycle Codes
  • 2.3 Proximal Policy Optimisation
  • 2.4 Dataset
  • 3 Framing as a sequential decision making problem
  • 3.1 Reward
  • 3.2 Framing as a Markov Decision Process
  • 4 Policy Architecture
  • 4.1 Multi-Layer-Perceptron (MLP)
  • 4.2 BB code construction aware encoder
  • 4.3 The logical-qubit features
  • 4.4 The fusion layer
  • 4.5 The policy head and the value head
  • 4.6 Pretraining the encoder
  • 4.6.1 Evaluation (of the encoder)
  • 5 Related Work
  • 6 Evaluation and results
  • 7 Discussion and further work
  • 7.1 Immediate further work
  • 7.2 Wider application
  • 7.3 Reproducibility
  • A Dataset
  • References

Knowls

  1. Knowl 1 — Markov Decision Process Formulation for Bivariate Bicycle Code Design

    model/method

    The task of designing Bivariate Bicycle (BB) quantum Low-Density Parity Check (qLDPC) codes tailored to a specific decoder is formulated as a discrete Markov Decision Process (MDP):

    • State Space: A state s∈Ss \in \mathcal{S} is defined by the binary parity check matrix pair HX=[A∣B]H_X = [A \mid B] and HZ=[B⊤∣A⊤]H_Z = [B^\top \mid A^\top], where A,B∈F2ℓm×ℓmA, B \in \mathbb{F}_2^{\ell m \times \ell m} are circulant block matrices constructed from four univariate polynomials a(x),b(x),a(y),b(y)∈F2[x,y]a(x), b(x), a(y), b(y) \in \mathbb{F}_2[x, y]. The cyclic shift generators satisfy X=Sℓ⊗ImX = S_\ell \otimes I_m and Y=Iℓ⊗SmY = I_\ell \otimes S_m, giving A=a(X)⊕a(Y)A = a(X) \oplus a(Y) and B=b(X)⊕b(Y)B = b(X) \oplus b(Y), with total physical qubit block length n=2ℓmn = 2\ell m.
    • Action Space: An action a=(Aa(x),Ab(x),Aa(y),Ab(y))a = (A^{a(x)}, A^{b(x)}, A^{a(y)}, A^{b(y)}) is a 4-tuple of discrete operations. Each component ApA^{p} either selects at most one monomial coefficient ci∈{0,1}c_i \in \{0, 1\} of polynomial pp to flip via ci←ci⊕1c_i \leftarrow c_i \oplus 1, or executes a no-op.
    • Transition Function: State transitions are deterministic: applying action aa updates the polynomial coefficients and reconstructs the modified check matrices HX′,HZ′H_X', H_Z'.
    • Episode Horizon: Episodes initialize with random polynomials containing at most three non-zero coefficients and terminate via environment truncation after a horizon of T∈O(ℓ+m)T \in \mathcal{O}(\ell + m) steps (typically T=30T = 30).
  2. Knowl 2 — Logical Qubit Thresholded Physical-to-Logical Error Reward Function

    equation

    For a Calderbank-Steane-Shor (CSS) quantum parity check matrix pair (HX,HZ)(H_X, H_Z) with block length n=2ℓmn = 2\ell m and logical dimension k=n−rank(HX)−rank(HZ)k = n - \text{rank}(H_X) - \text{rank}(H_Z), the scalar reward r(HX,HZ)r(H_X, H_Z) over a physical depolarizing error rate interval [pmin⁡,pmax⁡][p_{\min}, p_{\max}] is defined as:

    r(HX,HZ)={1pmax⁡−pmin⁡∫pmin⁡pmax⁡(1−f^(p;H)) dpif k≥kmin⁡,k−kmin⁡kmin⁡if k<kmin⁡r(H_X, H_Z) = \begin{cases} \displaystyle \frac{1}{p_{\max} - p_{\min}} \int_{p_{\min}}^{p_{\max}} (1 - \hat{f}(p; H)) \, dp & \text{if } k \ge k_{\min}, \\ \displaystyle \frac{k - k_{\min}}{k_{\min}} & \text{if } k < k_{\min} \end{cases}

    where:

    • kmin⁡∈Z+k_{\min} \in \mathbb{Z}^+ is a prescribed minimum threshold for the number of encoded logical qubits.
    • f^(p;H)∈[0,1]\hat{f}(p; H) \in [0, 1] is the empirical combined failure rate (including decoder failures and undetected logical errors) computed via Monte Carlo decoding trials at physical error rate pp.
    • The integral is evaluated numerically via the trapezoid rule across a discrete grid of physical error rates spanning [pmin⁡,pmax⁡][p_{\min}, p_{\max}].
    • When k<kmin⁡k < k_{\min}, expensive decoder simulations are bypassed, and a deterministic linear penalty in [−1,0)[-1, 0) proportional to the logical qubit deficit is returned.
  3. Knowl 3 — Reinforcement Learning Algorithm for Decoder-Tailored Quantum Code Co-Design

    algorithm

    The algorithm trains a generative policy using Proximal Policy Optimization (PPO) to discover Bivariate Bicycle quantum error correcting codes that maximize decoding performance under a specified syndrome decoder.

    Input: Decoder function decode, target logical dimension kmin⁡k_{\min}, batch size FF, truncation horizon TT, error grid P=[pmin⁡,…,pmax⁡]P = [p_{\min}, \dots, p_{\max}]
    Output: Best parity check matrices HX∗,HZ∗H_X^*, H_Z^* and trained policy πθ\pi_\theta
    Initialize policy network πθ\pi_\theta and value network VϕV_\phi
    Initialize best reward r∗←−∞r^* \leftarrow -\infty, best matrices H∗←nullH^* \leftarrow \text{null}
    for each training iteration do
        Initialize batch transition buffer B←∅B \leftarrow \emptyset
        while ∣B∣<F|B| < F do
            Sample initial polynomials a(x),a(y),b(x),b(y)a(x), a(y), b(x), b(y) with at most 3 non-zero terms
            Construct initial state s=(HX,HZ)s = (H_X, H_Z)
            for t=1t = 1 to TT do
                Sample action a=(Aa(x),Ab(x),Aa(y),Ab(y))∼πθ(⋅∣s)a = (A^{a(x)}, A^{b(x)}, A^{a(y)}, A^{b(y)}) \sim \pi_\theta(\cdot \mid s)
                Apply action aa to update polynomials and construct candidate matrices H′=(HX′,HZ′)H' = (H_X', H_Z')
                Compute logical dimension k←2ℓm−rank(HX′)−rank(HZ′)k \leftarrow 2\ell m - \text{rank}(H_X') - \text{rank}(H_Z')
                if k≥kmin⁡k \ge k_{\min} then
                    Evaluate failure curve f^(p;H′)\hat{f}(p; H') across p∈Pp \in P via Monte Carlo decoding with decode
                    Compute reward r←1pmax⁡−pmin⁡∫pmin⁡pmax⁡(1−f^(p;H′)) dpr \leftarrow \frac{1}{p_{\max}-p_{\min}} \int_{p_{\min}}^{p_{\max}} (1 - \hat{f}(p; H')) \, dp
                else
                    Compute penalty reward r←(k−kmin⁡)/kmin⁡r \leftarrow (k - k_{\min}) / k_{\min}
                end if
                if r>r∗r > r^* then
                    r∗←rr^* \leftarrow r
                    H∗←H′H^* \leftarrow H'
                end if
                Store transition (s,a,r,H′,truncated=(t==T))(s, a, r, H', \text{truncated} = (t == T)) in BB
                s←H′s \leftarrow H'
            end for
        end while
        Update policy parameters θ\theta and critic parameters ϕ\phi using PPO clipped surrogate objective on BB
    end for
    return H∗,πθH^*, \pi_\theta
  4. Knowl 4 — Monte Carlo Logical Error Rate Evaluation for Quantum Parity Check Codes

    algorithm

    To evaluate candidate quantum parity check codes (HX,HZ)(H_X, H_Z) under depolarizing noise, Monte Carlo sampling computes empirical logical error rates against a syndrome decoding routine.

    Input: Parity check matrices HX∈F2mX×n,HZ∈F2mZ×nH_X \in \mathbb{F}_2^{m_X \times n}, H_Z \in \mathbb{F}_2^{m_Z \times n}, error range PP, number of samples per point NsN_s, decoder decode
    Output: Logical error rate estimate f^(p;H)\hat{f}(p; H) for each p∈Pp \in P
    Compute XX-type and ZZ-type logical operator generator matrices LX,LZL_X, L_Z from HX,HZH_X, H_Z
    Initialize output array f^\hat{f} of length ∣P∣|P|
    for each error rate index j=1j = 1 to ∣P∣|P| with error rate p=P[j]p = P[j] do
        Initialize error counter c←0c \leftarrow 0
        for i=1i = 1 to NsN_s do
            Sample Pauli error E=(E1,…,En)E = (E_1, \dots, E_n) with depolarizing rate pp
            Extract binary representations PX(E),PZ(E)∈F2nP_X(E), P_Z(E) \in \mathbb{F}_2^n
            Compute syndromes sX←PZ(E)⋅HX⊤s_X \leftarrow P_Z(E) \cdot H_X^\top and sZ←PX(E)⋅HZ⊤s_Z \leftarrow P_X(E) \cdot H_Z^\top
            Execute decoder (E^X,E^Z)←decode(HX,HZ,sX,sZ)(\hat{E}_X, \hat{E}_Z) \leftarrow \text{decode}(H_X, H_Z, s_X, s_Z)
            Compute residual errors RX←(PX(E)⊕E^X)R_X \leftarrow (P_X(E) \oplus \hat{E}_X) and RZ←(PZ(E)⊕E^Z)R_Z \leftarrow (P_Z(E) \oplus \hat{E}_Z)
            if HX⋅RZ⊤≠0H_X \cdot R_Z^\top \neq 0 or HZ⋅RX⊤≠0H_Z \cdot R_X^\top \neq 0 then
                c←c+1c \leftarrow c + 1
            else if LX⋅RZ⊤≠0L_X \cdot R_Z^\top \neq 0 or LZ⋅RX⊤≠0L_Z \cdot R_X^\top \neq 0 then
                c←c+1c \leftarrow c + 1
            end if
        end for
        f^[j]←c/Ns\hat{f}[j] \leftarrow c / N_s
    end for
    return f^\hat{f}
  5. Knowl 5 — Multi-Branch Neural Network Policy and Critic Architecture

    model/method

    The reinforcement learning framework parametrizes policy πθ(a∣s)\pi_\theta(a \mid s) and value estimator Vϕ(s)V_\phi(s) via separate instances of a three-branch neural network:

    1. Dense Matrix MLP Branch: Parity submatrices A,B∈{0,1}ℓm×ℓmA, B \in \{0, 1\}^{\ell m \times \ell m} are converted to bipolar values {−1,+1}\{-1, +1\} and concatenated into a 2(ℓm)22(\ell m)^2-dimensional flat vector. A three-layer MLP with tanh⁡\tanh activations projects this input to a 256-dimensional representation.
    2. Transferable Transformer Encoder Branch: Polynomial coefficients from a(x),a(y),b(x),b(y)a(x), a(y), b(x), b(y) are converted into 2ℓ+2m2\ell + 2m tokens of width 9+2H9 + 2H (where HH is the number of Fourier harmonics). Tokens are projected to dimension d=64d = 64, passed through 2 Transformer encoder layers (d=64d = 64, 4 heads, feed-forward dimension 128, GELU activations), and pooled via attention pooling to a 64-dimensional vector.
    3. Logical Invariant Branch: A 3-dimensional vector [min⁡(k−kmin⁡,0)/kmin⁡, 1k≥kmin⁡, log⁡(1+k)][ \min(k - k_{\min}, 0)/k_{\min}, \, \mathbf{1}_{k \ge k_{\min}}, \, \log(1 + k)] directly provides rank-derived logical dimension information.
    4. Fusion and Heads: The branches are concatenated (64+256+3=32364 + 256 + 3 = 323) and passed through a dense tanh⁡\tanh layer to produce a 256-dimensional fused state vector.
      • Policy Head: Four independent linear categorical heads output logits of sizes ℓ+1,ℓ+1,m+1,m+1\ell+1, \ell+1, m+1, m+1 corresponding to monomial flips plus a no-op for each polynomial. The joint action log-probability is the sum of the four categorical log-probabilities.
      • Value Head: A single linear layer maps the 256-dimensional fused vector to a scalar state-value estimate V^(s)\hat{V}(s).
  6. Knowl 6 — Cyclic Harmonic Positional Tokenization for Bivariate Polynomials

    model/method

    To represent Bivariate Bicycle code polynomials a(x),b(x)∈F2[x]/⟨xℓ−1⟩a(x), b(x) \in \mathbb{F}_2[x]/\langle x^\ell - 1 \rangle and a(y),b(y)∈F2[y]/⟨ym−1⟩a(y), b(y) \in \mathbb{F}_2[y]/\langle y^m - 1 \rangle in a size-transferable format, each monomial coefficient is mapped to a feature token. A code with parameters (ℓ,m)(\ell, m) produces a sequence of 2ℓ+2m2\ell + 2m tokens ordered as [a(x)∣a(y)∣b(x)∣b(y)][a(x) \mid a(y) \mid b(x) \mid b(y)].

    For monomial index k∈{0,…,p−1}k \in \{0, \dots, p-1\} in a polynomial with cyclic period p∈{ℓ,m}p \in \{\ell, m\}, the token feature vector has fixed width F=9+2HF = 9 + 2H regardless of ℓ,m\ell, m:

    1. Coefficient Bit (width 1): Binary indicator I(ck=1)\mathbb{I}(c_k = 1) stating if monomial term xkx^k or yky^k is non-zero.
    2. Polynomial Identifier (width 4): One-hot categorical indicator for a(x),a(y),b(x),a(x), a(y), b(x), or b(y)b(y).
    3. Cyclic Fourier Harmonics (width 2H2H): Coordinate pairs (sin⁡(hθk),cos⁡(hθk))(\sin(h \theta_k), \cos(h \theta_k)) for integer harmonic orders h∈{1,…,H}h \in \{1, \dots, H\} at unit circle angle θk=2πk/p\theta_k = 2\pi k / p. Setting H≥⌊p/2⌋H \ge \lfloor p/2 \rfloor provides a complete Fourier basis for positions on the period-pp cycle.
    4. Linear Position (width 1): Normalized position scalar k/p∈[0,1)k / p \in [0, 1).
    5. Global Dimension Features (width 3): Global parameters (log⁡(ℓ),log⁡(m),1.0)(\log(\ell), \log(m), 1.0) shared across all tokens.
  7. Knowl 7 — Multi-Task Pretraining Loss for Code Performance Surrogate Models

    equation

    To warm-start the policy and value networks before reinforcement learning, the Transformer encoder trunk is pretrained as a surrogate model to predict both the Monte Carlo physical-to-logical error curve and the logical dimension kk. The multi-task pretraining objective is:

    L=1M∑i=1MNs⋅BCE(p^i,ciNs)+λk(log⁡(1+k^)−log⁡(1+k))2\mathcal{L} = \frac{1}{M} \sum_{i=1}^M N_s \cdot \text{BCE}\left(\hat{p}_i, \frac{c_i}{N_s}\right) + \lambda_k \left( \log(1 + \hat{k}) - \log(1 + k) \right)^2

    where:

    • MM is the number of evaluated physical error rates on the grid (e.g., M=5M = 5).
    • NsN_s is the number of Monte Carlo decoding trials per grid point (e.g., Ns=50N_s = 50).
    • ci∈{0,…,Ns}c_i \in \{0, \dots, N_s\} is the observed count of logical decoding failures at error rate pip_i.
    • p^i∈[0,1]\hat{p}_i \in [0, 1] is the predicted failure rate output by a 5-dimensional sigmoid head (64→64→564 \to 64 \to 5).
    • BCE(p,y)=−ylog⁡(p)−(1−y)log⁡(1−p)\text{BCE}(p, y) = -y \log(p) - (1 - y) \log(1 - p) is binary cross-entropy (binomial negative log-likelihood).
    • k∈Nk \in \mathbb{N} is the exact logical dimension computed via Gaussian elimination, k^\hat{k} is the predicted logical dimension from a scalar regression head (64→64→164 \to 64 \to 1), and λk>0\lambda_k > 0 is a loss weighting coefficient.

    After pretraining, both prediction heads are discarded, and the encoder trunk weights are transferred to initialize the RL policy and critic networks.

  8. Knowl 8 — Tension and Negative Rank Correlation in Zero-Shot Encoder Transfer

    empirical result

    Surrogate Transformer encoders pretrained on small code dimensions (such as (ℓ,m)=(6,6)(\ell, m) = (6, 6)) exhibit a fundamental tension when transferring zero-shot to out-of-distribution, unseen code dimensions (such as (ℓ,m)∈{(5,15),(3,27),(21,18)}(\ell, m) \in \{(5, 15), (3, 27), (21, 18)\}).

    Across 300 trained encoder instances spanning varying harmonic orders H∈{3,4,6,7,10}H \in \{3, 4, 6, 7, 10\} and random initializations, the Spearman rank correlation between the logical qubit prediction error (kk-MAE) and the error curve prediction error (Curve-MAE) becomes consistently negative on unseen code sizes, with values falling between −0.10-0.10 and −0.62-0.62. This negative correlation demonstrates that encoder representations fine-tuned or grounded solely to predict algebraic structural properties (such as logical dimension kk) fail to simultaneously extrapolate decoder-dependent error rate curves to larger code families.

  9. Knowl 9 — Empirical Policy Convergence and Reward Ascent in BB Code Optimization

    empirical result

    When training a PPO agent over Bivariate Bicycle codes evaluated with a BP+OSD decoder, policy evaluation metrics demonstrate robust convergence from random states to high-performing codes:

    • In early training batches, deterministic policy evaluation trajectories spend the vast majority of steps in the negative penalty region (k<kmin⁡k < k_{\min}) due to initializing from sparse random polynomials.
    • Across training iterations (evaluated every 10 collector batches), the 80% reward percentile band contracts and lifts entirely out of the penalty region toward maximum attainable rewards.
    • For [[108,8,10]][[108, 8, 10]] codes (l=9,m=6l=9, m=6 with reference reward 0.5460.546), the rolling median within-episode reward gain—the difference between the initial reset state reward and the final state reward—rises monotonically from ≈0.0\approx 0.0 to over 1.01.0 across 28 independent training runs before plateauing near collector batch 50.
    • In mature policy stages, best-of-evaluation codes consistently match or surpass the rewards achieved by published reference codes.
  10. Knowl 10 — Bivariate Bicycle Code Benchmark Specifications and Reference Rewards

    data/table

    The reinforcement learning environment is evaluated across seven Bivariate Bicycle code configurations parameterized by cyclic dimensions (ℓ,m)(\ell, m). The physical qubit block size is n=2ℓmn = 2\ell m, generator matrices A,BA, B have dimension ℓm×ℓm\ell m \times \ell m, and reference rewards are computed using width-normalized trapezoidal integration over a 5-point geometric physical error grid p∈[0.001,0.1]p \in [0.001, 0.1] under BP+OSD decoding.

    ll mm n=2lmn = 2lm A,BA, B dimensions Reference code [[n,k,d]][[n, k, d]] kmin⁡k_{\min} Reference reward
    6 6 72 36×3636 \times 36 [[72,12,6]][[72, 12, 6]] 12 0.346
    15 3 90 45×4545 \times 45 [[90,8,10]][[90, 8, 10]] 8 0.506
    9 6 108 54×5454 \times 54 [[108,8,10]][[108, 8, 10]] 8 0.546
    12 6 144 72×7272 \times 72 [[144,12,12]][[144, 12, 12]] 12 0.554
    5 15 150 75×7575 \times 75 — — —
    3 27 162 81×8181 \times 81 — — —
    21 18 756 378×378378 \times 378 [[756,16,≤34]][[756, 16, \le 34]] 16 0.600

    The threshold parameter kmin⁡k_{\min} dictates the minimum number of logical qubits required before non-penalty decoder simulation rewards are computed. Reference codes from prior literature provide performance baselines for normalized reward comparison.

Coverage note — Omitted the summary table of code-evaluation sample counts across grid types (Table 5) and empirical coefficient distribution heatmaps (Appendix A) as they document offline dataset sizing and logging histograms rather than core algorithmic or theoretical contributions.

References

  1. 1.Pavel Panteleev and Gleb Kalachev. Asymptotically good quantum and locally testable classical ldpc codes. In Proceedings of the 54th Annual ACM SIGACT Symposium on Theory of Computing (STOC), pages 375–388, 2022. DOI: 10.1145/3519935.3520017. URL https://doi.org/10.1145/3519935.3520017.
  2. 2.Sergey Bravyi, Andrew W Cross, Jay M Gambetta, Dmitri Maslov, Patrick Rall, and Theodore J Yoder. High-threshold and low-overhead fault-tolerant quantum memory. Nature, 627(8005):778–782, 2024. DOI: 10.1038/s41586-024-07107-7. URL https://doi.org/10.1038/s41586-024-07107-7.
  3. 3.Judea Pearl. Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference. Morgan Kaufmann, 1988.
  4. 4.Omer S Sella, Andrew W Moore, and Noa Zilberman. Fec killed the cut-through switch. In Proceedings of the 2018 Workshop on Networking for Emerging Applications and Technologies, pages 15–20, 2018. DOI: 10.1145/3229574.3229577.
  5. 5.R Michael Tanner. A recursive approach to low complexity codes. IEEE Transactions on Information Theory, 27(5):533–547, 1981. DOI: 10.1109/tit.1981.1056404. URL https://doi.org/10.1109/TIT.1981.1056404.
  6. 6.Marc PC Fossorier and Shu Lin. Soft-decision decoding of linear block codes based on ordered statistics. IEEE Transactions on Information Theory, 41(5):1379–1396, 1995. DOI: 10.1109/isit.1994.394624. URL https://doi.org/10.1109/18.412683.
  7. 7.Daniel Gottesman. Surviving as a quantum computer in a classical world. Textbook manuscript preprint, 2026. URL https://www.cs.umd.edu/~dgottesm/QECCbook-2026.pdf.
  8. 8.A Robert Calderbank and Peter W Shor. Good quantum error-correcting codes exist. Physical Review A, 54(2):1098, 1996. DOI: 10.1103/physreva.54.1098. URL https://doi.org/10.1103/PhysRevA.54.1098.
  9. 9.Andrew M Steane. Error correcting codes in quantum theory. Physical Review Letters, 77(5):793, 1996. DOI: 10.1103/PhysRevLett.77.793. URL https://doi.org/10.1103/PhysRevLett.77.793.
  10. 10.David J. C. MacKay, Graeme Mitchison, and Paul L. McFadden. Sparse-graph codes for quantum error correction. IEEE Transactions on Information Theory, 50(10): 2315–2330, 2004. DOI: 10.1109/TIT.2004.834737.
  11. 11.Lukas Voss, Sim Jian Xian, Tobias Haug, and Kishor Bharti. Multivariate bicycle codes, 2024.
  12. 12.Lukas Voss, Sim Jian Xian, Tobias Haug, and Kishor Bharti. Multivariate bicycle codes. Physical Review A, 111(6):L060401, 2025. DOI: 10.1103/ll5p-z88p.
  13. 13.Pavel Panteleev and Gleb Kalachev. Degenerate quantum ldpc codes with good finite length performance. Quantum, 5:585, 2021. DOI: 10.22331/q-2021-11-22-585. URL https://doi.org/10.22331/q-2021-11-22-585.
  14. 14.Ming Wang and Frank Mueller. Coprime bivariate bicycle codes. arXiv preprint arXiv:2408.10001, 2024.
  15. 15.Jens Niklas Eberhardt and Vincent Steffan. Logical operators and fold-transversal gates of bivariate bicycle codes. IEEE Transactions on Information Theory, 71(2): 1140–1152, 2024. DOI: 10.48550/arXiv.2407.03973.
  16. 16.Alexey A. Kovalev and Leonid P. Pryadko. Quantum kronecker sum-product low-density parity-check codes with finite rate. Physical Review A, 88(1):012311, 2013. DOI: 10.1103/PhysRevA.88.012311.
  17. 17.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017. DOI: 10.48550/arXiv.1707.06347. URL https://doi.org/10.48550/arXiv.1707.06347.
  18. 18.Albert Bou, Matteo Bettini, Sebastian Dittert, Vikash Kumar, Shagun Sodhani, Xiaomeng Yang, Gianni De Fabritiis, and Vincent Moens. TorchRL: A data-driven decision-making library for PyTorch. In International Conference on Learning Representations (ICLR), 2024. DOI: 10.48550/arXiv.2306.00577.
  19. 19.Omer S. Sella. bbCodesDataset: decoder-evaluation data for bivariate bicycle codes, 2026. URL https://github.com/Omer-Sella/bbCodesDataset. Dataset and figure-reproduction scripts.
  20. 20.Mark Towers, Ariel Kwiatkowski, Jordan Terry, John U. Balis, Gianluca De Cola, Tristan Deleu, Manuel Goulão, Andreas Kallinteris, Markus Krimmel, Arjun KG, et al. Gymnasium: A standard interface for reinforcement learning environments. arXiv preprint arXiv:2407.17032, 2024. DOI: 10.48550/arXiv.2407.17032.
  21. 21.Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. MIT Press, 2 edition, 2018.
  22. 22.Joschka Roffe, David R. White, Simon Burton, and Earl Campbell. Decoding across the quantum low-density parity-check code landscape. Physical Review Research, 2 (4), Dec 2020. ISSN 2643-1564. DOI: 10.1103/physrevresearch.2.043423. URL http://dx.doi.org/10.1103/PhysRevResearch.2.043423.
  23. 23.Omer Shimon Sella and Thomas Heinis. A mapping of the min-sum decoder to reduction operations, and its implementation using cuda kernels. arXiv preprint arXiv:2507.10424, 2025. DOI: 10.48550/arXiv.2507.10424.
  24. 24.Ming Wang and Frank Mueller. Coprime bivariate bicycle codes and their layouts on cold atoms. Quantum, 10:2009, 2026. DOI: 10.22331/q-2026-02-23-2009.
  25. 25.Shai Evra, Tali Kaufman, and Gilles Zémor. Decodable quantum ldpc codes beyond the n\sqrt{n} distance barrier using high-dimensional expanders. SIAM Journal on Computing, 53(6):FOCS20–276, 2022. DOI: 10.1137/20M1383689. URL https://doi.org/10.1137/20M1383689.
  26. 26.Naren Manjunath, Vieri Mattei, Apoorv Tiwari, and Tyler D Ellison. Universal quantum computation with group surface codes. arXiv preprint arXiv:2603.05502, 2026. DOI: 10.48550/arXiv.2603.05502. URL https://doi.org/10.48550/arXiv.2603.05502.
  27. 27.Ryan Sweke, Markus S Kesselring, Evert PL van Nieuwenburg, and Jens Eisert. Reinforcement learning decoders for fault-tolerant quantum computation. Machine Learning: Science and Technology, 2(2):025005, 2020. DOI: 10.1088/2632-2153/abc609. URL https://doi.org/10.1088/2632-2153/abc609.
  28. 28.John Blue, Harshil Avlani, Zhiyang He, Liu Ziyin, and Isaac L Chuang. Machine learning decoding of circuit-level noise for bivariate bicycle codes. Quantum, 10:2149, 2026. DOI: 10.22331/q-2026-06-30-2149.
  29. 29.Hendrik Poulsen Nautrup, Nicolas Delfosse, Vedran Dunjko, Hans J Briegel, and Nicolai Friis. Optimizing quantum error correction codes with reinforcement learning. Quantum, 3:215, 2019. DOI: 10.22331/q-2019-12-16-215. URL https://doi.org/10.22331/q-2019-12-16-215.
  30. 30.Caroline Mauron, Terry Farrelly, and Thomas M Stace. Optimization of tensor network codes with reinforcement learning. New Journal of Physics, 26:023024, 2024. DOI: 10.1088/1367-2630/ad23a6. URL https://doi.org/10.1088/1367-2630/ad23a6.
  31. 31.Austin Yubo He and Zi-Wen Liu. Discovering highly efficient low-weight quantum error-correcting codes with reinforcement learning. arXiv preprint arXiv:2502.14372, 2025. DOI: 10.48550/arXiv.2502.14372. URL https://doi.org/10.48550/arXiv.2502.14372.
  32. 32.Juan Cruz-Benito, Andrew W Cross, David Kremer, and Ismael Faro. Evolutionary discovery of bivariate bicycle codes with LLM-guided search. arXiv preprint arXiv:2606.02418, 2026. DOI: 10.48550/arXiv.2606.02418. URL https://doi.org/10.48550/arXiv.2606.02418.
  33. 33.Yihua Chengyu, Richard Meister, Conor Carty, Sheng-Ku Lin, and Roberto Bondesan. Bayesian optimization for quantum error-correcting code discovery. arXiv preprint arXiv:2601.18562, 2026. DOI: 10.48550/arXiv.2601.18562. URL https://doi.org/10.48550/arXiv.2601.18562.
  34. 34.Vincent Paul Su, C Cao, Hong-Ye Hu, Yariv Yanay, Charles Tahan, and Brian Swingle. Discovery of optimal quantum error correcting codes via reinforcement learning (2023). arXiv preprint arXiv:2305.06378. DOI: 10.48550/arXiv.2305.06378.
  35. 35.ChunJun Cao and Brad Lackey. Quantum lego: Building quantum error correction codes from tensor networks. PRX Quantum, 3(2):020332, 2022. DOI: 10.1103/PRXQuantum.3.020332.
  36. 36.Omer Sella. Coding for emerging archival storage media. PhD thesis, University of Cambridge, 2024.
  37. 37.Fabio Pardo, Arash Tavakoli, Vitaly Levdik, and Petar Kormushev. Time limits in reinforcement learning. In Proceedings of the 35th International Conference on Machine Learning (ICML), volume 80 of PMLR, pages 4045–4054, 2018. DOI: 10.48550/arXiv.1712.00378.
  38. 38.Sham Kakade and John Langford. Approximately optimal approximate reinforcement learning. In Proceedings of the 19th International Conference on Machine Learning (ICML), pages 267–274, 2002.
  39. 39.Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O. Stanley, and Jeff Clune. First return, then explore. Nature, 590(7847):580–586, 2021. DOI: 10.1038/s41586-020-03157-9.
  40. 40.Austin G Fowler, Matteo Mariantoni, John M Martinis, and Andrew N Cleland. Surface codes: Towards practical large-scale quantum computation. Physical Review A, 86(3):032324, 2012. DOI: 10.1103/physreva.86.032324. URL https://doi.org/10.1103/PhysRevA.86.032324.
  41. 41.Daniel Litinski. A game of surface codes: Large-scale quantum computing with lattice surgery. Quantum, 3:128, 2019. DOI: 10.22331/q-2019-03-05-128. URL https://doi.org/10.22331/q-2019-03-05-128.
  42. 42.Tristan Müller, Thomas Alexander, Michael E Beverland, Markus Bühler, Blake R Johnson, Thilo Maurer, and Drew Vandeth. Improved belief propagation is sufficient for real-time decoding of quantum memory. arXiv preprint arXiv:2506.01779, 2025. DOI: 10.48550/arXiv.2506.01779. URL https://arxiv.org/pdf/2506.01779.
  43. 43.Omer S. Sella. bbCodeSurrogates: pretrained encoders for bivariate bicycle codes, 2026. URL https://github.com/Omer-Sella/bbCodeSurrogates. 300 encoder checkpoints (6 recipes × 5 harmonics × 10 seeds) and the transfer-evaluation census.

Citation

MLA
Sella, O. S., et al. “Designing Quantum Error Correcting Codes to Fit Decoders via Reinforcement Learning”. arXiv, 2026, https://doi.org/10.48550/arxiv.2608.15754.
APA
Sella, O. S., Pinsler, R., & Heinis, T. (2026). Designing Quantum Error Correcting Codes to fit decoders via Reinforcement Learning. arXiv. https://doi.org/10.48550/arxiv.2608.15754
Chicago
Sella, O. S., R. Pinsler, and T. Heinis. 2026. “Designing Quantum Error Correcting Codes to Fit Decoders via Reinforcement Learning”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2608.15754.
Harvard
Sella, O.S., Pinsler, R. and Heinis, T. (2026) “Designing Quantum Error Correcting Codes to fit decoders via Reinforcement Learning”. arXiv. Available at: https://doi.org/10.48550/arxiv.2608.15754.
Vancouver
1. Sella OS, Pinsler R, Heinis T (2026) Designing Quantum Error Correcting Codes to fit decoders via Reinforcement Learning. https://doi.org/10.48550/arxiv.2608.15754

BibTeX

@misc{https://doi.org/10.48550/arxiv.2608.15754,
  doi = {10.48550/ARXIV.2608.15754},
  url = {https://arxiv.org/abs/2608.15754},
  author = {Sella, Omer S. and Pinsler, Robert and Heinis, Thomas},
  keywords = {Quantum Physics (quant-ph), Information Theory (cs.IT), FOS: Physical sciences, FOS: Computer and information sciences},
  title = {Designing Quantum Error Correcting Codes to fit decoders via Reinforcement Learning},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/