Designing Quantum Error Correcting Codes to fit decoders via Reinforcement Learning
Omer S. SellaRobert PinslerThomas Heinis
Develops a reinforcement learning framework using Proximal Policy Optimization to generate Bivariate Bicycle quantum error-correcting codes tailored specifically to maximize the performance of chosen decoder architectures.
Practical quantum computing requires quantum error correction to suppress physical noise faster and more reliably than errors accumulate. Traditional approaches typically design quantum error-correcting codes and decoding algorithms separately. However, heuristic decoders depend heavily on specific code structures, and poorly matched combinations lead to excessive decoding latency or high logical error rates. The article evaluates whether reinforcement learning can co-design quantum codes to directly optimize the performance of a fixed decoder architecture.
To achieve this, the authors formulate the design of Bivariate Bicycle codes as a sequential decision-making process. The system uses Proximal Policy Optimization to train an agent that modifies code polynomials by flipping individual coefficients. The reward function combines two objectives: penalizing codes that fail to meet a required minimum number of logical qubits and integrating the area under the physical-to-logical error rate curve across simulated depolarizing noise channels. To improve sampling efficiency, the authors incorporate a transformer-based encoder pretrained on a dataset of over 39 million code evaluation records across various code dimensions.
The findings demonstrate that reinforcement learning successfully navigates the code design space to discover high-performing quantum codes tailored to the chosen decoder. Across multiple training runs, the agent progressively shifts from generating invalid or low-capacity codes to consistently discovering codes that match or exceed published benchmark performance, reaching reference reward levels such as 0.546 for specific code sizes. The analysis also reveals an underlying tension in transferability: while the pretrained encoder accurately captures logical qubit counts and error curves on seen dimensions, attempts to extrapolate both metrics simultaneously to larger, unseen code sizes showed negative rank correlation.
These results show that automated, decoder-aware code generation is viable and can replace laborious manual code construction. Tailoring codes to specific decoding algorithms and hardware constraints reduces data corruption risks and operational latencies, which is essential for scaling fault-tolerant architectures. The authors recommend expanding this framework to incorporate circuit-level noise simulations, applying execution budgets to limit decoding costs during training, and exploring simultaneous optimization of both code structures and parameterized decoder settings.
Confidence in the findings is supported by extensive empirical evaluations across 300 encoder models and multi-seed training runs. Nevertheless, the study remains limited by its reliance on a simplified, symmetric depolarizing noise model rather than full hardware-level circuit noise, and the current framework operates on central processing units rather than fully integrated real-time quantum control hardware.
- Paper: Proximal Policy Optimization Algorithms, John Schulman et al. (2017). This chapter introduces Proximal Policy Optimization (PPO), which serves as the core reinforcement learning algorithm adapted by the source paper to optimize quantum code stabilizer sets.
- Paper: Quantum Anticodes, Chunjun Cao et al. (2025). This paper establishes foundational algebraic and structural perspectives on quantum error-correcting codes, providing essential domain context for analyzing and designing stabilizer codes.
- Paper: Trust Region Policy Optimization, John Schulman et al. (2015). This work develops Trust Region Policy Optimization, establishing the theoretical policy-constraint foundations that directly motivated PPO and stable policy-gradient optimization.
- Paper: Parameterized quantum circuits as machine learning models, Marcello Benedetti et al. (2019). This paper surveys parameterized quantum circuits as trainable machine learning models, offering relevant background on integrating learning algorithms with quantum state and circuit optimization.
No sufficiently relevant recommendations were found.
