Jump-Start Reinforcement Learning

Ikechukwu UchenduTed XiaoYao LuBanghua ZhuMengyuan YanJoséphine SimonMatthew BenniceChuyuan FuCong MaJiantao Jiao

article2023ICML157 citations

Proposes a meta-algorithm that bootstraps value-based reinforcement learning using a prior guide-policy to initialize exploration trajectories, reducing sample complexity from exponential to polynomial and outperforming standard imitation and RL baselines.

Listen

Reinforcement learning enables autonomous systems to optimize performance through trial and error, but training policies from scratch is notoriously inefficient due to the difficulty of exploring complex environments with sparse rewards. While pre-existing policies, human demonstrations, or historical datasets can provide helpful starting points, naively using them to initialize modern value-based reinforcement learning algorithms often causes initial performance to collapse as value estimators struggle to adapt to new state distributions.

The article evaluates Jump-Start Reinforcement Learning, a meta-algorithm designed to bootstrap any reinforcement learning method using an existing, sub-optimal guide policy. Its objective is to demonstrate that systematically guiding an exploration policy can substantially reduce the sample complexity of reinforcement learning across standard benchmarks and complex continuous control tasks.

The approach operates by pairing a fixed guide policy with an actively learning exploration policy. At the start of training, the guide policy controls the agent for an initial sequence of steps to place it near promising states, after which the exploration policy takes over to complete the task. The authors implement two switching mechanisms: a backward curriculum that gradually reduces the guide policy's step count as performance improves, and a random switching schedule. They evaluate this framework across benchmark maze navigation and dexterous hand manipulation tasks, as well as simulated vision-based robotic grasping problems requiring continuous 3D control from raw pixel inputs.

The findings show that Jump-Start Reinforcement Learning significantly improves data efficiency and final task performance. On challenging robot manipulation tasks, the method learned effectively with as few as 20 demonstrations—a hundredfold reduction compared to the 2,000 demonstrations typically needed by baseline algorithms. In benchmark navigation tasks with limited data (10,000 transitions), the approach achieved success rates between 71% and 73%, whereas standard offline-to-online methods scored near 0% to 33%. Theoretical analysis confirms that using a guide policy improves sample complexity from exponential to polynomial in relation to the task horizon. Furthermore, policies pre-trained on simpler tasks successfully generalized to guide agents in more complex environments.

These results demonstrate that organizations can deploy reinforcement learning more rapidly and at lower operational cost by eliminating the need for massive initial demonstration datasets. Bootstrapping RL with simple heuristics or sub-optimal prior policies reduces training timelines and avoids the sample-inefficient exploration that often limits real-world robotic deployments.

For practical implementation, engineering teams should leverage available sub-optimal controllers or small demonstration datasets to construct guide policies rather than collecting exhaustive datasets upfront. While a structured curriculum offers superior sample efficiency during early training stages, a simple random switching strategy provides a viable, low-overhead alternative. Teams should conduct pilot testing on target workflows before full deployment, particularly to verify whether a cold start (initializing the exploration agent from scratch) or a warm start (copying prior network weights) best suits the dataset quality.

Confidence in these findings is strong for simulated environments and established benchmarks. However, stakeholders should note key limitations: the framework's effectiveness depends on the guide policy providing reasonable state space coverage. A severely biased or adversarial guide policy can constrain exploration and delay convergence, requiring domain experts to curate and validate initial guidance behaviors before training in safety-critical settings.

arXiv: 2204.02372
Cover for Jump-Start Reinforcement Learning

Abstract

Reinforcement learning (RL) provides a theoretical framework for continuously improving an agent's behavior via trial and error. However, efficiently learning policies from scratch can be very difficult, particularly for tasks that present exploration challenges. In such settings, it might be desirable to initialize RL with an existing policy, offline data, or demonstrations. However, naively performing such initialization in RL often works poorly, especially for value-based methods. In this paper, we present a meta algorithm that can use offline data, demonstrations, or a pre-existing policy to initialize an RL policy, and is compatible with any RL approach. In particular, we propose Jump-Start Reinforcement Learning (JSRL), an algorithm that employs two policies to solve tasks: a guide-policy, and an exploration-policy. By using the guide-policy to form a curriculum of starting states for the exploration-policy, we are able to efficiently improve performance on a set of simulated robotic tasks. We show via experiments that it is able to significantly outperform existing imitation and reinforcement learning algorithms, particularly in the small-data regime. In addition, we provide an upper bound on the sample complexity of JSRL and show that with the help of a guide-policy, one can improve the sample complexity for non-optimism exploration methods from exponential in horizon to polynomial.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 4. Jump-Start Reinforcement Learning
  • 4.1. Rolling In With Two Policies
  • 4.2. Algorithm
  • 4.3. Theoretical Analysis
  • 5. Experiments
  • 5.1. Comparison with IL+RL baselines
  • 5.2. Vision-Based Robotic Tasks
  • 5.3. Initial Dataset Sensitivity
  • 5.4. JSRL-Curriculum vs. JSRL-Random Switching
  • 5.5. Guide-Policy Generalization
  • 6. Conclusion
  • 7. Limitations
  • Acknowledgements
  • References
  • A. Appendix
  • A.1. Imitation and Reinforcement Learning (IL+RL)
  • A.1.1. D4RL
  • A.1.2. SIMULATED ROBOTIC GRASPING
  • A.2. Experiment Implementation Details
  • A.2.1. D4RL: ANT MAZE AND ADROIT
  • A.2.2. SIMULATED ROBOTIC MANIPULATION
  • A.3. Additional Experiments
  • A.4. Hyperparameters of JSRL
  • A.5. Theoretical Analysis for JSRL
  • A.5.1. SETUP AND NOTATIONS
  • A.5.2. PROOF SKETCH FOR THEOREM 4.1
  • A.5.3. UPPER BOUND OF JSRL
  • A.5.4. PROOF OF THEOREM A.3 AND COROLLARIES

Knowls

  1. Knowl 1 — Jump-Start Reinforcement Learning Framework

    model/method

    Jump-Start Reinforcement Learning (JSRL) is a meta-algorithm designed to bootstrap reinforcement learning (RL) agents—particularly value-based algorithms—using an existing prior policy, called the guide-policy πg(a∣s)\pi^g(a|s). JSRL addresses the initialization failure of value-based methods (which struggle when an actor is initialized with prior knowledge but the critic is uninitialized or poorly fitted) by deploying two policies sequentially in each episode.

    For an episodic Markov Decision Process (MDP) with total horizon HH:

    1. The guide-policy πg\pi^g executes for the first hh steps (h≤Hh \le H), guiding the agent from the initial state distribution to high-value or task-relevant state distributions.
    2. The exploration-policy πe\pi^e (the downstream RL policy being trained) takes control for the remaining H−hH - h steps.

    Over the course of training, the number of guide steps hh is gradually decreased from an initial value H1≤HH_1 \le H down to 0 according to a performance-driven curriculum (or sampled randomly via JSRL-Random). This reduces the state-distribution shift between πg\pi^g and πe\pi^e progressively: at each curriculum stage, πe\pi^e only needs to master the incremental sub-horizon necessary to reach the states visited in the preceding stage. Once h=0h = 0, πe\pi^e acts for the entire horizon independently.

  2. Knowl 2 — Jump-Start Reinforcement Learning Meta-Algorithm

    algorithm

    The JSRL meta-algorithm coordinates data collection and updates between the guide-policy πg\pi^g and the exploration-policy πe\pi^e across a descending sequence of guide horizons.

    Input: guide-policy πg\pi^g, performance threshold β\beta, task horizon HH, guide-step sequence (H1,H2,…,Hn)(H_1, H_2, \dots, H_n) where Hi∈{1,…,H}H_i \in \{1, \dots, H\}.
    Output: exploration-policy πe\pi^e and learned action-value function Q^\hat{Q}.
    Initialize exploration-policy πe\pi^e (either cold-started from scratch or warm-started as πe←πg\pi^e \leftarrow \pi^g), initialize Q-function Q^\hat{Q}, initialize replay buffer D←∅\mathcal{D} \leftarrow \emptyset.
    for each current guide-step h∈(H1,H2,…,Hn)h \in (H_1, H_2, \dots, H_n) do
        Define non-stationary composite policy π\pi by π1:h=πg\pi_{1:h} = \pi^g and πh+1:H=πe\pi_{h+1:H} = \pi^e
        repeat
            Roll out composite policy π\pi in the environment to obtain trajectory τ=((s1,a1,r1),…,(sH,aH,rH))\tau = ((s_1, a_1, r_1), \dots, (s_H, a_H, r_H))
            Append trajectory τ\tau to dataset D\mathcal{D}
            πe,Q^←TRAINPOLICY(πe,Q^,D)\pi^e, \hat{Q} \leftarrow \text{TRAINPOLICY}(\pi^e, \hat{Q}, \mathcal{D})
        until EVALUATEPOLICY(π)≥β\text{EVALUATEPOLICY}(\pi) \ge \beta
    end for
    return πe,Q^\pi^e, \hat{Q}

    Key execution details:

    • Initial guide-steps (H1H_1): Set to the average episode length required by πg\pi^g to solve the task (computed over prior rollouts).
    • Curriculum schedule: Guide-steps decrease in increments of H1/nH_1 / n, evaluating h∈{H1−H1n,H1−2H1n,…,0}h \in \{H_1 - \frac{H_1}{n}, H_1 - 2\frac{H_1}{n}, \dots, 0\}. In the random variant (JSRL-Random), hh is sampled uniformly from {H1,…,Hn}\{H_1, \dots, H_n\} in every episode.
    • Advancement threshold (β\beta): EVALUATEPOLICY(π)\text{EVALUATEPOLICY}(\pi) tracks a moving average of returns over recent evaluation intervals (e.g., 5 evaluations for benchmark tasks, 3 for robotic grasping) and advances when performance comes within a specified tolerance percentage of the historical best moving average.
    • Policy updates: TRAINPOLICY\text{TRAINPOLICY} updates πe\pi^e and Q^\hat{Q} via off-policy/actor-critic updates using transitions stored in D\mathcal{D} (e.g., Implicit Q-Learning or QT-Opt).
  3. Knowl 3 — State Feature Coverage Assumption for JSRL Guide-Policy

    assumption

    Let M=(S,A,P,R,p0,γ,H)\mathcal{M} = (\mathcal{S}, \mathcal{A}, P, R, p_0, \gamma, H) be an episodic Markov Decision Process. Let dhπ(s)d_h^{\pi}(s) denote the marginalized state visitation distribution at step h∈[H]h \in [H] under policy π\pi, and let π∗\pi^* denote an optimal policy.

    Assume that states are parameterized by a feature mapping ϕ:S→Rd\phi : \mathcal{S} \to \mathbb{R}^d such that for any policy π\pi, the Q-function Qπ(s,a)Q^\pi(s, a) and policy π(s)\pi(s) depend on state ss only through ϕ(s)\phi(s). The guide-policy πg\pi^g is assumed to cover the state features visited by the optimal policy:

    sup⁡s∈S,h∈[H]dhπ∗(ϕ(s))dhπg(ϕ(s))≤C\sup_{s \in \mathcal{S}, h \in [H]} \frac{d_h^{\pi^*}(\phi(s))}{d_h^{\pi^g}(\phi(s))} \le C

    for a finite constant C≥1C \ge 1.

    This assumption requires only that πg\pi^g visits the feature states encountered by the optimal policy with non-zero probability. It is strictly weaker than standard offline RL concentratability assumptions (which require coverage over all optimal state-action pairs (s,a)(s, a)), because πg\pi^g can select suboptimal actions at those states.

  4. Knowl 4 — Exponential Lower Bound for Non-Optimistic Exploration from Scratch

    theoretical result

    For finite-horizon episodic Markov Decision Processes without access to a guide-policy, non-optimism-based exploration algorithms that explore uniformly when estimated action values are zero (such as 0-initialized ϵ\epsilon-greedy and FALCON+) suffer from an expected sample complexity that is exponential in the horizon length HH.

    Specifically, on an HH-step combination lock MDP where a specific optimal action ah∗a_h^* at state sh∗s_h^* is required at each step h∈{1,…,H}h \in \{1, \dots, H\} to remain on the rewarding trajectory, any deviation transitions the agent to an absorbing unrewarded path. With zero-initialized values and zero intermediate rewards, uniform random exploration yields an arrival probability at the goal state sH∗s_H^* of (1/∣A∣)H=2−H(1/|\mathcal{A}|)^H = 2^{-H} (for ∣A∣=2|\mathcal{A}| = 2). Consequently, finding a policy with suboptimality Es0∼p0[V∗(s0)−Vπ(s0)]<0.5\mathbb{E}_{s_0 \sim p_0}[V^*(s_0) - V^\pi(s_0)] < 0.5 requires at least Ω(2H)\Omega(2^H) samples in expectation.

  5. Knowl 5 — Polynomial Suboptimality Upper Bound for JSRL

    theoretical result

    Under the guide-policy feature coverage condition sup⁡s,hdhπ∗(ϕ(s))dhπg(ϕ(s))≤C\sup_{s, h} \frac{d_h^{\pi^*}(\phi(s))}{d_h^{\pi^g}(\phi(s))} \le C and assuming a contextual bandit exploration oracle with regret bounded by ∑t=1TEs∼p0[r(s,π∗(s))−r(s,πt(s))]≤f(T,R)\sum_{t=1}^T \mathbb{E}_{s \sim p_0}[r(s, \pi^*(s)) - r(s, \pi^t(s))] \le f(T, R) for rewards bounded in [0,R][0, R], a backward curriculum JSRL algorithm executed over total budget TT samples satisfies:

    Es0∼p0[V0∗(s0)−V0π(s0)]≤C∑h=0H−1f(TH,H−h)\mathbb{E}_{s_0 \sim p_0}\left[V_0^*(s_0) - V_0^\pi(s_0)\right] \le C \sum_{h=0}^{H-1} f\left(\frac{T}{H}, H - h\right)

    Specialized sample complexity and regret rates:

    1. Tabular MDP with ϵ\epsilon-greedy oracle: Using regret bound f(T,R)=R(SA/T)1/3f(T, R) = R (SA / T)^{1/3}, the suboptimality is bounded by: O(CH7/3S1/3A1/3T−1/3)O\left(C H^{7/3} S^{1/3} A^{1/3} T^{-1/3}\right)
    2. Tabular MDP with FALCON+ oracle: Using regret bound f(T,R)=R(SA2/T)1/2f(T, R) = R (S A^2 / T)^{1/2}, the suboptimality is bounded by: O(CH5/2S1/2AT−1/2)O\left(C H^{5/2} S^{1/2} A T^{-1/2}\right)
    3. General Function Approximation with FALCON+ oracle: Given an offline regression oracle with mean squared error EF(n)\mathcal{E}_{\mathcal{F}}(n) on function class F\mathcal{F}, the suboptimality rate is: O~(C∑h=1HAEF(T/H))\tilde{O}\left(C \sum_{h=1}^H \sqrt{A \mathcal{E}_{\mathcal{F}}(T/H)}\right)

    Thus, JSRL provably reduces the exploration sample complexity of non-optimistic methods from exponential in HH to polynomial in HH without needing explicit uncertainty quantification bonuses.

  6. Knowl 6 — Experimental Benchmarks for JSRL Evaluation

    experimental setup

    JSRL is evaluated across two benchmark domains spanning continuous state-action spaces and high-dimensional pixel observations:

    1. D4RL Benchmarks:

      • Tasks: Ant Maze navigation (umaze, umaze-diverse, medium-play, medium-diverse, large-play, large-diverse) and Adroit dexterous manipulation (door-binary-v0, pen-binary-v0, relocate-binary-v0).
      • Data ablation: Initial offline dataset sizes are varied across 100, 1k, 10k, 100k, and 1m transitions.
      • Setup: Guide-policies are pre-trained with Implicit Q-Learning (IQL) on offline datasets. Fine-tuning uses IQL updates combined with JSRL curricula. The replay buffer mixes 75% online samples and 25% offline samples during updates.
    2. Simulated Vision-Based Robotic Grasping:

      • Setup: A 7-DoF robotic arm positioned above three bins filled with diverse objects, observed via an over-the-shoulder RGB camera.
      • Actions: 4D continuous end-effector Cartesian displacements (Δx,Δy,Δz,Δyaw)(\Delta x, \Delta y, \Delta z, \Delta \text{yaw}) plus discrete gripper open/close commands.
      • Tasks: Indiscriminate Grasping (sparse binary reward for lifting any object) and Instance Grasping (sparse binary reward only when a designated target object is grasped).
      • Data ablation: 20, 200, 2k, and 20k offline demonstrations used to pre-train Behavioral Cloning (BC) guide-policies. Online fine-tuning is conducted using QT-Opt as the exploration engine for up to 100,000 steps.
  7. Knowl 7 — Empirical Performance of IQL+JSRL on D4RL Offline-to-Online Benchmarks

    data/table

    The table below reports average normalized scores after 1 million online fine-tuning steps across D4RL Ant Maze and Adroit tasks across various offline dataset sizes. IQL+JSRL with curriculum and random switching is compared against AWAC, BC, CQL, and vanilla IQL.

    Environment Dataset AWAC BC CQL IQL IQL+JSRL (Ours)
    Curriculum Random
    antmaze-umaze-v0 1k 0.0±0.00.0 \pm 0.0 0.00.0 0.0±0.00.0 \pm 0.0 0.2±0.50.2 \pm 0.5 15.6±19.915.6 \pm 19.9 10.4±9.610.4 \pm 9.6
    10k 0.0±0.00.0 \pm 0.0 1.01.0 0.0±0.00.0 \pm 0.0 55.5±12.555.5 \pm 12.5 71.7±14.571.7 \pm 14.5 52.3±26.752.3 \pm 26.7
    100k 0.0±0.00.0 \pm 0.0 62.062.0 0.0±0.00.0 \pm 0.0 74.2±25.674.2 \pm 25.6 93.7±4.293.7 \pm 4.2 92.1±2.892.1 \pm 2.8
    1m 93.67±1.8993.67 \pm 1.89 61.061.0 64.33±45.5864.33 \pm 45.58 97.6±3.297.6 \pm 3.2 98.1±1.498.1 \pm 1.4 95.0±3.095.0 \pm 3.0
    antmaze-umaze-diverse-v0 1k 0.0±0.00.0 \pm 0.0 0.00.0 0.0±0.00.0 \pm 0.0 0.0±0.00.0 \pm 0.0 3.1±8.03.1 \pm 8.0 1.9±4.81.9 \pm 4.8
    10k 0.0±0.00.0 \pm 0.0 1.01.0 0.0±0.00.0 \pm 0.0 33.1±10.733.1 \pm 10.7 72.6±12.272.6 \pm 12.2 39.4±20.139.4 \pm 20.1
    100k 0.0±0.00.0 \pm 0.0 13.013.0 0.0±0.00.0 \pm 0.0 29.9±23.129.9 \pm 23.1 81.3±23.081.3 \pm 23.0 82.3±14.282.3 \pm 14.2
    1m 46.67±3.6846.67 \pm 3.68 80.080.0 0.50±0.500.50 \pm 0.50 53.0±30.553.0 \pm 30.5 88.6±16.388.6 \pm 16.3 89.8±10.089.8 \pm 10.0
    antmaze-medium-play-v0 10k 0.0±0.00.0 \pm 0.0 0.00.0 0.0±0.00.0 \pm 0.0 0.1±0.30.1 \pm 0.3 16.7±12.916.7 \pm 12.9 3.8±5.03.8 \pm 5.0
    100k 0.0±0.00.0 \pm 0.0 0.00.0 0.0±0.00.0 \pm 0.0 32.8±32.632.8 \pm 32.6 86.7±3.786.7 \pm 3.7 56.2±28.856.2 \pm 28.8
    1m 0.0±0.00.0 \pm 0.0 0.00.0 0.0±0.00.0 \pm 0.0 92.8±2.792.8 \pm 2.7 91.1±3.991.1 \pm 3.9 87.8±4.287.8 \pm 4.2
    antmaze-medium-diverse-v0 10k 0.0±0.00.0 \pm 0.0 0.00.0 0.0±0.00.0 \pm 0.0 0.0±0.00.0 \pm 0.0 16.6±11.716.6 \pm 11.7 5.1±8.25.1 \pm 8.2
    100k 0.0±0.00.0 \pm 0.0 0.00.0 0.0±0.00.0 \pm 0.0 15.7±17.715.7 \pm 17.7 81.5±18.881.5 \pm 18.8 67.0±17.467.0 \pm 17.4
    1m 0.0±0.00.0 \pm 0.0 0.00.0 0.0±0.00.0 \pm 0.0 92.4±4.592.4 \pm 4.5 93.1±3.193.1 \pm 3.1 86.3±5.986.3 \pm 5.9
    antmaze-large-play-v0 100k 0.0±0.00.0 \pm 0.0 0.00.0 0.0±0.00.0 \pm 0.0 2.6±8.22.6 \pm 8.2 36.3±16.436.3 \pm 16.4 17.7±13.417.7 \pm 13.4
    1m 0.0±0.00.0 \pm 0.0 0.00.0 0.0±0.00.0 \pm 0.0 62.4±12.462.4 \pm 12.4 62.9±11.362.9 \pm 11.3 48.6±10.048.6 \pm 10.0
    antmaze-large-diverse-v0 100k 0.0±0.00.0 \pm 0.0 0.00.0 0.0±0.00.0 \pm 0.0 4.1±10.44.1 \pm 10.4 34.4±23.034.4 \pm 23.0 22.4±15.422.4 \pm 15.4
    1m 0.0±0.00.0 \pm 0.0 0.00.0 0.0±0.00.0 \pm 0.0 68.3±8.968.3 \pm 8.9 68.3±8.868.3 \pm 8.8 58.3±6.558.3 \pm 6.5
    pen-binary-v0 100 3.13±4.433.13 \pm 4.43 0.00.0 31.46±9.9931.46 \pm 9.99 18.8±11.618.8 \pm 11.6 24.3±12.124.3 \pm 12.1 29.1±7.629.1 \pm 7.6
    1k 1.43±1.101.43 \pm 1.10 0.00.0 54.50±0.054.50 \pm 0.0 30.1±10.230.1 \pm 10.2 36.7±7.936.7 \pm 7.9 46.3±6.346.3 \pm 6.3
    10k 2.21±1.302.21 \pm 1.30 0.00.0 51.36±4.3451.36 \pm 4.34 38.4±11.238.4 \pm 11.2 44.3±6.244.3 \pm 6.2 52.1±3.352.1 \pm 3.3
    100k 1.23±1.081.23 \pm 1.08 0.00.0 59.58±1.4359.58 \pm 1.43 65.0±2.965.0 \pm 2.9 62.6±3.662.6 \pm 3.6 60.6±2.760.6 \pm 2.7

    The data shows that while JSRL matches competitive offline-to-online methods in the large-data regime (1m transitions), it delivers substantial gains in small-to-medium data regimes (10k and 100k transitions), where baseline algorithms fail to learn due to sparse rewards and poor state coverage.

  8. Knowl 8 — Grasping Performance Under Constrained Offline Demonstrations

    data/table

    The table below evaluates vision-based robotic grasping success rates under varying demonstration counts (20 to 20,000 demonstrations) on simulated Indiscriminate Grasping and Instance Grasping tasks after fine-tuning.

    Environment Demos AW-Opt BC QT-Opt QT-Opt+JSRL QT-Opt+JSRL Random
    Indiscriminate Grasping 20 0.33±0.430.33 \pm 0.43 0.19±0.040.19 \pm 0.04 0.00±0.000.00 \pm 0.00 0.91±0.010.91 \pm 0.01 0.89±0.000.89 \pm 0.00
    Indiscriminate Grasping 200 0.93±0.020.93 \pm 0.02 0.23±0.000.23 \pm 0.00 0.92±0.020.92 \pm 0.02 0.92±0.000.92 \pm 0.00 0.92±0.010.92 \pm 0.01
    Indiscriminate Grasping 2k 0.93±0.010.93 \pm 0.01 0.40±0.060.40 \pm 0.06 0.92±0.010.92 \pm 0.01 0.93±0.020.93 \pm 0.02 0.94±0.020.94 \pm 0.02
    Indiscriminate Grasping 20k 0.93±0.040.93 \pm 0.04 0.92±0.000.92 \pm 0.00 0.93±0.000.93 \pm 0.00 0.95±0.010.95 \pm 0.01 0.94±0.000.94 \pm 0.00
    Instance Grasping 20 0.44±0.050.44 \pm 0.05 0.05±0.030.05 \pm 0.03 0.29±0.200.29 \pm 0.20 0.54±0.020.54 \pm 0.02 0.53±0.020.53 \pm 0.02
    Instance Grasping 200 0.44±0.040.44 \pm 0.04 0.16±0.010.16 \pm 0.01 0.44±0.040.44 \pm 0.04 0.52±0.010.52 \pm 0.01 0.55±0.020.55 \pm 0.02
    Instance Grasping 2k 0.42±0.020.42 \pm 0.02 0.30±0.010.30 \pm 0.01 0.15±0.220.15 \pm 0.22 0.52±0.020.52 \pm 0.02 0.57±0.020.57 \pm 0.02
    Instance Grasping 20k 0.55±0.010.55 \pm 0.01 0.48±0.010.48 \pm 0.01 0.27±0.200.27 \pm 0.20 0.55±0.010.55 \pm 0.01 0.56±0.020.56 \pm 0.02

    Key takeaways:

    • In the small-data regime of only 20 demonstrations (100×\times fewer than the standard 2,000 demonstrations), QT-Opt+JSRL achieves 91% success on Indiscriminate Grasping and 54% on Instance Grasping, whereas baseline QT-Opt fails completely (0% on Indiscriminate Grasping) and AW-Opt/BC achieve markedly lower performance.
    • When demonstration count is abundant (20,000 demos), baseline methods close the gap as exploration ceases to be the primary bottleneck.
  9. Knowl 9 — Curriculum versus Random Guide-Step Switching Dynamics

    empirical result

    Comparing the structured backward curriculum (JSRL) against uniform random guide-step sampling (JSRL-Random) shows that:

    1. Asymptotic Convergence: Both JSRL-Curriculum and JSRL-Random converge to comparable final performance across both D4RL maze tasks and vision-based robotic manipulation tasks.
    2. Early-Stage Sample Efficiency: JSRL-Curriculum exhibits faster learning during the early stages of online training than JSRL-Random. Rolling in backwards from the goal enables the exploration policy πe\pi^e to reliably encounter high-reward states early on.
    3. Core Mechanism: These findings demonstrate that the primary benefit of JSRL is driven by the state visitation distribution induced by the guide-policy πg\pi^g (placing the agent in high-value feature states), while the structured curriculum provides an additional acceleration in early sample efficiency.
  10. Knowl 10 — Cross-Task Generalization of Guide-Policies

    empirical result

    JSRL can accelerate learning even when the guide-policy πg\pi^g is trained on an easier or related task rather than the exact target task:

    1. Indiscriminate to Instance Grasping: An indiscriminate grasping guide policy (trained online with QT-Opt to 90% indiscriminate success, achieving only 5% instance grasping success) was used as the guide policy for instance grasping. QT-Opt+JSRL with the indiscriminate guide achieved over 50% instance grasping success within 40k steps, outperforming vanilla QT-Opt.
    2. Ant Maze Play to Diverse Transfer: Guide-policies trained offline on simpler antmaze-*-play environments (fixed start and goal positions) were transferred to initialize fine-tuning on antmaze-*-diverse environments (random starts and goals). IQL+JSRL initialized with the play-task guide policy generalized better to unseen goals and initial states compared to vanilla IQL fine-tuning from scratch.
  11. Knowl 11 — Limitations and Failure Modes of JSRL

    limitation

    The effectiveness of Jump-Start Reinforcement Learning is bounded by several specific conditions:

    1. Inherited Bias and Suboptimality: JSRL remains susceptible to biases present in the prior data or guide policy. In safety-critical settings like physical robotics, a poor or misaligned guide-policy could execute hazardous actions during the roll-in phase.
    2. Adversarial or Pathological Guide-Policies: If πg\pi^g actively guides the agent away from goals or remains static in a confined region, exploration is constrained during the hh initial steps. Learning can become strictly slower than random exploration until the curriculum successfully reduces hh to 0.
    3. Hyperparameter Sensitivity of Stage Progression (β\beta): Advancing curriculum stages based on noisy evaluation metrics with excessive tolerance allows the algorithm to reduce hh before πe\pi^e has sufficiently mastered the current sub-horizon, causing policy degradation.

Coverage note — None omitted; all core theoretical bounds (exponential lower bound, polynomial upper bounds for tabular and function approximation), algorithm variants (curriculum and random switching), experimental domains (D4RL and vision-based manipulation), ablation results, and stated limitations are included.

References

  1. 1.Agarwal, A., Henaff, M., Kakade, S., and Sun, W. Pc-pg: Policy cover directed exploration for provable policy gradient learning. arXiv preprint arXiv:2007.08459, 2020.
  2. 2.Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. On the theory of policy gradient methods: Optimality, approximation, and distribution shift. Journal of Machine Learning Research, 22(98):1–76, 2021.
  3. 3.Bagnell, J., Kakade, S. M., Schneider, J., and Ng, A. Policy search by dynamic programming. Advances in neural information processing systems, 16, 2003.
  4. 4.Bagnell, J. A. Learning decisions: Robustness, uncertainty, and approximation. Carnegie Mellon University, 2004.
  5. 5.Chen, J. and Jiang, N. Information-theoretic considerations in batch reinforcement learning. arXiv preprint arXiv:1905.00360, 2019.
  6. 6.Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems, 34, 2021.
  7. 7.Chu, W., Li, L., Reyzin, L., and Schapire, R. Contextual bandits with linear payoff functions. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, pp. 208–214. JMLR Workshop and Conference Proceedings, 2011.
  8. 8.Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J. First return, then explore. Nature, 590(7847): 580–586, 2021.
  9. 9.Florensa, C., Held, D., Wulfmeier, M., Zhang, M., and Abbeel, P. Reverse curriculum generation for reinforcement learning. In Conference on robot learning, pp. 482–495. PMLR, 2017.
  10. 10.Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. D4RL: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219, 2020.
  11. 11.Hosu, I.-A. and Rebedea, T. Playing atari games with deep reinforcement learning and human checkpoint replay. arXiv preprint arXiv:1607.05077, 2016.
  12. 12.Ivanovic, B., Harrison, J., Sharma, A., Chen, M., and Pavone, M. Barc: Backward reachability curriculum for robotic reinforcement learning. In 2019 International Conference on Robotics and Automation (ICRA), pp. 15–21. IEEE, 2019.
  13. 13.Janner, M., Li, Q., and Levine, S. Offline reinforcement learning as one big sequence modeling problem. Advances in neural information processing systems, 34, 2021.
  14. 14.Jiang, N. On value functions and the agent-environment boundary. arXiv preprint arXiv:1905.13341, 2019.
  15. 15.Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. Contextual decision processes with low bellman rank are pac-learnable. In International Conference on Machine Learning, pp. 1704–1713. PMLR, 2017.
  16. 16.Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. Is Q-learning provably efficient? In Proceedings of the 32nd International Conference on Neural Information Processing Systems, pp. 4868–4878, 2018.
  17. 17.Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. Provably efficient reinforcement learning with linear function approximation. In Conference on Learning Theory, pp. 2137–2143. PMLR, 2020.
  18. 18.Kakade, S. and Langford, J. Approximately optimal approximate reinforcement learning. In ICML, volume 2, pp. 267–274, 2002.
  19. 19.Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al. Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation. arXiv preprint arXiv:1806.10293, 2018.
  20. 20.Kober, J., Mohler, B., and Peters, J. Imitation and reinforcement learning for motor primitives with perceptual coupling. In From motor learning to interaction learning in robots, pp. 209–225. Springer, 2010.
  21. 21.Koenig, S. and Simmons, R. G. Complexity analysis of real-time reinforcement learning. In AAAI, pp. 99–107, 1993.
  22. 22.Kostrikov, I., Nair, A., and Levine, S. Offline reinforcement learning with implicit q-learning. arXiv preprint arXiv:2110.06169, 2021.
  23. 23.Krishnamurthy, A., Langford, J., Slivkins, A., and Zhang, C. Contextual bandits with continuous actions: Smoothing, zooming, and adapting. In Conference on Learning Theory, pp. 2025–2027. PMLR, 2019.
  24. 24.Kumar, A., Zhou, A., Tucker, G., and Levine, S. Conservative q-learning for offline reinforcement learning. arXiv preprint arXiv:2006.04779, 2020.
  25. 25.Langford, J. and Zhang, T. The epoch-greedy algorithm for contextual multi-armed bandits. Advances in neural information processing systems, 20(1):96–1, 2007.
  26. 26.Liao, P., Qi, Z., and Murphy, S. Batch policy learning in average reward Markov decision processes. arXiv preprint arXiv:2007.11771, 2020.
  27. 27.Lin, L.-J. Self-improving reactive agents based on reinforcement learning, planning and teaching. Machine learning, 8(3-4):293–321, 1992.
  28. 28.Liu, B., Cai, Q., Yang, Z., and Wang, Z. Neural trust region/proximal policy optimization attains globally optimal policy. In Neural Information Processing Systems, 2019.
  29. 29.Lu, Y., Hausman, K., Chebotar, Y., Yan, M., Jang, E., Herzog, A., Xiao, T., Irpan, A., Khansari, M., Kalashnikov, D., and Levine, S. Aw-opt: Learning robotic skills with imitation andreinforcement at scale. In 2021 Conference on Robot Learning (CoRL), 2021.
  30. 30.McAleer, S., Agostinelli, F., Shmakov, A., and Baldi, P. Solving the rubik’s cube without human knowledge. 2019.
  31. 31.Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. Playing Atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602, 2013.
  32. 32.Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P. Overcoming exploration in reinforcement learning with demonstrations. In 2018 IEEE International Conference on Robotics and Automation (ICRA), pp. 6292–6299. IEEE, 2018.
  33. 33.Nair, A., Gupta, A., Dalal, M., and Levine, S. Awac: Accelerating online reinforcement learning with offline datasets. 2020.
  34. 34.Osband, I. and Van Roy, B. Model-based reinforcement learning and the eluder dimension. arXiv preprint arXiv:1406.1853, 2014.
  35. 35.Ouyang, Y., Gagrani, M., Nayyar, A., and Jain, R. Learning unknown markov decision processes: A thompson sampling approach. arXiv preprint arXiv:1709.04570, 2017.
  36. 36.Peng, X. B., Abbeel, P., Levine, S., and van de Panne, M. Deepmimic: Example-guided deep reinforcement learning of physics-based character skills. ACM Trans. Graph., 37(4), July 2018.
  37. 37.Peng, X. B., Kumar, A., Zhang, G., and Levine, S. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. arXiv preprint arXiv:1910.00177, 2019.
  38. 38.Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S. Learning complex dexterous manipulation with deep reinforcement learning and demonstrations. arXiv preprint arXiv:1709.10087, 2017.
  39. 39.Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. Bridging offline reinforcement learning and imitation learning: A tale of pessimism. arXiv preprint arXiv:2103.12021, 2021.
  40. 40.Resnick, C., Raileanu, R., Kapoor, S., Peysakhovich, A., Cho, K., and Bruna, J. Backplay:” man muss immer umkehren”. arXiv preprint arXiv:1807.06919, 2018.
  41. 41.Ross, S. and Bagnell, D. Efficient reductions for imitation learning. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pp. 661–668, 2010.
  42. 42.Salimans, T. and Chen, R. Learning montezuma’s revenge from a single demonstration. arXiv preprint arXiv:1812.03381, 2018.
  43. 43.Schaal, S. et al. Learning from demonstration. Advances in neural information processing systems, pp. 1040–1046, 1997.
  44. 44.Scherrer, B. Approximate policy iteration schemes: A comparison. In International Conference on Machine Learning, pp. 1314–1322, 2014.
  45. 45.Simchi-Levi, D. and Xu, Y. Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability. Available at SSRN 3562765, 2020.
  46. 46.Smart, W. and Pack Kaelbling, L. Effective reinforcement learning for mobile robots. In Proceedings 2002 IEEE International Conference on Robotics and Automation (Cat. No.02CH37292), volume 4, pp. 3404–3410 vol.4, 2002. doi: 10.1109/ROBOT.2002.1014237.
  47. 47.Vecerik, M., Hester, T., Scholz, J., Wang, F., Pietquin, O., Piot, B., Heess, N., Rothorl, T., Lampe, T., and Riedmiller, M. Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards, 2018.
  48. 48.Wang, L., Cai, Q., Yang, Z., and Wang, Z. Neural policy gradient methods: Global optimality and rates of convergence. In International Conference on Learning Representations, 2019.
  49. 49.Xie, T., Jiang, N., Wang, H., Xiong, C., and Bai, Y. Policy finetuning: Bridging sample-efficient offline and online reinforcement learning. arXiv preprint arXiv:2106.04895, 2021.
  50. 50.Zanette, A., Lazaric, A., Kochenderfer, M., and Brunskill, E. Learning near optimal policies with low inherent bellman error. In International Conference on Machine Learning, pp. 10978–10989. PMLR, 2020.
  51. 51.Zhang, J., Koppel, A., Bedi, A. S., Szepesvari, C., and Wang, M. Variational policy gradient method for reinforcement learning with general utilities. arXiv preprint arXiv:2007.02151, 2020a.
  52. 52.Zhang, Z., Zhou, Y., and Ji, X. Almost optimal model-free reinforcement learning via reference-advantage decomposition. Advances in Neural Information Processing Systems, 33, 2020b.
  53. 53.Zheng, Q., Zhang, A., and Grover, A. Online decision transformer. arXiv preprint arXiv:2202.05607, 2022.

Citation

MLA
Uchendu, I., et al. “Jump-Start Reinforcement Learning”. International Conference on Machine Learning, vol. 202, 2023, pp. 34556–83, https://proceedings.mlr.press/v202/uchendu23a.html.
APA
Uchendu, I., Xiao, T., Lu, Y., Zhu, B., Yan, M., Simon, J., Bennice, M., Fu, C., Ma, C., Jiao, J., Levine, S., & Hausman, K. (2023). Jump-Start Reinforcement Learning. International Conference on Machine Learning, 202, 34556–34583. https://proceedings.mlr.press/v202/uchendu23a.html
Chicago
Uchendu, I., T. Xiao, Y. Lu, et al. 2023. “Jump-Start Reinforcement Learning”. International Conference on Machine Learning 202: 34556–83. https://proceedings.mlr.press/v202/uchendu23a.html.
Harvard
Uchendu, I. et al. (2023) “Jump-Start Reinforcement Learning”, International Conference on Machine Learning. PMLR, pp. 34556–34583. Available at: https://proceedings.mlr.press/v202/uchendu23a.html.
Vancouver
1. Uchendu I, Xiao T, Lu Y, et al (2023) Jump-Start Reinforcement Learning. In: International Conference on Machine Learning. PMLR, pp 34556–34583

BibTeX

@InProceedings{pmlr-v202-uchendu23a,
  title = 	 {Jump-Start Reinforcement Learning},
  author =       {Uchendu, Ikechukwu and Xiao, Ted and Lu, Yao and Zhu, Banghua and Yan, Mengyuan and Simon, Jos\'{e}phine and Bennice, Matthew and Fu, Chuyuan and Ma, Cong and Jiao, Jiantao and Levine, Sergey and Hausman, Karol},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {34556--34583},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/uchendu23a/uchendu23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/uchendu23a.html},
  abstract = 	 {Reinforcement learning (RL) provides a theoretical framework for continuously improving an agent’s behavior via trial and error. However, efficiently learning policies from scratch can be very difficult, particularly for tasks that present exploration challenges. In such settings, it might be desirable to initialize RL with an existing policy, offline data, or demonstrations. However, naively performing such initialization in RL often works poorly, especially for value-based methods. In this paper, we present a meta algorithm that can use offline data, demonstrations, or a pre-existing policy to initialize an RL policy, and is compatible with any RL approach. In particular, we propose Jump-Start Reinforcement Learning (JSRL), an algorithm that employs two policies to solve tasks: a guide-policy, and an exploration-policy. By using the guide-policy to form a curriculum of starting states for the exploration-policy, we are able to efficiently improve performance on a set of simulated robotic tasks. We show via experiments that it is able to significantly outperform existing imitation and reinforcement learning algorithms, particularly in the small-data regime. In addition, we provide an upper bound on the sample complexity of JSRL and show that with the help of a guide-policy, one can improve the sample complexity for non-optimism exploration methods from exponential in horizon to polynomial.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/