How to Leverage Unlabeled Data in Offline Reinforcement Learning

Tianhe YuAviral KumarYevgen ChebotarKarol HausmanChelsea FinnSergey Levine

article2022ICML82 citations

Demonstrates that assigning a constant zero reward to unlabeled offline data—when combined with conservative sample reweighting—outperforms learned reward models across robotic control tasks by effectively balancing reward bias against sample complexity.

Listen

Offline reinforcement learning enables systems to learn effective control policies from pre-collected datasets without requiring active, potentially risky or costly exploration. However, existing methods heavily rely on datasets where every transition is explicitly annotated with task-specific rewards. In real-world applications, such as robotics, collecting vast amounts of unlabeled background interaction data is inexpensive, whereas manually annotating or programmatically computing task rewards is prohibitively difficult and expensive. Previous approaches that attempt to utilize unlabeled data typically train separate reward prediction models or inverse reinforcement learning classifiers to infer missing labels, but these methods add significant modeling complexity and frequently fail due to reward overestimation errors.

The main objective of the article is to demonstrate and theoretically analyze a remarkably simple alternative: labeling all unlabeled prior data with a constant minimum reward (such as zero) to safely share data across single-task and multi-task offline reinforcement learning problems. The article evaluates this baseline strategy, termed unlabeled data sharing, and assesses an enhanced variant that incorporates conservative reweighting to minimize distributional shift and reward bias.

To establish these insights, the authors develop theoretical performance bounds that formalize the trade-off between reward bias, sample complexity, and distributional shift. They test the approach across standard simulated benchmarks, including robotic locomotion, maze navigation, and both state-based and vision-based robotic manipulation tasks. The evaluations compare the proposed methods against training solely on labeled data, complex reward predictors, inverse reinforcement learning techniques, representation learning methods, and oracle baselines that utilize ground-truth reward functions.

The investigation yields several key findings:

  1. The simple zero-reward labeling strategy consistently outperforms conventional reward prediction models and inverse reinforcement learning methods across virtually all evaluated benchmarks. In single-task navigation, for example, zero-reward labeling achieved success rates above 80%, whereas reward prediction and classifier-based methods failed completely (0.0%).
  2. The simple approach regularly approaches the performance of oracle baselines that possess perfect programmatic knowledge of the reward function. Diagnostic tests revealed this occurred even when roughly 60% of the unlabeled transitions were actually successful demonstrations of the task, proving that explicit reward accuracy is less critical than the reduction in sampling error.
  3. Combining zero-reward labeling with conservative reweighting (which prioritizes unlabeled transitions based on conservative value estimates) substantially improves performance further. In multi-task manipulation, conservative reweighting increased average success rates from 56.4% to 71.2%, matching the oracle baseline (70.1%). In vision-based robotic manipulation, it achieved a 75.0% success rate across ten tasks, improving by roughly 14 percentage points over training without data sharing (60.8%).
  4. The simple strategy thrives when labeled data is limited in quantity or narrow in coverage and unlabeled data is abundant. However, it degrades when unlabeled datasets are very small or when labeled data already possesses broad coverage and medium quality, where the introduced reward bias outweighs any reduction in sampling error.

These findings indicate that organizations developing autonomous systems can bypass the engineering overhead, instability, and expense of designing complex reward estimation pipelines. Instead, practitioners can leverage large repositories of uncurated, unlabeled interaction data simply by assigning baseline zero rewards. The results counter the standard expectation that inaccurate reward annotations inherently degrade reinforcement learning performance, demonstrating instead that conservative offline learning algorithms can safely exploit the broader state coverage provided by biased data.

For practical implementation, teams should adopt zero-reward data sharing when target task demonstrations are scarce but general interaction data is plentiful. In more complex multi-task or high-dimensional environments, decision-makers should integrate conservative reweighting schemes to filter out detrimental distribution shifts. Conversely, if unlabeled datasets are tiny or the target dataset already covers the state space comprehensively, sharing should be avoided unless combined with reweighting.

Confidence in these findings is supported by rigorous mathematical proofs and consistent results across diverse continuous control domains. However, readers should note that evaluations were conducted in simulated benchmark environments and state-based robotic setups. Further validation on physical, real-world robotic systems and broader industrial workflows is recommended to confirm transferability and determine exact hyperparameters under real-world noise.

arXiv: 2202.01741
Cover for How to Leverage Unlabeled Data in Offline Reinforcement Learning

Abstract

Offline reinforcement learning (RL) can learn control policies from static datasets but, like standard RL methods, it requires reward annotations for every transition. In many cases, labeling large datasets with rewards may be costly, especially if those rewards must be provided by human labelers, while collecting diverse unlabeled data might be comparatively inexpensive. How can we best leverage such unlabeled data in offline RL? One natural solution is to learn a reward function from the labeled data and use it to label the unlabeled data. In this paper, we find that, perhaps surprisingly, a much simpler method that simply applies zero rewards to unlabeled data leads to effective data sharing both in theory and in practice, without learning any reward model at all. While this approach might seem strange (and incorrect) at first, we provide extensive theoretical and empirical analysis that illustrates how it trades off reward bias, sample complexity and distributional shift, often leading to good results. We characterize conditions under which this simple strategy is effective, and further show that extending it with a simple reweighting approach can further alleviate the bias introduced by using incorrect reward labels. Our empirical evaluation confirms these findings in simulated robotic locomotion, navigation, and manipulation settings.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 4 How To Use Unlabeled Data in Offline RL
  • 4.1 Theoretical Analysis of UDS in Offline RL
  • 4.2 Reducing Reward Bias by Reweighting Data
  • 4.3 Why Can We Expect UDS to Outperform Reward Prediction Methods?
  • 5 Experiments
  • 5.1 Results of Empirical Evaluations
  • 5.2 Empirical Analysis of UDS and CDS+UDS
  • 6 Discussion
  • Acknowledgements
  • References
  • A Proofs for Theoretical Analysis of UDS and Optimized Reweighting Schemes
  • A.1 Notation and Assumptions
  • A.2 Performance Guarantee for UDS: Proof for Theorem 4.1
  • A.3 When is Reward Bias Small? Proof of Theorem 4.2
  • A.4 When is The Bound in Theorem 4.1 Tightest?: Proof of Theorem 4.3
  • B Additional empirical analysis of the reason that CDS+UDS and UDS work
  • B.1 Meta-World Domains
  • B.2 D4RL Hopper Diagnostic Study on Varying Unlabeled Dataset Size
  • B.3 D4RL Hopper Ablation Study on Reward Learning Methods with Varying Labeled Dataset Size and Quality
  • B.4 Takeaways from the empirical analysis
  • C Additional details on the quality of data shared from other tasks in the multi-task offline RL setting
  • D Empirical Results of UDS and CDS+UDS in multi-task locomotion domain with dense rewards
  • E Comparisons of CDS+UDS and UDS to Multi-Task Model-Based Offline RL Approaches
  • F Comparisons to Pre-Trained Representation Learning on Unlabeled Data
  • G Details of UDS and CDS+UDS
  • G.1 Details on the training procedure
  • G.2 Details on the environment and the datasets
  • G.3 Computation Complexity

Knowls

  1. Knowl 1 — Unlabeled Data Sharing by Minimum-Reward Relabeling

    model/method

    The paper introduces unlabeled data sharing (UDS) for offline reinforcement learning. Let DLD_L be a dataset of transitions (s,a,s′,r)(s,a,s',r) with task reward labels and let DUD_U be a dataset of transitions without labels. UDS constructs the effective training dataset

    Deff=DL∪{(s,a,s′,0):(s,a,s′)∈DU}.D_{\mathrm{eff}}=D_L\cup\{(s,a,s',0):(s,a,s')\in D_U\}.

    The zero is the minimum reward after affine rescaling, so the same procedure can use any minimum reward. UDS then trains an ordinary conservative offline-RL algorithm on DeffD_{\mathrm{eff}} without learning a reward model or requiring the programmatic form of the reward function. In a multi-task problem, transitions from the target task retain their labels, while transitions collected for other tasks are relabeled with zero for the target task.

  2. Knowl 2 — Theoretical Setting for UDS Analysis

    assumption

    The theoretical analysis considers a finite, tabular Markov decision process M=(S,A,P,γ,r)M=(\mathcal S,\mathcal A,P,\gamma,r) with discount factor γ∈[0,1)\gamma\in[0,1) and rewards r(s,a)∈[0,1]r(s,a)\in[0,1]. A dataset DD is generated by a behavior policy πβ(a∣s)\pi_\beta(a\mid s), and r^\widehat r and P^\widehat P denote empirical reward and transition estimates. The analysis assumes every state-action pair is observed and that, with probability at least 1−δ1-\delta,

    ∣r^(s,a)−r(s,a)∣≤Cr,δ∣D(s,a)∣,∥P^(⋅∣s,a)−P(⋅∣s,a)∥1≤CP,δ∣D(s,a)∣,|\widehat r(s,a)-r(s,a)|\leq\sqrt{\frac{C_{r,\delta}}{|D(s,a)|}},\qquad \|\widehat P(\cdot\mid s,a)-P(\cdot\mid s,a)\|_1\leq\sqrt{\frac{C_{P,\delta}}{|D(s,a)|}},

    where ∣D(s,a)∣|D(s,a)| is the number of transitions with state-action pair (s,a)(s,a) and Cr,δ,CP,δC_{r,\delta},C_{P,\delta} are concentration constants. The paper abstracts conservative offline RL as optimizing

    π∗=arg⁡max⁡π[JD(π)−α1−γD(π,πβ)],\pi^*=\arg\max_{\pi}\left[J_D(\pi)-\frac{\alpha}{1-\gamma}D(\pi,\pi_\beta)\right],

    where JD(π)J_D(\pi) is the policy return in the empirical MDP, α≥0\alpha\geq0 is the conservatism coefficient, and DD is a divergence between the learned policy and the dataset behavior policy. For the conservative-Q divergence used in the analysis, DCQL(p,q)=∑ap(a)[p(a)/q(a)−1]D_{\mathrm{CQL}}(p,q)=\sum_a p(a)[p(a)/q(a)-1] for action distributions pp and qq with common support.

  3. Knowl 3 — Safe Policy Improvement Guarantee for UDS

    theoretical result

    Under the finite-MDP, reward-boundedness, coverage, and concentration assumptions, the policy πUDS∗\pi^*_{\mathrm{UDS}} learned from the zero-reward effective dataset is, with probability at least 1−δ1-\delta, a safe improvement over the effective behavior policy πβeff\pi^{\mathrm{eff}}_\beta:

    J(πUDS∗)≥J(πβeff)−ζerr+α1−γD(πUDS∗,πβeff).J(\pi^*_{\mathrm{UDS}})\geq J(\pi^{\mathrm{eff}}_\beta)-\zeta_{\mathrm{err}}+\frac{\alpha}{1-\gamma}D(\pi^*_{\mathrm{UDS}},\pi^{\mathrm{eff}}_\beta).

    Here J(π)J(\pi) is the true discounted return, and the error term decomposes into reward bias and sampling error,

    ζerr=Breward+O ⁣(γ(1−γ)2)Es∼d^πUDS∗[DCQL(πUDS∗,πβeff)(s)∣Deff(s)∣],\zeta_{\mathrm{err}}=B_{\mathrm{reward}}+O\!\left(\frac{\gamma}{(1-\gamma)^2}\right)\mathbb E_{s\sim\widehat d^{\pi^*_{\mathrm{UDS}}}}\left[\sqrt{\frac{D_{\mathrm{CQL}}(\pi^*_{\mathrm{UDS}},\pi^{\mathrm{eff}}_\beta)(s)}{|D_{\mathrm{eff}}(s)|}}\right],

    with

    Breward=11−γ∑s,a(d^πβeff(s,a)−d^πUDS∗(s,a))(1−f(s,a))r(s,a),f(s,a)=∣DL(s,a)∣∣Deff(s,a)∣.B_{\mathrm{reward}}=\frac{1}{1-\gamma}\sum_{s,a}\left(\widehat d^{\pi^{\mathrm{eff}}_\beta}(s,a)-\widehat d^{\pi^*_{\mathrm{UDS}}}(s,a)\right)(1-f(s,a))r(s,a), \qquad f(s,a)=\frac{|D_L(s,a)|}{|D_{\mathrm{eff}}(s,a)|}.

    The quantities d^π(s,a)\widehat d^\pi(s,a) and d^π(s)\widehat d^\pi(s) are state-action and state visitation marginals in the empirical MDP, ∣Deff(s)∣|D_{\mathrm{eff}}(s)| is the number of effective-data transitions from state ss, and DCQL(⋅,⋅)(s)D_{\mathrm{CQL}}(\cdot,\cdot)(s) is the conservative-Q divergence between the action distributions at state ss. The result formalizes a tradeoff: zero labels increase reward bias through 1−f(s,a)1-f(s,a), but additional transitions reduce sampling error and can increase the effective behavior quality.

  4. Knowl 4 — Conditions Determining Whether UDS Helps

    theoretical result

    The UDS bound identifies three regimes in which adding incorrectly zero-labeled data can be beneficial. First, if labeled and unlabeled data have the same state-action distribution, then f(s,a)f(s,a) is constant across state-action pairs. The reward-bias contribution reduces to a difference between the learned policy and the effective behavior policy in the empirical MDP, which is nonpositive when offline RL improves that behavior policy; the larger dataset then reduces sampling error without an additional reward-bias penalty.

    Second, UDS is favored when labeled data are high-quality but limited and unlabeled data are abundant, low- or medium-quality, and provide broader coverage. The zero labels are then roughly consistent with the unlabeled data being worse than the labeled demonstrations, while the additional coverage reduces sampling error. Even high-quality unlabeled data can help when it substantially improves the effective behavior policy.

    Third, for long-horizon tasks with effective horizon H=1/(1−γ)H=1/(1-\gamma), the sampling-error term can dominate the reward-bias term. When the effective state-wise dataset size satisfies ∣Deff(s)∣=Ω(H2)∣DL(s)∣|D_{\mathrm{eff}}(s)|=\Omega(H^2)|D_L(s)|, the additional data remove one factor of HH from the sampling error, whereas the reward bias grows only linearly with HH.

    UDS is expected to fail when the unlabeled set is not large enough to reduce sampling error, does not improve state coverage or behavior quality, and still introduces incorrect rewards. The paper's diagnostic examples include medium-quality labeled data combined with random unlabeled data and broad labeled datasets for which additional unlabeled data provide little new coverage.

  5. Knowl 5 — Optimal Reweighting Distribution for Reducing UDS Bias

    theoretical result

    The paper derives a distributional reweighting principle for zero-labeled data. Let dL(s,a)d_L(s,a) be the normalized labeled-data density, dπ(s,a)d^\pi(s,a) the state-action marginal of the policy learned from the reweighted dataset, and d^πβeff(s,a)\widehat d^{\pi^{\mathrm{eff}}_\beta}(s,a) the state-action distribution of the resulting effective behavior policy. The effective behavior distribution that minimizes the reward-bias term alone satisfies

    d^πβeff(s,a)∝dL(s,a)dπ(s,a).\widehat d^{\pi^{\mathrm{eff}}_\beta}(s,a)\propto\sqrt{d_L(s,a)d^\pi(s,a)}.

    Thus, reweighting should favor state-action pairs that are both represented in the labeled data and likely under the learned policy.

    When reward bias and sampling error are optimized jointly, the optimal effective distribution is the solution p∗p^* of

    p∗=arg⁡min⁡p∈Δ∣S∣∣A∣∑s,a[C1d^π(s,a)p(s,a)+C2∣dL(s,a)∣d^π(s,a)p(s,a)],p^*=\arg\min_{p\in\Delta^{|\mathcal S||\mathcal A|}}\sum_{s,a}\left[C_1\frac{\widehat d^\pi(s,a)}{\sqrt{p(s,a)}}+C_2|d_L(s,a)|\frac{\widehat d^\pi(s,a)}{p(s,a)}\right],

    where p(s,a)p(s,a) is a probability distribution over state-action pairs, Δ∣S∣∣A∣\Delta^{|\mathcal S||\mathcal A|} is the corresponding simplex, C1=γCP,δ/[(1−γ)2∣Deff∣]C_1=\gamma C_{P,\delta}/[(1-\gamma)^2\sqrt{|D_{\mathrm{eff}}|}], and C2=∣DL∣/[(1−γ)∣Deff∣]C_2=|D_L|/[(1-\gamma)|D_{\mathrm{eff}}|]. The first term penalizes sampling error and the second penalizes reward bias. Both terms favor reducing distributional shift between the learned policy and the effective behavior distribution, explaining why conservative data-sharing reweighting is useful even when the shared rewards are all zero.

  6. Knowl 6 — CDS+UDS Conservative Reweighting Procedure

    algorithm

    The paper implements the theoretically motivated reweighting with conservative data sharing (CDS) combined with UDS. Every unlabeled transition is first assigned reward zero. For a single-task target, compute a conservative policy value Q^π(s,a)\widehat Q^\pi(s,a) for each unlabeled transition and compare it with the kkth-percentile value among labeled transitions. Define

    Δ(s,a;U ⁣→ ⁣L)=Q^π(s,a)−Pk%{Q^π(s′,a′):(s′,a′)∼DL},\Delta(s,a;U\!\to\!L)=\widehat Q^\pi(s,a)-P_{k\%}\{\widehat Q^\pi(s',a'):(s',a')\sim D_L\},

    and apply the soft weight

    wCDS(s,a;U ⁣→ ⁣L)=σ ⁣(Δ(s,a;U ⁣→ ⁣L)τ),w_{\mathrm{CDS}}(s,a;U\!\to\!L)=\sigma\!\left(\frac{\Delta(s,a;U\!\to\!L)}{\tau}\right),

    where σ\sigma is the sigmoid and τ\tau is a temperature. The weight is used in both the critic loss and policy objective, so transitions estimated to be less useful than the labeled-data threshold contribute less. The multi-task version computes the same quantity for a transition from task jj when relabeled for target task ii and applies the weight primarily to cross-task data.

    The experiments use k=50k=50 in all single-task domains, k=90k=90 in multi-task Meta-World, and k=50k=50 in the other multi-task domains. The temperature is an exponential running average of Δ\Delta with decay 0.9950.995; for state-based domains it is clipped to the interval [1,∞)[1,\infty). The underlying offline learner is CQL. In hopper experiments with random unlabeled data, the CQL coefficient is 1.01.0 and the standard Q-averaging term is removed; other hopper settings use coefficient 5.05.0. AntMaze uses the Lagrange CQL variant with target constraint τ=10.0\tau=10.0.

  7. Knowl 7 — Single-Task Locomotion and Navigation Results

    data/table

    The single-task evaluation combines 10,000 labeled transitions with large unlabeled datasets: hopper uses 1 million random or medium transitions, while AntMaze uses expert labeled transitions and medium-play or large-play unlabeled transitions. Scores are benchmark normalized returns, averaged over three random seeds. UDS improves over using labeled data only in three of four settings, while CDS+UDS improves over UDS in all four and remains close to oracle methods that use true rewards for relabeling.

    Environment; labeled data; unlabeled dataCDS+UDSUDSNo SharingReward PredictorVICERCECDS (oracle)Sharing All (oracle)
    Hopper; expert; random81.578.677.167.6n/an/a83.386.1
    Hopper; expert; medium78.364.477.151.7n/an/a82.564.6
    AntMaze; expert; medium-play82.682.717.20.00.00.083.583.1
    AntMaze; expert; large-play47.133.10.70.00.00.046.150.2

    "No Sharing" uses only labeled data; Reward Predictor learns a reward model; VICE and RCE infer rewards or values from labeled and unlabeled data; CDS and Sharing All have access to the true reward function. The results show that zero-reward sharing can match or exceed reward-learning methods despite requiring no reward model, and CDS reweighting is especially useful in cases where naive UDS incurs substantial reward bias.

  8. Knowl 8 — Multi-Task Manipulation, Navigation, and Vision Results

    data/table

    The multi-task evaluations use binary task rewards unless otherwise noted. The reported scores are success rates averaged over tasks; low-dimensional results average six seeds and image-based results average three seeds. UDS and CDS+UDS use no true rewards for cross-task relabeling, whereas CDS and Sharing All are oracle-reward baselines.

    Domain and aggregate task setCDS+UDSUDSVICERCENo SharingReward PredictorCDS (oracle)Sharing All (oracle)
    Meta-World, four tasks, average71.2%±11.3%71.2\%\pm11.3\%56.4%±12.8%56.4\%\pm12.8\%21.5%±0.7%21.5\%\pm0.7\%0.7%±0.4%0.7\%\pm0.4\%33.4%±8.3%33.4\%\pm8.3\%41.0%±11.9%41.0\%\pm11.9\%70.1%±8.1%70.1\%\pm8.1\%59.4%±5.7%59.4\%\pm5.7\%
    AntMaze, medium maze, 3 tasks31.5%±3.0%31.5\%\pm3.0\%26.5%±9.1%26.5\%\pm9.1\%2.9%±1.0%2.9\%\pm1.0\%0.0%±0.0%0.0\%\pm0.0\%21.6%±7.1%21.6\%\pm7.1\%3.8%±3.8%3.8\%\pm3.8\%36.7%±6.2%36.7\%\pm6.2\%22.9%±3.6%22.9\%\pm3.6\%
    AntMaze, large maze, 7 tasks18.4%±6.1%18.4\%\pm6.1\%14.2%±3.9%14.2\%\pm3.9\%2.5%±1.1%2.5\%\pm1.1\%0.0%±0.0%0.0\%\pm0.0\%13.3%±8.6%13.3\%\pm8.6\%5.9%±4.1%5.9\%\pm4.1\%22.8%±4.5%22.8\%\pm4.5\%16.7%±7.0%16.7\%\pm7.0\%
    Image-based manipulation, 10 tasks, average75.0%±3.3%75.0\%\pm3.3\%67.3%±0.8%67.3\%\pm0.8\%n/an/a60.8%±7.5%60.8\%\pm7.5\%n/a74.8%±6.4%74.8\%\pm6.4\%66.4%±7.2%66.4\%\pm7.2\%

    UDS beats No Sharing on the Meta-World and AntMaze aggregates and on the image-based manipulation average; CDS+UDS improves further and is competitive with oracle reward-sharing methods. In the separate dense-reward multi-task walker experiment, the three-task average scores were 1028.6±76.81028.6\pm76.8 for CDS+UDS, 796.7±106.3796.7\pm106.3 for UDS, 926.6±37.7926.6\pm37.7 for No Sharing, 506.7±343.6506.7\pm343.6 for Reward Predictor, 1013.6±71.51013.6\pm71.5 for oracle CDS, and 781.0±100.8781.0\pm100.8 for oracle Sharing All.

  9. Knowl 9 — Sensitivity to Data Quality, Coverage, and Unlabeled-Set Size

    empirical result

    A controlled hopper study varied the quality and amount of labeled and unlabeled data. Scores are normalized returns averaged over ten random seeds. UDS benefits most when unlabeled data add coverage or improve the effective behavior policy; it hurts when medium-quality labeled data already provide reasonable coverage and random unlabeled data add little useful information. CDS+UDS improves over UDS in five of seven listed compositions.

    Labeled data / sizeUnlabeled data / sizeCDS+UDSUDSNo SharingCDS (oracle)Sharing All (oracle)
    Expert / 10kRandom / 1M82.178.877.183.386.1
    Expert / 10kMedium / 1M78.164.877.182.564.6
    Expert / 10kExpert / 990k106.1108.477.1106.6112.3
    Medium / 10kRandom / 1M33.29.928.741.238.9
    Medium / 10kExpert / 1M108.9106.728.7111.1107.5
    Random / 10kMedium / 1M63.547.19.692.369.8
    Random / 10kExpert / 1M101.295.99.6110.9102.8

    In a separate size ablation with random labeled data and expert unlabeled data, the scores for (CDS+UDS,UDS,Sharing All)(\mathrm{CDS+UDS},\mathrm{UDS},\mathrm{Sharing\ All}) were (10.1,10.0,71.9)(10.1,10.0,71.9) with 10k unlabeled transitions, (105.8,81.8,96.3)(105.8,81.8,96.3) with 100k, and (102.3,97.0,102.8)(102.3,97.0,102.8) with 1M. Thus, UDS needs enough unlabeled data for sampling-error reduction to outweigh reward bias; CDS+UDS can mitigate this tradeoff at intermediate dataset sizes but does not solve the most data-limited case.

  10. Knowl 10 — Why Zero Labels Can Beat Learned Reward Models

    theoretical result

    For an arbitrary reward-labeling method, let Δr(s,a)\Delta r(s,a) be the error between the true reward and the reward used on unlabeled transitions. The reward-bias contribution has the generic form

    RewardBias⁡(π,πβeff)=11−γ∑s,aΔ ⁣(d^πβeff,d^π)(s,a) Δr(s,a),\operatorname{RewardBias}(\pi,\pi^{\mathrm{eff}}_\beta)=\frac{1}{1-\gamma}\sum_{s,a}\Delta\!\left(\widehat d^{\pi^{\mathrm{eff}}_\beta},\widehat d^\pi\right)(s,a)\,\Delta r(s,a),

    where Δ(d^πβeff,d^π)=d^πβeff−d^π\Delta(\widehat d^{\pi^{\mathrm{eff}}_\beta},\widehat d^\pi)=\widehat d^{\pi^{\mathrm{eff}}_\beta}-\widehat d^\pi. For UDS, Δr(s,a)=(1−f(s,a))r(s,a)≥0\Delta r(s,a)=(1-f(s,a))r(s,a)\geq0 because the unlabeled reward is set to zero. Consequently, state-action pairs visited more often by the learned policy than by the effective behavior policy have negative distribution difference and contribute in the direction of reducing reward-bias suboptimality.

    For a learned reward predictor, Δr(s,a)=r(s,a)−r^ϕ(s,a)\Delta r(s,a)=r(s,a)-\widehat r_\phi(s,a) can have either sign. Policy optimization may exploit state-action pairs whose predicted rewards are overestimated, causing those same pairs to have negative distribution difference and negative reward error; their product then increases the bias rather than canceling it. This provides the paper's explanation for why simple underestimation by UDS can outperform a learned reward model whose errors are selectively exploited, particularly when the labeled dataset is small.

Coverage note — Detailed architectural and dataset-provenance appendices, representation-learning comparisons, and secondary COMBO/ACL baselines were omitted because they support rather than alter the central UDS, reweighting, theory, and primary empirical conclusions.

References

  1. 1.Abbeel, P. and Ng, A. Y. Apprenticeship learning via inverse reinforcement learning. In Proceedings of the twenty-first international conference on Machine learning, pp. 1, 2004.
  2. 2.Achiam, J., Held, D., Tamar, A., and Abbeel, P. Constrained policy optimization. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pp. 22–31. JMLR. org, 2017.
  3. 3.Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W. Hindsight experience replay. arXiv preprint arXiv:1707.01495, 2017.
  4. 4.Cabi, S., Colmenarejo, S. G., Novikov, A., Konyushkova, K., Reed, S., Jeong, R., Zolna, K., Aytar, Y., Budden, D., Vecerik, M., et al. Scaling data-driven robotics with reward sketching and batch reinforcement learning. arXiv preprint arXiv:1909.12200, 2019.
  5. 5.Chebotar, Y., Hausman, K., Lu, Y., Xiao, T., Kalashnikov, D., Varley, J., Irpan, A., Eysenbach, B., Julian, R., Finn, C., and Levine, S. Actionable models: Unsupervised offline reinforcement learning of robotic skills. arXiv preprint arXiv:2104.07749, 2021.
  6. 6.Dasari, S., Ebert, F., Tian, S., Nair, S., Bucher, B., Schmeckpeper, K., Singh, S., Levine, S., and Finn, C. Robonet: Large-scale multi-robot learning, 2020.
  7. 7.de Lima, L. M. and Krohling, R. A. Discovering an aid policy to minimize student evasion using offline reinforcement learning. arXiv preprint arXiv:2104.10258, 2021.
  8. 8.Dorfman, R. and Tamar, A. Offline meta reinforcement learning. arXiv preprint arXiv:2008.02598, 2020.
  9. 9.Duan, Y., Jia, Z., and Wang, M. Minimax-optimal off-policy evaluation with linear function approximation. In International Conference on Machine Learning, pp. 2701–2709. PMLR, 2020.
  10. 10.Ernst, D., Geurts, P., and Wehenkel, L. Tree-based batch mode reinforcement learning. Journal of Machine Learning Research, 6:503–556, 2005.
  11. 11.Eysenbach, B., Geng, X., Levine, S., and Salakhutdinov, R. Rewriting history with inverse rl: Hindsight inference for policy improvement. arXiv preprint arXiv:2002.11089, 2020.
  12. 12.Eysenbach, B., Levine, S., and Salakhutdinov, R. Replacing rewards with examples: Example-based policy search via recursive classification. arXiv preprint arXiv:2103.12656, 2021.
  13. 13.Finn, C., Levine, S., and Abbeel, P. Guided cost learning: Deep inverse optimal control via policy optimization. In International conference on machine learning, pp. 49–58. PMLR, 2016a.
  14. 14.Finn, C., Tan, X. Y., Duan, Y., Darrell, T., Levine, S., and Abbeel, P. Deep spatial autoencoders for visuomotor learning. In 2016 IEEE International Conference on Robotics and Automation (ICRA), pp. 512–519. IEEE, 2016b.
  15. 15.Fu, J., Luo, K., and Levine, S. Learning robust rewards with adversarial inverse reinforcement learning. International Conference on Learning Representations, 2018a.
  16. 16.Fu, J., Singh, A., Ghosh, D., Yang, L., and Levine, S. Variational inverse control with events: A general framework for data-driven reward definition. Conference on Neural Information Processing Systems, 2018b.
  17. 17.Fu, J., Singh, A., Ghosh, D., Yang, L., and Levine, S. Variational inverse control with events: A general framework for data-driven reward definition. arXiv preprint arXiv:1805.11686, 2018c.
  18. 18.Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. D4rl: Datasets for deep data-driven reinforcement learning, 2020.
  19. 19.Fujimoto, S. and Gu, S. S. A minimalist approach to offline reinforcement learning. arXiv preprint arXiv:2106.06860, 2021.
  20. 20.Fujimoto, S., Meger, D., and Precup, D. Off-policy deep reinforcement learning without exploration. arXiv preprint arXiv:1812.02900, 2018.
  21. 21.Ghasemipour, S. K. S., Schuurmans, D., and Gu, S. S. Emaq: Expected-max q-learning operator for simple yet effective offline and online rl. In International Conference on Machine Learning, pp. 3682–3691. PMLR, 2021.
  22. 22.Ho, J. and Ermon, S. Generative adversarial imitation learning. Conference on Neural Information Processing Systems, 2016.
  23. 23.Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R. Way off-policy batch deep reinforcement learning of implicit human preferences in dialog. arXiv preprint arXiv:1907.00456, 2019.
  24. 24.Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al. Scalable deep reinforcement learning for vision-based robotic manipulation. In Conference on Robot Learning, pp. 651–673. PMLR, 2018.
  25. 25.Kalashnikov, D., Varley, J., Chebotar, Y., Swanson, B., Jonschkowski, R., Finn, C., Levine, S., and Hausman, K. Mt-opt: Continuous multi-task robotic reinforcement learning at scale. Conference on Robot Learning (CoRL), 2021.
  26. 26.Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T. Morel: Model-based offline reinforcement learning. arXiv preprint arXiv:2005.05951, 2020.
  27. 27.Konyushkova, K., Zolna, K., Aytar, Y., Novikov, A., Reed, S., Cabi, S., and de Freitas, N. Semi-supervised reward learning for offline reinforcement learning. arXiv preprint arXiv:2012.06899, 2020.
  28. 28.Kostrikov, I., Fergus, R., Tompson, J., and Nachum, O. Offline reinforcement learning with fisher divergence critic regularization. In International Conference on Machine Learning, pp. 5774–5783. PMLR, 2021.
  29. 29.Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S. Stabilizing off-policy q-learning via bootstrapping error reduction. In Advances in Neural Information Processing Systems, pp. 11761–11771, 2019.
  30. 30.Kumar, A., Zhou, A., Tucker, G., and Levine, S. Conservative q-learning for offline reinforcement learning. arXiv preprint arXiv:2006.04779, 2020.
  31. 31.Lange, S., Gabel, T., and Riedmiller, M. A. Batch reinforcement learning. In Reinforcement Learning, volume 12. Springer, 2012.
  32. 32.Laroche, R., Trichelair, P., and Des Combes, R. T. Safe policy improvement with baseline bootstrapping. In International Conference on Machine Learning, pp. 3652–3661. PMLR, 2019.
  33. 33.Lee, B.-J., Lee, J., and Kim, K.-E. Representation balancing offline model-based reinforcement learning. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=QpNz8r_Ri2Y.
  34. 34.Levine, S., Kumar, A., Tucker, G., and Fu, J. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643, 2020.
  35. 35.Li, A. C., Pinto, L., and Abbeel, P. Generalized hindsight for reinforcement learning. arXiv preprint arXiv:2002.11708, 2020.
  36. 36.Li, J., Vuong, Q., Liu, S., Liu, M., Ciosek, K., Ross, K., Christensen, H. I., and Su, H. Multi-task batch reinforcement learning with metric learning. arXiv preprint arXiv:1909.11373, 2019.
  37. 37.Lin, X., Baweja, H. S., and Held, D. Reinforcement learning without ground-truth state. arXiv preprint arXiv:1905.07866, 2019.
  38. 38.Liu, H., Trott, A., Socher, R., and Xiong, C. Competitive experience replay. arXiv preprint arXiv:1902.00528, 2019.
  39. 39.Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E. Provably good batch reinforcement learning without great exploration. arXiv preprint arXiv:2007.08202, 2020.
  40. 40.Mitchell, E., Rafailov, R., Peng, X. B., Levine, S., and Finn, C. Offline meta-reinforcement learning with advantage weighting. In International Conference on Machine Learning, pp. 7780–7791. PMLR, 2021.
  41. 41.Ng, A. Y. and Russell, S. J. Algorithms for inverse reinforcement learning. In Proceedings of the Seventeenth International Conference on Machine Learning, ICML ’00, 2000.
  42. 42.Pomerleau, D. A. Alvinn: an autonomous land vehicle in a neural network. In Proceedings of the 1st International Conference on Neural Information Processing Systems, pp. 305–313, 1988.
  43. 43.Rafailov, R., Yu, T., Rajeswaran, A., and Finn, C. Offline reinforcement learning from images with latent space models. Learning for Decision Making and Control (L4DC), 2021.
  44. 44.Riedmiller, M. Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method. In European Conference on Machine Learning, pp. 317–328. Springer, 2005.
  45. 45.Ross, S. and Bagnell, D. Agnostic system identification for model-based reinforcement learning. In ICML, 2012.
  46. 46.Shortreed, S. M., Laber, E., Lizotte, D. J., Stroup, T. S., Pineau, J., and Murphy, S. A. Informing sequential clinical decision-making through reinforcement learning: an empirical study. Machine learning, 84(1-2):109–136, 2011.
  47. 47.Singh, A., Yang, L., Hartikainen, K., Finn, C., and Levine, S. End-to-end robotic reinforcement learning without reward engineering. arXiv preprint arXiv:1904.07854, 2019.
  48. 48.Singh, A., Yu, A., Yang, J., Zhang, J., Kumar, A., and Levine, S. Cog: Connecting new skills to past experience with offline reinforcement learning. arXiv preprint arXiv:2010.14500, 2020.
  49. 49.Sun, H., Li, Z., Liu, X., Lin, D., and Zhou, B. Policy continuation with hindsight inverse dynamics. arXiv preprint arXiv:1910.14055, 2019.
  50. 50.Swazinna, P., Udluft, S., and Runkler, T. Overcoming model bias for robust offline deep reinforcement learning. arXiv preprint arXiv:2008.05533, 2020.
  51. 51.Wang, L., Zhang, W., He, X., and Zha, H. Supervised reinforcement learning with recurrent neural network for dynamic treatment recommendation. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2018.
  52. 52.Wu, Y., Tucker, G., and Nachum, O. Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361, 2019.
  53. 53.Xie, A., Singh, A., Levine, S., and Finn, C. Few-shot goal inference for visuomotor learning and planning. In Conference on Robot Learning, pp. 40–52. PMLR, 2018.
  54. 54.Xie, A., Ebert, F., Levine, S., and Finn, C. Improvisation through physical understanding: Using novel objects as tools with visual foresight. Robotics: Science and Systems (RSS), 2019.
  55. 55.Yang, M. and Nachum, O. Representation matters: Offline pretraining for sequential decision making. arXiv preprint arXiv:2102.05815, 2021.
  56. 56.Yang, M., Levine, S., and Nachum, O. Trail: Near-optimal imitation learning with suboptimal data. arXiv preprint arXiv:2110.14770, 2021.
  57. 57.Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. In Conference on Robot Learning, pp. 1094–1100. PMLR, 2020a.
  58. 58.Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. In Conference on Robot Learning, pp. 1094–1100. PMLR, 2020b.
  59. 59.Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T. Mopo: Model-based offline policy optimization. arXiv preprint arXiv:2005.13239, 2020c.
  60. 60.Yu, T., Kumar, A., Chebotar, Y., Hausman, K., Levine, S., and Finn, C. Conservative data sharing for multi-task offline reinforcement learning. arXiv preprint arXiv:2109.08128, 2021a.
  61. 61.Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C. Combo: Conservative offline model-based policy optimization. arXiv preprint arXiv:2102.08363, 2021b.
  62. 62.Zhan, X., Xu, H., Zhang, Y., Huo, Y., Zhu, X., Yin, H., and Zheng, Y. Deepthermal: Combustion optimization for thermal power generating units using offline reinforcement learning. arXiv preprint arXiv:2102.11492, 2021.
  63. 63.Zhou, W., Bajracharya, S., and Held, D. Plas: Latent action space for offline reinforcement learning. arXiv preprint arXiv:2011.07213, 2020.
  64. 64.Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K. Maximum entropy inverse reinforcement learning. In Aaai, volume 8, pp. 1433–1438. Chicago, IL, USA, 2008.

Citation

MLA
Yu, T., et al. “How to Leverage Unlabeled Data in Offline Reinforcement Learning”. International Conference on Machine Learning, vol. 162, 2022, pp. 25611–35, https://proceedings.mlr.press/v162/yu22c.html.
APA
Yu, T., Kumar, A., Chebotar, Y., Hausman, K., Finn, C., & Levine, S. (2022). How to Leverage Unlabeled Data in Offline Reinforcement Learning. International Conference on Machine Learning, 162, 25611–25635. https://proceedings.mlr.press/v162/yu22c.html
Chicago
Yu, T., A. Kumar, Y. Chebotar, K. Hausman, C. Finn, and S. Levine. 2022. “How to Leverage Unlabeled Data in Offline Reinforcement Learning”. International Conference on Machine Learning 162: 25611–35. https://proceedings.mlr.press/v162/yu22c.html.
Harvard
Yu, T. et al. (2022) “How to Leverage Unlabeled Data in Offline Reinforcement Learning”, International Conference on Machine Learning. PMLR, pp. 25611–25635. Available at: https://proceedings.mlr.press/v162/yu22c.html.
Vancouver
1. Yu T, Kumar A, Chebotar Y, Hausman K, Finn C, Levine S (2022) How to Leverage Unlabeled Data in Offline Reinforcement Learning. In: International Conference on Machine Learning. PMLR, pp 25611–25635

BibTeX

@InProceedings{pmlr-v162-yu22c,
  title = 	 {How to Leverage Unlabeled Data in Offline Reinforcement Learning},
  author =       {Yu, Tianhe and Kumar, Aviral and Chebotar, Yevgen and Hausman, Karol and Finn, Chelsea and Levine, Sergey},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {25611--25635},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/yu22c/yu22c.pdf},
  url = 	 {https://proceedings.mlr.press/v162/yu22c.html},
  abstract = 	 {Offline reinforcement learning (RL) can learn control policies from static datasets but, like standard RL methods, it requires reward annotations for every transition. In many cases, labeling large datasets with rewards may be costly, especially if those rewards must be provided by human labelers, while collecting diverse unlabeled data might be comparatively inexpensive. How can we best leverage such unlabeled data in offline RL? One natural solution is to learn a reward function from the labeled data and use it to label the unlabeled data. In this paper, we find that, perhaps surprisingly, a much simpler method that simply applies zero rewards to unlabeled data leads to effective data sharing both in theory and in practice, without learning any reward model at all. While this approach might seem strange (and incorrect) at first, we provide extensive theoretical and empirical analysis that illustrates how it trades off reward bias, sample complexity and distributional shift, often leading to good results. We characterize conditions under which this simple strategy is effective, and further show that extending it with a simple reweighting approach can further alleviate the bias introduced by using incorrect reward labels. Our empirical evaluation confirms these findings in simulated robotic locomotion, navigation, and manipulation settings.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/