Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time Guarantees

Siliang ZengChenliang LiAlfredo GarcíaMingyi Hong

article2022NeurIPS51 citations

Develops an efficient single-loop maximum-likelihood inverse reinforcement learning algorithm with finite-time convergence guarantees under nonlinear reward parameterization, enabling accurate reward recovery and superior policy transfer in continuous control tasks.

Listen

Training automated systems by observing expert demonstrations is critical across robotics, autonomous driving, and healthcare. Inverse reinforcement learning aims to recover both the expert's decision policy and the underlying reward function that drives it. While mimicking behavior alone is useful, recovering an accurate reward function is essential for transferring learned behaviors to new tasks or adapting to altered physical environments. Traditional methods suffer from severe computational bottlenecks because they repeatedly calculate optimal policies inside an inner loop while refining rewards in an outer loop. Recent alternatives reduce this computational burden but sacrifice reward recovery accuracy, which causes systems to fail when applied to new conditions.

The article evaluates a single-loop inverse reinforcement learning framework based on maximum likelihood estimation that recovers accurate reward functions and decision policies without nested computational loops. To establish credibility, the authors conduct mathematical convergence proofs and evaluate the method against established benchmarks using high-dimensional robotics control tasks in the MuJoCo simulation environment across six random trials, focusing on data-scarce settings with single expert demonstrations.

The findings show that the proposed algorithm delivers superior performance and strong stability. First, the single-loop design provably converges to high-quality stationary solutions in finite time, requiring a bounded number of iterations even when dealing with complex, non-linear reward functions. Second, in standard robotics imitation benchmarks, the algorithm consistently matches or exceeds state-of-the-art baselines. Third, in transfer learning scenarios where the physical dynamics of the agent change between training and testing, the state-only reward version significantly outperforms competing models. For instance, in an altered robotics simulation, the recovered reward enabled downstream policy performance of 187.69 compared to scores of 156.45 or lower from alternative methods, while competing single-level techniques completely failed with negative cumulative scores.

These results demonstrate that organizations can reduce the computational costs and instability of training imitation learning systems without sacrificing the generalizability of the learned reward functions. Decoupling the reward function from specific transition dynamics lowers the risk of deployment failure when deploying robotic systems into environments that differ from demonstration settings. Decision-makers in automated control and robotics should consider adopting this single-loop maximum likelihood approach for imitation tasks requiring environmental transfer.

Future efforts should evaluate this framework in offline settings, as the current method relies on online environment interactions during training. Practitioners should also exercise caution to ensure expert demonstration datasets are properly curated, as inverse reinforcement learning will systematically propagate any negative biases or suboptimal behaviors present in the source training data.

arXiv: 2210.01282
  • Paper: Algorithms for Inverse Reinforcement Learning, Andrew Y. Ng et al. (2000). This seminal paper introduces the fundamental problem formulation of inverse reinforcement learning via iterative reward fitting that the source directly seeks to make single-loop and computationally tractable.
  • Paper: Maximum Entropy Inverse Reinforcement Learning, Brian D. Ziebart et al. (2008). It formulates the probabilistic maximum-entropy and maximum-likelihood framework for inverse reinforcement learning that underlies the source paper's objective.
  • Paper: Apprenticeship learning via inverse reinforcement learning, Pieter Abbeel et al. (2004). It provides the foundational framework and convergence analysis for apprenticeship learning via inverse reinforcement learning upon which modern IRL algorithms build.
  • Paper: Generative Adversarial Imitation Learning, Jonathan Ho et al. (2016). It establishes the adversarial, occupancy-measure matching paradigm for scaling imitation learning and highlights the nested-loop computational bottlenecks the source aims to overcome.
  • Paper: Policy Gradient Methods for Reinforcement Learning with Function Approximation, Richard S. Sutton et al. (1999). It introduces the foundational policy gradient theorem essential for optimizing parameterized policies within modern likelihood-based and single-loop RL methods.
  • Paper: Trust Region Policy Optimization, John Schulman et al. (2015). It presents the monotonic policy improvement guarantees and trust-region optimization concepts used to analyze convergence in continuous robotic control.
  • Paper: Transfer Learning for Reinforcement Learning Domains: A Survey, Matthew E. Taylor et al. (2009). It provides a comprehensive taxonomy of transfer learning in reinforcement learning, explaining the challenges of domain and dynamics shifts central to the source paper's evaluation.
  • Paper: Efficient Diffusion Policies For Offline Reinforcement Learning, Bingyi Kang et al. (2023). Extends the principle of efficient maximum likelihood decision-making to expressive diffusion policies for offline reinforcement learning.
  • Paper: Actor Prioritized Experience Replay, Baturay Saglam et al. (2023). Addresses gradient instability and sample efficiency in continuous control actor-critic methods through prioritized replay mechanisms.
  • Paper: Flow Q-Learning, Seohong Park et al. (2025). Applies decoupled policy extraction and generative modeling to offline reinforcement learning to scale beyond traditional nested optimization.
  • Paper: On Training in Imagination, Nadav Timor et al. (2026). Analyzes the theoretical impact of reward-model and dynamics-model errors when policies are optimized using learned reward representations.
Cover for Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time Guarantees

Abstract

Inverse reinforcement learning (IRL) aims to recover the reward function and the associated optimal policy that best fits observed sequences of states and actions implemented by an expert. Many algorithms for IRL have an inherently nested structure: the inner loop finds the optimal policy given parametrized rewards while the outer loop updates the estimates towards optimizing a measure of fit. For high dimensional environments such nested-loop structure entails a significant computational burden. To reduce the computational burden of a nested loop, novel methods such as SQIL [1] and IQ-Learn [2] emphasize policy estimation at the expense of reward estimation accuracy. However, without accurate estimated rewards, it is not possible to do counterfactual analysis such as predicting the optimal policy under different environment dynamics and/or learning new tasks. In this paper we develop a novel single-loop algorithm for IRL that does not compromise reward estimation accuracy. In the proposed algorithm, each policy improvement step is followed by a stochastic gradient step for likelihood maximization. We show that the proposed algorithm provably converges to a stationary solution with a finite-time guarantee. If the reward is parameterized linearly, we show the identified solution corresponds to the solution of the maximum entropy IRL problem. Finally, by using robotics control problems in MuJoCo and their transfer settings, we show that the proposed algorithm achieves superior performance compared with other IRL and imitation learning benchmarks.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Dynamic Discrete Choice Model
  • 2.2 Maximum Likelihood Inverse Reinforcement Learning (ML-IRL)
  • 2.3 Computational Effort and Estimation Quality of Existing Algorithms
  • 3 Problem Approximation in High Dimensional State Space
  • 4 The Proposed Algorithm
  • 5 Theoretical Analysis
  • 6 The Linearly Parameterized Reward Function Case
  • 7 The Case with State-only Dependent Rewards
  • 8 Extension to the Offline Setting
  • 9 Testbed
  • 10 Conclusions
  • 11 Auxiliary Lemmas
  • 12 Proof of Theorem
  • 12.1 Proof of Relation ()
  • 12.2 Proof of relation ()
  • 13 Proof of Theorem
  • References
  • 14 Supplementary Experiment
  • 15 Proof of Lemma
  • 16 Proof of Lemma
  • 17 Proof of Lemma
  • 17.1 Proof of Inequality ()
  • 17.2 Proof of Inequality ()
  • 18 Proof of Lemma
  • 19 Proof of Lemma
  • 20 Proof of Lemma
  • 20.1 Proof of Inequality ()
  • 20.2 Proof of Inequality ()

Knowls

  1. Knowl 1 — Maximum Log-Likelihood Inverse Reinforcement Learning Formulation

    model/method

    Consider a Markov Decision Process (MDP) defined by (S,A,P,η,r,γ)(\mathcal{S}, \mathcal{A}, P, \eta, r, \gamma), where S\mathcal{S} is the state space, A\mathcal{A} is the action space, P(s′∣s,a)P(s' \mid s, a) is the transition probability kernel, η(⋅)\eta(\cdot) is the initial state distribution, r(s,a;θ)r(s, a; \theta) is a parameterized reward function with parameter θ∈Rd\theta \in \mathbb{R}^d, and γ∈(0,1)\gamma \in (0, 1) is the discount factor.

    Given expert trajectories τ={(st,at)}t≥0\tau = \{(s_t, a_t)\}_{t \ge 0} generated by an expert policy πE\pi^E, the Maximum Log-Likelihood Inverse Reinforcement Learning (ML-IRL) problem is formulated as the bi-level optimization problem:

    max⁡θL(θ):=Eτ∼πE[∑t=0∞γtlog⁡πθ(at∣st)]\max_\theta L(\theta) := \mathbb{E}_{\tau \sim \pi^E} \left[ \sum_{t=0}^\infty \gamma^t \log \pi_\theta(a_t \mid s_t) \right]

    subject to the lower-level entropy-regularized MDP policy optimization:

    πθ:=arg⁡max⁡πEτ∼π[∑t=0∞γt(r(st,at;θ)+H(π(⋅∣st)))]\pi_\theta := \arg\max_\pi \mathbb{E}_{\tau \sim \pi} \left[ \sum_{t=0}^\infty \gamma^t \left( r(s_t, a_t; \theta) + \mathcal{H}(\pi(\cdot \mid s_t)) \right) \right]

    where H(π(⋅∣s)):=−∑a∈Aπ(a∣s)log⁡π(a∣s)\mathcal{H}(\pi(\cdot \mid s)) := -\sum_{a \in \mathcal{A}} \pi(a \mid s) \log \pi(a \mid s) is the Shannon entropy of policy π\pi at state ss. The upper-level objective optimizes the reward parameter θ\theta to maximize the likelihood of the expert trajectories under the induced optimal policy πθ\pi_\theta, while the entropy regularization in the lower level guarantees the uniqueness of the optimal policy πθ\pi_\theta for any fixed θ\theta.

  2. Knowl 2 — Strong Duality between MaxEnt-IRL and ML-IRL under Linear Reward Parameterization

    theoretical result

    Let an MDP be defined by state space S\mathcal{S}, action space A\mathcal{A}, transition distribution P(s′∣s,a)P(s' \mid s, a), initial state distribution η\eta, and discount factor γ∈(0,1)\gamma \in (0, 1). Suppose the reward function is parameterized linearly with a feature mapping ϕ:S×A→Rd\phi: \mathcal{S} \times \mathcal{A} \to \mathbb{R}^d such that r(s,a;θ)=ϕ(s,a)Tθr(s, a; \theta) = \phi(s, a)^T \theta for all s∈Ss \in \mathcal{S} and a∈Aa \in \mathcal{A}.

    Consider the Maximum Entropy Inverse Reinforcement Learning (MaxEnt-IRL) problem:

    max⁡πH(π):=Eτ∼π[−∑t=0∞γtlog⁡π(at∣st)]\max_\pi H(\pi) := \mathbb{E}_{\tau \sim \pi} \left[ -\sum_{t=0}^\infty \gamma^t \log \pi(a_t \mid s_t) \right]

    subject to the feature expectation matching constraint:

    Eτ∼π[∑t=0∞γtϕ(st,at)]=Eτ∼πE[∑t=0∞γtϕ(st,at)]\mathbb{E}_{\tau \sim \pi} \left[ \sum_{t=0}^\infty \gamma^t \phi(s_t, a_t) \right] = \mathbb{E}_{\tau \sim \pi^E} \left[ \sum_{t=0}^\infty \gamma^t \phi(s_t, a_t) \right]

    Let θ\theta denote the Lagrange dual variables corresponding to this constraint. Then:

    1. The Maximum Log-Likelihood IRL (ML-IRL) formulation:
    max⁡θL(θ):=Eτ∼πE[∑t=0∞γtlog⁡πθ(at∣st)]\max_\theta L(\theta) := \mathbb{E}_{\tau \sim \pi^E} \left[ \sum_{t=0}^\infty \gamma^t \log \pi_\theta(a_t \mid s_t) \right]

    with πθ:=arg⁡max⁡πEτ∼π[∑t=0∞γt(ϕ(st,at)Tθ+H(π(⋅∣st)))]\pi_\theta := \arg\max_\pi \mathbb{E}_{\tau \sim \pi} \left[ \sum_{t=0}^\infty \gamma^t (\phi(s_t, a_t)^T \theta + \mathcal{H}(\pi(\cdot \mid s_t))) \right], is the exact Lagrangian dual problem of MaxEnt-IRL. 2. Strong duality holds:

    L(θ∗)=H(π∗)L(\theta^*) = H(\pi^*)

    where θ∗\theta^* and π∗\pi^* are the global optimal solutions to the ML-IRL and MaxEnt-IRL problems, respectively. Consequently, under linear reward parameterization, ML-IRL is a concave optimization problem, and any stationary point of L(θ)L(\theta) is a global optimal reward estimator.

  3. Knowl 3 — Exact Gradient of the Log-Likelihood Function in ML-IRL

    theoretical result

    In the Maximum Log-Likelihood Inverse Reinforcement Learning framework with parameterized reward r(s,a;θ)r(s, a; \theta), expert policy πE\pi^E, and entropy-regularized optimal policy πθ=arg⁡max⁡πEτ∼π[∑t=0∞γt(r(st,at;θ)+H(π(⋅∣st)))]\pi_\theta = \arg\max_\pi \mathbb{E}_{\tau \sim \pi} [ \sum_{t=0}^\infty \gamma^t (r(s_t, a_t; \theta) + \mathcal{H}(\pi(\cdot \mid s_t))) ], the objective function is:

    L(θ)=Eτ∼πE[∑t=0∞γtlog⁡πθ(at∣st)]L(\theta) = \mathbb{E}_{\tau \sim \pi^E} \left[ \sum_{t=0}^\infty \gamma^t \log \pi_\theta(a_t \mid s_t) \right]

    The gradient of L(θ)L(\theta) with respect to θ\theta is given by the difference between the expected discounted cumulative reward gradient under the expert policy and under the induced optimal policy:

    ∇L(θ)=Eτ∼πE[∑t≥0γt∇θr(st,at;θ)]−Eτ∼πθ[∑t≥0γt∇θr(st,at;θ)]\nabla L(\theta) = \mathbb{E}_{\tau \sim \pi^E} \left[ \sum_{t \ge 0} \gamma^t \nabla_\theta r(s_t, a_t; \theta) \right] - \mathbb{E}_{\tau \sim \pi_\theta} \left[ \sum_{t \ge 0} \gamma^t \nabla_\theta r(s_t, a_t; \theta) \right]

    where τ={(st,at)}t≥0\tau = \{(s_t, a_t)\}_{t \ge 0} denotes a state-action trajectory.

  4. Knowl 4 — Single-Loop Maximum-Likelihood Inverse Reinforcement Learning Algorithm

    algorithm

    The Maximum-Likelihood Inverse Reinforcement Learning (ML-IRL) algorithm avoids nested-loop computation by alternating single policy improvement steps and stochastic gradient reward update steps using two-timescale stochastic approximation.

    Input: Initial reward parameter θ0\theta_0, initial policy π0\pi_0, reward stepsize α\alpha, total iterations KK.
    for k=0,1,…,K−1k = 0, 1, \dots, K-1 do
        Compute soft Q-function Qrθk,πksoft(s,a)=r(s,a;θk)+γEs′∼P(⋅∣s,a)[Vrθk,πksoft(s′)]Q^{\text{soft}}_{r_{\theta_k}, \pi_k}(s, a) = r(s, a; \theta_k) + \gamma \mathbb{E}_{s' \sim P(\cdot \mid s, a)} [V^{\text{soft}}_{r_{\theta_k}, \pi_k}(s')]
        Update policy: πk+1(a∣s)∝exp⁡(Qrθk,πksoft(s,a))\pi_{k+1}(a \mid s) \propto \exp(Q^{\text{soft}}_{r_{\theta_k}, \pi_k}(s, a)) for all s∈S,a∈As \in \mathcal{S}, a \in \mathcal{A}
        Sample an expert trajectory τkE={(stE,atE)}t≥0\tau_k^E = \{(s_t^E, a_t^E)\}_{t \ge 0} from the expert dataset
        Sample a trajectory τkA={(stA,atA)}t≥0\tau_k^A = \{(s_t^A, a_t^A)\}_{t \ge 0} from the current policy πk+1\pi_{k+1}
        Compute gradient estimator: gk=h(θk;τkE)−h(θk;τkA)g_k = h(\theta_k; \tau_k^E) - h(\theta_k; \tau_k^A), where h(θ;τ)=∑t≥0γt∇θr(st,at;θ)h(\theta; \tau) = \sum_{t \ge 0} \gamma^t \nabla_\theta r(s_t, a_t; \theta)
        Update reward parameter: θk+1=θk+αgk\theta_{k+1} = \theta_k + \alpha g_k
    end for
    return Final reward parameter θK\theta_K and policy πK\pi_K.

    In continuous or high-dimensional implementations, the exact soft policy evaluation and improvement can be replaced with mini-batch updates via Soft Actor-Critic (SAC) or Soft Q-learning.

  5. Knowl 5 — Regularity and Ergodicity Assumptions for ML-IRL Convergence Analysis

    assumption

    The non-asymptotic theoretical analysis of the single-loop ML-IRL algorithm relies on two structural conditions on the MDP and the parameterized reward function:

    1. Ergodicity and Geometric Mixing: For any policy π\pi, the induced Markov chain with transition probability kernel PP on state space S\mathcal{S} is irreducible and aperiodic. Furthermore, there exist constants κ>0\kappa > 0 and ρ∈(0,1)\rho \in (0, 1) such that for all states s∈Ss \in \mathcal{S} and time steps t≥0t \ge 0:
    sup⁡s∈S∥P(st∈⋅∣s0=s,π)−μπ(⋅)∥TV≤κρt\sup_{s \in \mathcal{S}} \|P(s_t \in \cdot \mid s_0 = s, \pi) - \mu_\pi(\cdot)\|_{\text{TV}} \le \kappa \rho^t

    where ∥⋅∥TV\|\cdot\|_{\text{TV}} is the total variation norm and μπ\mu_\pi denotes the stationary state distribution under policy π\pi.

    1. Reward Function Smoothness and Boundedness: For all states s∈Ss \in \mathcal{S}, actions a∈Aa \in \mathcal{A}, and reward parameters θ,θ1,θ2∈Rd\theta, \theta_1, \theta_2 \in \mathbb{R}^d, the gradient of the reward function ∇θr(s,a;θ)\nabla_\theta r(s, a; \theta) satisfies:
    ∥∇θr(s,a;θ)∥≤Lr\|\nabla_\theta r(s, a; \theta)\| \le L_r ∥∇θr(s,a;θ1)−∇θr(s,a;θ2)∥≤Lg∥θ1−θ2∥\|\nabla_\theta r(s, a; \theta_1) - \nabla_\theta r(s, a; \theta_2)\| \le L_g \|\theta_1 - \theta_2\|

    where Lr>0L_r > 0 and Lg>0L_g > 0 are finite constants.

  6. Knowl 6 — Finite-Time Convergence Guarantees for Single-Loop ML-IRL

    theoretical result

    Suppose the ergodicity and reward smoothness assumptions hold. Let the ML-IRL algorithm be run for KK iterations with the reward parameter stepsize chosen as:

    α:=α0Kσ\alpha := \frac{\alpha_0}{K^\sigma}

    for constants α0>0\alpha_0 > 0 and σ∈(0,1)\sigma \in (0, 1).

    Then the policy tracking error and the expected squared gradient of the log-likelihood function satisfy the non-asymptotic bounds:

    1K∑k=0K−1E[∥log⁡πk+1−log⁡πθk∥∞]=O(K−1)+O(K−σ)\frac{1}{K} \sum_{k=0}^{K-1} \mathbb{E} \left[ \|\log \pi_{k+1} - \log \pi_{\theta_k}\|_{\infty} \right] = \mathcal{O}(K^{-1}) + \mathcal{O}(K^{-\sigma}) 1K∑k=0K−1E[∥∇L(θk)∥2]=O(K−σ)+O(K−1+σ)+O(K−1)\frac{1}{K} \sum_{k=0}^{K-1} \mathbb{E} \left[ \|\nabla L(\theta_k)\|^2 \right] = \mathcal{O}(K^{-\sigma}) + \mathcal{O}(K^{-1+\sigma}) + \mathcal{O}(K^{-1})

    where ∥log⁡πk+1−log⁡πθk∥∞:=max⁡s∈S,a∈A∣log⁡πk+1(a∣s)−log⁡πθk(a∣s)∣\|\log \pi_{k+1} - \log \pi_{\theta_k}\|_{\infty} := \max_{s \in \mathcal{S}, a \in \mathcal{A}} |\log \pi_{k+1}(a \mid s) - \log \pi_{\theta_k}(a \mid s)|.

    In particular, setting σ=1/2\sigma = 1/2 yields an optimal convergence rate of O(K−1/2)\mathcal{O}(K^{-1/2}) for both quantities. Consequently, the algorithm requires O(ϵ−2)\mathcal{O}(\epsilon^{-2}) policy improvement and reward update steps to find an ϵ\epsilon-approximate stationary solution under nonlinear reward parameterizations.

  7. Knowl 7 — Equivalent Value-Gap Formulation for State-Only Reward ML-IRL

    theoretical result

    When the parameterized reward function depends solely on the state, r(s;θ)r(s; \theta), and expert trajectories consist only of visited state sequences generated by an expert policy πE\pi^E, the ML-IRL objective:

    max⁡θEτ∼πE[∑t=0∞γtlog⁡πθ(at∣st)]\max_\theta \mathbb{E}_{\tau \sim \pi^E} \left[ \sum_{t=0}^\infty \gamma^t \log \pi_\theta(a_t \mid s_t) \right]

    is mathematically equivalent to minimizing the expected soft value gap between the optimal induced policy πθ\pi_\theta and the expert policy πE\pi^E:

    min⁡θ(Es0∼η(⋅)[Vrθ,πθsoft(s0)]−Es0∼η(⋅)[Vrθ,πEsoft(s0)])\min_\theta \left( \mathbb{E}_{s_0 \sim \eta(\cdot)} \left[ V^{\text{soft}}_{r_\theta, \pi_\theta}(s_0) \right] - \mathbb{E}_{s_0 \sim \eta(\cdot)} \left[ V^{\text{soft}}_{r_\theta, \pi^E}(s_0) \right] \right)

    subject to:

    πθ:=arg⁡max⁡πEπ[∑t=0∞γt(r(st;θ)+H(π(⋅∣st)))]\pi_\theta := \arg\max_\pi \mathbb{E}_\pi \left[ \sum_{t=0}^\infty \gamma^t \left( r(s_t; \theta) + \mathcal{H}(\pi(\cdot \mid s_t)) \right) \right]

    where Vrθ,πsoft(s)=Eπ[∑t=0∞γt(r(st;θ)+H(π(⋅∣st)))∣s0=s]V^{\text{soft}}_{r_\theta, \pi}(s) = \mathbb{E}_\pi [ \sum_{t=0}^\infty \gamma^t (r(s_t; \theta) + \mathcal{H}(\pi(\cdot \mid s_t))) \mid s_0 = s ], and η(⋅)\eta(\cdot) is the initial state distribution.

    This equivalence enables ML-IRL to operate in state-only demonstration settings by computing the stochastic gradient estimator using only the visited states: gk=∑t≥0γt∇θr(stE;θk)−∑t≥0γt∇θr(stA;θk)g_k = \sum_{t \ge 0} \gamma^t \nabla_\theta r(s_t^E; \theta_k) - \sum_{t \ge 0} \gamma^t \nabla_\theta r(s_t^A; \theta_k).

  8. Knowl 8 — Imitation Performance of ML-IRL on MuJoCo Benchmarks with Single Trajectory

    empirical result

    In continuous control robotics tasks from MuJoCo, ML-IRL was evaluated against Behavioral Cloning (BC), Generative Adversarial Imitation Learning (GAIL), Inverse Soft-Q Learning (IQ-Learn), and ff-IRL under a limited-data regime where the expert dataset contains only a single demonstration trajectory. All algorithms used Soft Actor-Critic (SAC) as the underlying RL solver. Cumulative returns are averaged over 6 random seeds (evaluating 20 trajectories per seed after convergence):

    Task BC GAIL IQ-Learn ff-IRL ML-IRL (State-Only) ML-IRL (State-Action) Expert
    Hopper 20.49 2815.59 2981.01 3074.55 3089.79 3121.68 3592.63
    Half-Cheetah -1.87 3301.52 4175.88 4375.88 4472.85 4086.92 5098.30
    Walker -14.01 1112.79 3961.42 4464.20 4380.17 4504.88 5344.21
    Ant 760.46 1154.27 4362.90 4571.71 4675.34 4984.34 5926.18
    Humanoid 78.48 3016.40 5227.10 5243.90 5390.31 5240.57 5351.08

    Behavioral Cloning fails under single-trajectory demonstrations due to error accumulation. ML-IRL (both state-action and state-only variants) outperforms GAIL, IQ-Learn, and ff-IRL across all tasks, matching or approaching expert performance.

  9. Knowl 9 — Transfer Learning Performance across Changing Dynamics in Ant Environments

    empirical result

    To evaluate reward transferability across changing environment dynamics, algorithms were trained on a single expert trajectory collected in a source environment (Custom-Ant) and tested on a target environment with altered dynamics (Disabled-Ant). Two settings were evaluated:

    1. Data Transfer: Training IRL agents directly in Disabled-Ant using expert demonstrations from Custom-Ant.
    2. Reward Transfer: Inferring the reward function in Custom-Ant, then training a Soft Actor-Critic (SAC) policy in Disabled-Ant using the recovered reward function.
    Setting IQ-Learn AIRL ff-IRL ML-IRL (State-Only) Ground-Truth SAC
    Data Transfer -11.78 -5.39 188.85 221.51 320.15
    Reward Transfer -1.04 130.30 156.45 187.69 320.15

    IQ-Learn and AIRL struggle in transfer settings because their implicit reward representations are entangled with source environment dynamics. ML-IRL (State-Only) achieves the highest scores in both settings by isolating state-based preference signals that remain valid under altered transition dynamics.

  10. Knowl 10 — Requirement of Online Environment Interaction in ML-IRL

    limitation

    A primary limitation of the proposed ML-IRL algorithm and its non-asymptotic convergence guarantees is the requirement for online environment interaction. At each iteration, the algorithm samples on-policy trajectories τkA\tau_k^A from the current policy πk+1\pi_{k+1} to evaluate the stochastic gradient estimator gk=h(θk;τkE)−h(θk;τkA)g_k = h(\theta_k; \tau_k^E) - h(\theta_k; \tau_k^A) and to perform soft policy updates. The current framework and theoretical guarantees do not apply to purely offline IRL settings where no additional environment interactions are allowed.

Coverage note — Omitted materials are the detailed step-by-step mathematical proofs in Appendices D, E, F, G, and H, and the specific neural network hyperparameter tuning tables in Appendix B, as they represent supporting derivations and standard implementation configurations.

References

  1. 1.S. Reddy, A. D. Dragan, and S. Levine, ‘SQIL: Imitation learning via reinforcement learning with sparse rewards,’ in International Conference on Learning Representations, 2020. [Online]. Available: https://openreview.net/forum?id=S1xKd24twB
  2. 2.D. Garg, S. Chakraborty, C. Cundy, J. Song, and S. Ermon, ‘Iq-learn: Inverse soft-q learning for imitation,’ Advances in Neural Information Processing Systems, vol. 34, 2021.
  3. 3.T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, and J. Peters, An Algorithmic Perspective on Imitation Learning, ser. Foundations and Trends in Robotics, 2018, vol. 7.
  4. 4.K. Kim, S. Garg, K. Shiragur, and S. Ermon, ‘Reward identification in inverse reinforcement learning,’ in International Conference on Machine Learning. PMLR, 2021, pp. 5496–5505.
  5. 5.B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey et al., ‘Maximum entropy inverse reinforcement learning.’ in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 8. Chicago, IL, USA, 2008, pp. 1433–1438.
  6. 6.B. D. Ziebart, J. A. Bagnell, and A. K. Dey, ‘Modeling interaction via the principle of maximum causal entropy,’ in International Conference on Machine Learning, 2010.
  7. 7.M. Wulfmeier, P. Ondruska, and I. Posner, ‘Maximum entropy deep inverse reinforcement learning,’ arXiv preprint arXiv:1507.04888, 2015.
  8. 8.C. Finn, P. Christiano, P. Abbeel, and S. Levine, ‘A connection between generative adversarial networks, inverse reinforcement learning, and energy-based models,’ arXiv preprint arXiv:1611.03852, 2016.
  9. 9.C. Finn, S. Levine, and P. Abbeel, ‘Guided cost learning: Deep inverse optimal control via policy optimization,’ in International Conference on Machine Learning. PMLR, 2016, pp. 49–58.
  10. 10.J. Ho and S. Ermon, ‘Generative adversarial imitation learning,’ Advances in Neural Information Processing Systems, vol. 29, 2016.
  11. 11.J. Fu, K. Luo, and S. Levine, ‘Learning robust rewards with adversarial inverse reinforcement learning,’ arXiv preprint arXiv:1710.11248, 2017.
  12. 12.M. Chen, Y. Wang, T. Liu, Z. Yang, X. Li, Z. Wang, and T. Zhao, ‘On computation and generalization of generative adversarial imitation learning,’ arXiv preprint arXiv:2001.02792, 2020.
  13. 13.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, ‘Generative adversarial nets,’ Advances in Neural Information Processing Systems, vol. 27, 2014.
  14. 14.T. Ni, H. Sikchi, Y. Wang, T. Gupta, L. Lee, and B. Eysenbach, ‘f-irl: Inverse reinforcement learning via state marginal matching,’ arXiv preprint arXiv:2011.04709, 2020.
  15. 15.K. Kurach, M. Lucic, X. Zhai, M. Michalski, and S. Gelly, ‘The gan landscape: Losses, architectures, regularization, and normalization,’ 2018.
  16. 16.I. Kostrikov, K. K. Agrawal, D. Dwibedi, S. Levine, and J. Tompson, ‘Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning,’ in International Conference on Learning Representations, 2019. [Online]. Available: https://openreview.net/forum?id=Hk4fpoA5Km
  17. 17.E. T. Jaynes, ‘Information theory and statistical mechanics,’ Physical review, vol. 106, no. 4, p. 620, 1957.
  18. 18.B. D. Ziebart, J. A. Bagnell, and A. K. Dey, ‘The principle of maximum causal entropy for estimating interacting processes,’ IEEE Transactions on Information Theory, vol. 59, no. 4, pp. 1966–1980, 2013.
  19. 19.M. Bloem and N. Bambos, ‘Infinite time horizon maximum causal entropy inverse reinforcement learning,’ in 53rd IEEE conference on decision and control. IEEE, 2014, pp. 4911–4916.
  20. 20.Z. Zhou, M. Bloem, and N. Bambos, ‘Infinite time horizon maximum causal entropy inverse reinforcement learning,’ IEEE Transactions on Automatic Control, vol. 63, no. 9, pp. 2787–2802, 2017.
  21. 21.T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, ‘Reinforcement learning with deep energy-based policies,’ in International Conference on Machine Learning. PMLR, 2017, pp. 1352–1361.
  22. 22.T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, ‘Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,’ in International Conference on Machine Learning. PMLR, 2018, pp. 1861–1870.
  23. 23.M. Hong, H.-T. Wai, Z. Wang, and Z. Yang, ‘A two-timescale framework for bilevel optimization: Complexity analysis and application to actor-critic,’ arXiv preprint arXiv:2007.05170, 2020.
  24. 24.K. Ji, J. Yang, and Y. Liang, ‘Bilevel optimization: Convergence analysis and enhanced design,’ in International Conference on Machine Learning. PMLR, 2021, pp. 4882–4892.
  25. 25.P. Khanduri, S. Zeng, M. Hong, H.-T. Wai, Z. Wang, and Z. Yang, ‘A near-optimal algorithm for stochastic bilevel optimization via double-momentum,’ Advances in Neural Information Processing Systems, vol. 34, 2021.
  26. 26.V. Jain, P. Doshi, and B. Banerjee, ‘Model-free irl using maximum likelihood estimation,’ in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, no. 01, 2019, pp. 3951–3958.
  27. 27.N. Sanghvi, S. Usami, M. Sharma, J. Groeger, and K. Kitani, ‘Inverse reinforcement learning with explicit policy estimates,’ in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 11, 2021, pp. 9472–9480.
  28. 28.S. Cen, C. Cheng, Y. Chen, Y. Wei, and Y. Chi, ‘Fast global convergence of natural policy gradient methods with entropy regularization,’ Operations Research, 2021.
  29. 29.S. Cayci, N. He, and R. Srikant, ‘Linear convergence of entropy-regularized natural policy gradient with linear function approximation,’ arXiv preprint arXiv:2106.04096, 2021.
  30. 30.O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans, ‘Bridging the gap between value and policy based reinforcement learning,’ Advances in Neural Information Processing Systems, vol. 30, 2017.
  31. 31.V. Konda and J. Tsitsiklis, ‘Actor-critic algorithms,’ Advances in Neural Information Processing Systems, vol. 12, 1999.
  32. 32.Y. F. Wu, W. Zhang, P. Xu, and Q. Gu, ‘A finite-time analysis of two time-scale actor-critic methods,’ Advances in Neural Information Processing Systems, vol. 33, pp. 17 617–17 628, 2020.
  33. 33.V. S. Borkar, ‘Stochastic approximation with two time scales,’ Systems & Control Letters, vol. 29, no. 5, pp. 291–294, 1997.
  34. 34.J. Bhandari, D. Russo, and R. Singal, ‘A finite time analysis of temporal difference learning with linear function approximation,’ in Conference on learning theory. PMLR, 2018, pp. 1691–1692.
  35. 35.S. Zou, T. Xu, and Y. Liang, ‘Finite-sample analysis for sarsa with linear function approximation,’ Advances in Neural Information Processing Systems, vol. 32, 2019.
  36. 36.C. Jin, P. Netrapalli, and M. Jordan, ‘What is local optimality in nonconvex-nonconcave minimax optimization?’ in International Conference on Machine Learning. PMLR, 2020, pp. 4880–4889.
  37. 37.Z. Guan, T. Xu, and Y. Liang, ‘When will generative adversarial imitation learning algorithms attain global convergence,’ in International Conference on Artificial Intelligence and Statistics. PMLR, 2021, pp. 1117–1125.
  38. 38.T. Chen, Y. Sun, and W. Yin, ‘Closing the gap: Tighter analysis of alternating stochastic gradient methods for bilevel problems,’ Advances in Neural Information Processing Systems, vol. 34, pp. 25 294–25 307, 2021.
  39. 39.C. Yu, J. Liu, S. Nemati, and G. Yin, ‘Reinforcement learning in healthcare: A survey,’ ACM Computing Surveys (CSUR), vol. 55, no. 1, pp. 1–36, 2021.
  40. 40.B. R. Kiran, I. Sobh, V. Talpaert, P. Mannion, A. A. Al Sallab, S. Yogamani, and P. Pérez, ‘Deep reinforcement learning for autonomous driving: A survey,’ IEEE Transactions on Intelligent Transportation Systems, 2021.
  41. 41.T. Gangwani and J. Peng, ‘State-only imitation with transition dynamics mismatch,’ in International Conference on Learning Representations, 2020. [Online]. Available: https://openreview.net/forum?id=HJgLLyrYwB
  42. 42.L. Viano, Y.-T. Huang, P. Kamalaruban, A. Weller, and V. Cevher, ‘Robust inverse reinforcement learning under transition dynamics mismatch,’ Advances in Neural Information Processing Systems, vol. 34, 2021.
  43. 43.F. Torabi, G. Warnell, and P. Stone, ‘Generative adversarial imitation from observation,’ arXiv preprint arXiv:1807.06158, 2018.
  44. 44.E. Todorov, T. Erez, and Y. Tassa, ‘Mujoco: A physics engine for model-based control,’ in 2012 IEEE/RSJ international conference on intelligent robots and systems. IEEE, 2012, pp. 5026–5033.
  45. 45.D. A. Pomerleau, ‘ALVINN: An autonomous land vehicle in a neural network,’ Advances in Neural Information Processing Systems, vol. 1, 1988.
  46. 46.F. Liu, Z. Ling, T. Mu, and H. Su, ‘State alignment-based imitation learning,’ in International Conference on Learning Representations, 2020. [Online]. Available: https://openreview.net/forum?id=rylrdxHFDr
  47. 47.A. Kamoutsi, G. Banjac, and J. Lygeros, ‘Efficient performance bounds for primal-dual reinforcement learning from demonstrations,’ in International Conference on Machine Learning. PMLR, 2021, pp. 5257–5268.
  48. 48.H. Cao, S. Cohen, and L. Szpruch, ‘Identifiability in inverse reinforcement learning,’ Advances in Neural Information Processing Systems, vol. 34, 2021.
  49. 49.R. Wang, C. Ciliberto, P. V. Amadori, and Y. Demiris, ‘Random expert distillation: Imitation learning via expert policy support estimation,’ in International Conference on Machine Learning. PMLR, 2019, pp. 6536–6544.
  50. 50.D. S. Brown, W. Goo, and S. Niekum, ‘Better-than-demonstrator imitation learning via automatically-ranked demonstrations,’ in Conference on robot learning. PMLR, 2020, pp. 330–359.
  51. 51.P. Barde, J. Roy, W. Jeon, J. Pineau, C. Pal, and D. Nowrouzezahrai, ‘Adversarial soft advantage fitting: Imitation learning without policy optimization,’ Advances in Neural Information Processing Systems, vol. 33, pp. 12 334–12 344, 2020.
  52. 52.H. Kretzschmar, M. Spies, C. Sprunk, and W. Burgard, ‘Socially compliant mobile robot navigation via inverse reinforcement learning,’ The International Journal of Robotics Research, vol. 35, no. 11, pp. 1289–1307, 2016.
  53. 53.C. Xia and A. El Kamel, ‘Neural inverse reinforcement learning in autonomous navigation,’ Robotics and Autonomous Systems, vol. 84, pp. 1–14, 2016.
  54. 54.J. Rust and C. Phelan, ‘How social security and medicare affect retirement behavior in a world of incomplete markets,’ Econometrica: Journal of the Econometric Society, pp. 781–831, 1997.
  55. 55.Z. Eckstein and K. I. Wolpin, ‘Why youths drop out of high school: The impact of preferences, opportunities, and abilities,’ Econometrica, vol. 67, no. 6, pp. 1295–1339, 1999.
  56. 56.E. Duflo, R. Hanna, and S. P. Ryan, ‘Incentives work: Getting teachers to come to school,’ American Economic Review, vol. 102, no. 4, pp. 1241–78, 2012.
  57. 57.H. Fang and Y. Wang, ‘Estimating dynamic discrete choice models with hyperbolic discounting, with an application to mammography decisions,’ International Economic Review, vol. 56, no. 2, pp. 565–596, 2015.
  58. 58.L. Caliendo, M. Dvorkin, and F. Parro, ‘Trade and labor market dynamics: General equilibrium analysis of the china trade shock,’ Econometrica, vol. 87, no. 3, pp. 741–835, 2019.
  59. 59.C. Cirillo, R. Xu, and F. Bastin, ‘A dynamic formulation for car ownership modeling,’ Transportation Science, vol. 50, no. 1, pp. 322–335, 2016.
  60. 60.T. Xu, Z. Wang, and Y. Liang, ‘Improving sample complexity bounds for (natural) actor-critic algorithms,’ Advances in Neural Information Processing Systems, vol. 33, pp. 4358–4369, 2020.

Citation

MLA
Zeng, S., et al. “Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time Guarantees”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 10122–35, https://proceedings.neurips.cc/paper_files/paper/2022/file/41bd71e7bf7f9fe68f1c936940fd06bd-Paper-Conference.pdf.
APA
Zeng, S., Li, C., Garcia, A., & Hong, M. (2022). Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time Guarantees. Advances in Neural Information Processing Systems, 35, 10122–10135. https://proceedings.neurips.cc/paper_files/paper/2022/file/41bd71e7bf7f9fe68f1c936940fd06bd-Paper-Conference.pdf
Chicago
Zeng, S., C. Li, A. Garcia, and M. Hong. 2022. “Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time Guarantees”. Advances in Neural Information Processing Systems 35: 10122–35. https://proceedings.neurips.cc/paper_files/paper/2022/file/41bd71e7bf7f9fe68f1c936940fd06bd-Paper-Conference.pdf.
Harvard
Zeng, S. et al. (2022) “Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time Guarantees”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 10122–10135. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/41bd71e7bf7f9fe68f1c936940fd06bd-Paper-Conference.pdf.
Vancouver
1. Zeng S, Li C, Garcia A, Hong M (2022) Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time Guarantees. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 10122–10135

BibTeX

@inproceedings{zeng2022maximum,
  title = {Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time Guarantees},
  author = {Zeng, Siliang and Li, Chenliang and Garcia, Alfredo and Hong, Mingyi},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {10122-10135},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/41bd71e7bf7f9fe68f1c936940fd06bd-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors