Non-asymptotic and Accurate Learning of Nonlinear Dynamical Systems

Yahya SattarSamet Oymak

article2022JMLR63 citations

Establishes non-asymptotic sample complexity and convergence guarantees for learning nonlinear dynamical systems from a single finite trajectory using gradient descent by connecting temporally dependent data to independent samples through mixing-time arguments.

Listen

Modern sequential modeling and automated control systems—such as those used in speech processing, robotics, and cyber-physical infrastructure—increasingly rely on nonlinear dynamical systems. A persistent challenge in deploying these systems is system identification: accurately estimating unknown system parameters from real-world, operational time-series data. In real-world environments, practitioners typically have access to only a single finite trajectory of measurements where consecutive data points are statistically dependent, making it difficult to guarantee high accuracy, fast computational convergence, and minimal sample requirements.

The article establishes non-asymptotic statistical and computational guarantees for learning the unknown parameters of nonlinear dynamical systems from a single finite trajectory using standard first-order optimization. Specifically, it demonstrates that gradient descent can accurately and efficiently learn these dynamics in the presence of additive process noise.

The authors develop a theoretical framework that bridges time-dependent sequential trajectories and independent statistical learning theory. By assuming the closed-loop system is stable, the article utilizes a mixing-time argument showing that the system forgets its past states exponentially fast, allowing temporally dependent trajectory samples to be treated as independent approximations. To handle nonconvex optimization landscapes, the framework introduces a local one-point convexity and smoothness condition. The authors validate this mathematical foundation across standard linear setups and nonlinear activation functions, supported by numerical simulations evaluating the effects of noise levels, trajectory lengths, and degree of nonlinearity.

The investigation yields several key findings. First, gradient descent achieves linear computational convergence to the true parameters up to a residual statistical error that scales proportionally with the noise level and inversely with the square root of the trajectory length. Second, the required sample complexity scales linearly with the system dimension, establishing optimal estimation rates. Third, when systems exhibit separable structures across state updates, the sample complexity reduces further, scaling with the component dimension rather than the full parameter count. Finally, the analysis proves that the uniform convergence bounds explicitly capture the optimization error and background noise, removing common limitations from earlier literature that required strictly bounded nonlinear functions.

These results provide mathematical justification and operational confidence for using gradient descent in complex sequential modeling and control. For engineering workflows, these findings demonstrate that standard first-order training algorithms are robust against temporal correlations and noise, minimizing data collection costs and eliminating the need for multi-trajectory reset experiments. Additionally, the analysis demonstrates that stabilizing policies and certain nonlinearities can prevent divergence and ensure bounded states even when underlying linear dynamics are unstable.

Practitioners should prioritize verifying or enforcing stability via stabilizing control policies, as the required trajectory length and convergence rate degrade significantly as the system approaches marginal stability. When designing architectures, engineering teams should exploit separable state updates to minimize sample requirements. Furthermore, setting the gradient descent step size in proportion to the system's geometric conditioning is recommended to ensure linear convergence. Looking ahead, future research should explore learning mechanisms that do not rely on mixing-time assumptions and extend guarantees to active, data-driven control policy optimization in fully nonlinear regimes.

The primary limitation of this work lies in its dependence on the system stability decay parameter; as stability degrades toward the boundary, sample requirements grow substantially. The theoretical guarantees also assume additive random noise, independent random exploration inputs, and a known stabilizing control policy. Nevertheless, within these clearly specified operational boundaries, the article provides high theoretical and empirical confidence for the reliable estimation of nonlinear dynamical systems.

arXiv: 2002.08538

No sufficiently relevant recommendations were found.

Cover for Non-asymptotic and Accurate Learning of Nonlinear Dynamical Systems

Abstract

We consider the problem of learning a nonlinear dynamical system governed by a nonlinear state equation ht+1 = ϕ(ht, ut; θ) + wt. Here θ is the unknown system dynamics, ht is the state, ut is the input and wt is the additive noise vector. We study gradient based algorithms to learn the system dynamics θ from samples obtained from a single finite trajectory. If the system is run by a stabilizing input policy, then using a mixing-time argument we show that temporally-dependent samples can be approximated by i.i.d. samples. We then develop new guarantees for the uniform convergence of the gradient of the empirical loss induced by these i.i.d. samples. Unlike existing works, our bounds are noise sensitive which allows for learning the ground-truth dynamics with high accuracy and small sample complexity. When combined, our results facilitate efficient learning of a broader class of nonlinear dynamical systems as compared to the prior works. We specialize our guarantees to entrywise nonlinear activations and verify our theory in various numerical experiments.

Table of Contents

  • 1. Introduction
  • 2. Problem Setup
  • 2.1 Assumptions on the System and the Inputs
  • 2.2 Optimization Machinery
  • 3. Accurate Statistical Learning with Gradient Descent
  • 4. Learning from a Single Trajectory
  • 5. Main Results
  • 5.1 Non-asymptotic Identification of Nonlinear Systems
  • 5.2 Separable Dynamical Systems
  • 6. Applications
  • 6.1 Linear Dynamical Systems
  • 6.2 Nonlinear State Equations
  • 7. Numerical Experiments
  • 8. Related Work
  • 9. Conclusions
  • Acknowledgments
  • 10. Proofs of the Main Results
  • 10.1 Proof of Theorem 4
  • 10.2 Proof of Lemma 6
  • 10.3 Proof of Lemma 8
  • 10.4 Proof of Theorem 11
  • 10.5 Proof of Theorem 12
  • 10.6 Proof of Theorem 13
  • 10.7 Proof of Theorem 14
  • References
  • Appendix A. Proof of Corollaries 15 and 16
  • A.1 Application to Linear Dynamical Systems
  • A.1.1 Verification of Assumption 1
  • A.1.2 Verification of Assumption 2
  • A.1.3 Verification of Assumption 3
  • A.1.4 Verification of Assumption 4
  • A.1.5 Verification of Assumption 5
  • A.1.6 Proof of Corollary 15
  • A.2 Application to Nonlinear State Equations
  • A.2.1 Verification of Assumption 2
  • A.2.2 Verification of Assumption 3
  • A.2.3 Verification of Assumption 4
  • A.2.4 Verification of Assumption 5
  • A.2.5 Proof of Corollary 16

Knowls

  1. Knowl 1 — Single-trajectory system identification objective

    model/method

    The paper studies identification of an unknown nonlinear dynamical system

    ht+1=ϕ(ht,ut;θ⋆)+wt,h_{t+1}=\phi(h_t,u_t;\theta^\star)+w_t,

    where the unknown parameter is θ⋆∈Rd\theta^\star\in\mathbb{R}^d, the state is ht∈Rnh_t\in\mathbb{R}^n, the input is ut∈Rpu_t\in\mathbb{R}^p, and the process noise is wt∈Rnw_t\in\mathbb{R}^n. Inputs are generated by a fixed feedback policy with independent exploration,

    ut=π(ht)+zt,zt∼i.i.d.Dz,u_t=\pi(h_t)+z_t,\qquad z_t\stackrel{\mathrm{i.i.d.}}{\sim}D_z,

    so the closed-loop dynamics are

    ht+1=ϕ~(ht,zt;θ⋆)+wt,ϕ~(h,z;θ):=ϕ(h,π(h)+z;θ).h_{t+1}=\widetilde\phi(h_t,z_t;\theta^\star)+w_t, \qquad \widetilde\phi(h,z;\theta):=\phi(h,\pi(h)+z;\theta).

    Given one trajectory (ht,zt)t=0T−1(h_t,z_t)_{t=0}^{T-1}, the estimator minimizes the squared one-step prediction loss after discarding an integer churn period L≥1L\ge 1:

    L^(θ)=12(T−L)∑t=LT−1∥ht+1−ϕ~(ht,zt;θ)∥22,θ^∈arg⁡min⁡θ∈RdL^(θ).\widehat L(\theta)=\frac{1}{2(T-L)}\sum_{t=L}^{T-1}\left\|h_{t+1}-\widetilde\phi(h_t,z_t;\theta)\right\|_2^2, \qquad \widehat\theta\in\arg\min_{\theta\in\mathbb{R}^d}\widehat L(\theta).

    The optimization method is fixed-step gradient descent,

    θτ+1=θτ−η∇L^(θτ),\theta_{\tau+1}=\theta_\tau-\eta\nabla\widehat L(\theta_\tau),

    where η>0\eta>0 is a constant learning rate. The central difficulty is that successive terms in L^\widehat L are temporally dependent, despite the loss having the form of a standard empirical-risk objective.

  2. Knowl 2 — Stability and bounded-data assumptions

    assumption

    The closed-loop system is assumed to forget its initial state exponentially. For any initial state α∈Rn\alpha\in\mathbb{R}^n, let ht(α)h_t(\alpha) denote the state at time tt under the same excitation and noise sequences, initialized at h0=αh_0=\alpha. The system is (Cρ,ρ)(C_\rho,\rho)-stable if there are Cρ≥1C_\rho\ge 1 and ρ∈(0,1)\rho\in(0,1) such that

    ∥ht(α)−ht(0)∥2≤Cρρt∥α∥2\left\|h_t(\alpha)-h_t(0)\right\|_2\le C_\rho\rho^t\|\alpha\|_2

    for every admissible sequence of excitations and noises. This is a contraction of trajectories with different initial states, not necessarily convergence of either trajectory to a fixed point. For a stable linear system, the condition follows when the spectral radius of its state matrix is below one.

    The data-generation assumptions are h0=0h_0=0, independent identically distributed excitations ztz_t, and independent identically distributed process noises wtw_t. There are constants B,cw,σ>0B,c_w,\sigma>0 and p0∈[0,1)p_0\in[0,1) such that, with probability at least 1−p01-p_0 over the entire trajectory,

    ∥ϕ~(0,zt;θ⋆)∥2≤Bn,∥wt∥∞≤cwσ,0≤t≤T−1.\left\|\widetilde\phi(0,z_t;\theta^\star)\right\|_2\le B\sqrt n, \qquad \|w_t\|_\infty\le c_w\sigma, \qquad 0\le t\le T-1.

    Under these conditions, all states obey the high-probability bound

    ∥ht∥2≤β+n,β+:=Cρ(cwσ+B)1−ρ,0≤t≤T.\|h_t\|_2\le \beta_+\sqrt n, \qquad \beta_+:=\frac{C_\rho(c_w\sigma+B)}{1-\rho}, \qquad 0\le t\le T.

    Thus, stability controls both the dependence between observations and the magnitude of the states used in the loss.

  3. Knowl 3 — Auxiliary population loss and one-point optimization geometry

    assumption

    To analyze the dependent trajectory loss, the paper introduces the population loss obtained from a fresh length-LL trajectory. For h0=0h_0=0, define

    LD(θ):=E[ℓ(θ;hL,hL−1,zL−1)],L_D(\theta):=\mathbb{E}\left[\ell(\theta;h_L,h_{L-1},z_{L-1})\right],

    where

    ℓ(θ;hL,hL−1,zL−1):=12∥hL−ϕ~(hL−1,zL−1;θ)∥22.\ell(\theta;h_L,h_{L-1},z_{L-1}) :=\frac12\left\|h_L-\widetilde\phi(h_{L-1},z_{L-1};\theta)\right\|_2^2.

    The unknown parameter θ⋆\theta^\star is treated as the population minimizer. In a radius-rr ball around it, the population loss need not be convex globally; instead, it is assumed to satisfy one-point convexity and one-point smoothness. Specifically, there are constants β≥α>0\beta\ge\alpha>0 such that for every θ∈Bd(θ⋆,r)\theta\in B^d(\theta^\star,r),

    ⟨θ−θ⋆,∇LD(θ)⟩≥α∥θ−θ⋆∥22,\left\langle \theta-\theta^\star,\nabla L_D(\theta)\right\rangle \ge \alpha\|\theta-\theta^\star\|_2^2,

    and

    ∥∇LD(θ)∥2≤β∥θ−θ⋆∥2.\|\nabla L_D(\theta)\|_2 \le \beta\|\theta-\theta^\star\|_2.

    This one-point convexity and smoothness condition is weaker than global strong convexity and global smoothness. It is the optimization condition that converts a uniform empirical-gradient approximation into linear convergence of gradient descent up to a statistical error.

  4. Knowl 4 — Noise-sensitive uniform convergence for i.i.d. empirical gradients

    theoretical result

    For NN independent samples x1,…,xN∼Dx_1,\ldots,x_N\sim D, let

    L^S(θ):=1N∑i=1NL(θ,xi),LD(θ):=Ex∼D[L(θ,x)],\widehat L_S(\theta):=\frac1N\sum_{i=1}^N L(\theta,x_i), \qquad L_D(\theta):=\mathbb{E}_{x\sim D}[L(\theta,x)],

    and let θ⋆\theta^\star be the population minimizer. On the local ball Bd(θ⋆,r)B^d(\theta^\star,r), suppose both ∇L^S\nabla\widehat L_S and ∇LD\nabla L_D are LDL_D-Lipschitz with probability at least 1−p01-p_0. Suppose also that the centered single-sample gradient has subexponential norm

    ∥∇L(θ,x)−E[∇L(θ,x)]∥ψ1≤σ0+K∥θ−θ⋆∥2,\left\|\nabla L(\theta,x)-\mathbb{E}[\nabla L(\theta,x)]\right\|_{\psi_1} \le \sigma_0+K\|\theta-\theta^\star\|_2,

    where K,σ0>0K,\sigma_0>0 and, for a scalar random variable XX, ∥X∥ψ1:=sup⁡q≥1(E∣X∣q)1/q/q\|X\|_{\psi_1}:=\sup_{q\ge1}(\mathbb{E}|X|^q)^{1/q}/q; for a vector, the norm is the supremum over unit directional projections.

    Then there is an absolute constant c0>0c_0>0 such that, simultaneously for all θ∈Bd(θ⋆,r)\theta\in B^d(\theta^\star,r),

    ∥∇L^S(θ)−∇LD(θ)∥2≤c0(σ0+K∥θ−θ⋆∥2)log⁡ ⁣(3(LDNK+1))dN,\left\|\nabla\widehat L_S(\theta)-\nabla L_D(\theta)\right\|_2 \le c_0\left(\sigma_0+K\|\theta-\theta^\star\|_2\right) \log\!\left(3\left(\frac{L_DN}{K}+1\right)\right)\sqrt{\frac dN},

    with probability at least

    1−p0−log⁡ ⁣(Krσ0)e−100d.1-p_0-\log\!\left(\frac{Kr}{\sigma_0}\right)e^{-100d}.

    The bound separates irreducible gradient noise, represented by σ0\sigma_0, from optimization-dependent gradient variation, represented by K∥θ−θ⋆∥2K\|\theta-\theta^\star\|_2. Consequently, at the population minimizer the statistical term scales as σ0d/N\sigma_0\sqrt{d/N} rather than (σ0+rK)d/N(\sigma_0+rK)\sqrt{d/N}. The result only requires subexponential, rather than subgaussian, gradient fluctuations.

  5. Knowl 5 — Truncation converts a stable trajectory into independent sub-trajectories

    model/method

    For a state at time t≥Lt\ge L, define the LL-truncated state ht,Lh_{t,L} by resetting the state at time t−Lt-L to zero while preserving the excitations and noises from times t−Lt-L through t−1t-1. Equivalently, inputs and noises before t−Lt-L are replaced by zero and those afterward are kept unchanged. Stability gives the exponential approximation

    ∥ht−ht,L∥2≤CρρL∥ht−L∥2.\|h_t-h_{t,L}\|_2\le C_\rho\rho^L\|h_{t-L}\|_2.

    Set N=⌊(T−L)/L⌋N=\lfloor (T-L)/L\rfloor. For each offset τ∈{0,…,L−1}\tau\in\{0,\ldots,L-1\}, subsample times τ+L,τ+2L,…,τ+NL\tau+L,\tau+2L,\ldots,\tau+NL and form

    hˉ(i):=hτ+iL,L−1,z(i):=zτ+iL,i=1,…,N.\bar h^{(i)}:=h_{\tau+iL,L-1}, \qquad z^{(i)}:=z_{\tau+iL}, \qquad i=1,\ldots,N.

    Each truncated state hˉ(i)\bar h^{(i)} depends only on the noise and excitation block between consecutive sampling points. These blocks are disjoint, so the collections {hˉ(i)}i=1N\{\bar h^{(i)}\}_{i=1}^N, {z(i)}i=1N\{z^{(i)}\}_{i=1}^N, and the corresponding noises are mutually independent; the truncated states are identically distributed as hL−1h_{L-1}. The target states

    yˉ(i):=hτ+iL+1,L\bar y^{(i)}:=h_{\tau+iL+1,L}

    then produce independent and identically distributed triples (yˉ(i),hˉ(i),z(i))(\bar y^{(i)},\bar h^{(i)},z^{(i)}) with the same distribution as (hL,hL−1,zL−1)(h_L,h_{L-1},z_{L-1}).

    The resulting truncated empirical loss is

    L^tr(θ)=12(T−L)∑t=LT−1∥ht+1,L−ϕ~(ht,L−1,zt;θ)∥22=1L∑τ=0L−1ℓ^τtr(θ),\widehat L^{\mathrm{tr}}(\theta)=\frac{1}{2(T-L)}\sum_{t=L}^{T-1} \left\|h_{t+1,L}-\widetilde\phi(h_{t,L-1},z_t;\theta)\right\|_2^2 =\frac1L\sum_{\tau=0}^{L-1}\widehat\ell^{\mathrm{tr}}_\tau(\theta),

    where each ℓ^τtr\widehat\ell^{\mathrm{tr}}_\tau is an i.i.d.-sample loss with sample size NN. Thus, stability lets the dependent single trajectory be compared with LL independent-looking subproblems, at an exponentially small truncation cost.

  6. Knowl 6 — Uniform gradient approximation from one dependent trajectory

    theoretical result

    Consider the nonlinear system and trajectory loss

    L^(θ)=12(T−L)∑t=LT−1∥ht+1−ϕ~(ht,zt;θ)∥22,\widehat L(\theta)=\frac{1}{2(T-L)}\sum_{t=L}^{T-1} \|h_{t+1}-\widetilde\phi(h_t,z_t;\theta)\|_2^2,

    and the auxiliary population loss LDL_D defined from (hL,hL−1,zL−1)(h_L,h_{L-1},z_{L-1}). Assume stability and bounded data as above, the auxiliary-loss gradient conditions

    ∥∇L(θ,x)−E∇L(θ,x)∥ψ1≤σ0+K∥θ−θ⋆∥2,\|\nabla L(\theta,x)-\mathbb{E}\nabla L(\theta,x)\|_{\psi_1} \le \sigma_0+K\|\theta-\theta^\star\|_2,

    and local LDL_D-Lipschitzness of empirical and population gradients. Also assume that, uniformly for θ∈Bd(θ⋆,r)\theta\in B^d(\theta^\star,r) and admissible (h,z)(h,z),

    ∥∇θϕ~k(h,z;θ)∥2≤Cϕ~,∥∇h∇θϕ~k(h,z;θ)∥≤Dϕ~\|\nabla_\theta\widetilde\phi_k(h,z;\theta)\|_2\le C_{\widetilde\phi}, \qquad \|\nabla_h\nabla_\theta\widetilde\phi_k(h,z;\theta)\|\le D_{\widetilde\phi}

    for every output coordinate kk. Let β+=Cρ(cwσ+B)/(1−ρ)\beta_+=C_\rho(c_w\sigma+B)/(1-\rho) and

    Kϕ~:=2c0β+Dϕ~(cwσσ0∨Cϕ~K).K_{\widetilde\phi}:=\frac{2}{c_0}\beta_+D_{\widetilde\phi} \left(\frac{c_w\sigma}{\sigma_0}\vee\frac{C_{\widetilde\phi}}{K}\right).

    Choose

    N=⌊T−LL⌋,L≥⌈1+log⁡ ⁣(CρKϕ~nN/d)1−ρ⌉,N=\left\lfloor\frac{T-L}{L}\right\rfloor, \qquad L\ge \left\lceil 1+\frac{\log\!\left(C_\rho K_{\widetilde\phi}n\sqrt{N/d}\right)}{1-\rho}\right\rceil,

    where c0c_0 is the absolute constant in the i.i.d. gradient bound. Then, with probability at least

    1−2Lp0−Llog⁡ ⁣(Krσ0)e−100d,1-2Lp_0-L\log\!\left(\frac{Kr}{\sigma_0}\right)e^{-100d},

    all θ∈Bd(θ⋆,r)\theta\in B^d(\theta^\star,r) satisfy

    ∥∇L^(θ)−∇LD(θ)∥2≤2c0(σ0+K∥θ−θ⋆∥2)log⁡ ⁣(3(LDNK+1))dN.\left\|\nabla\widehat L(\theta)-\nabla L_D(\theta)\right\|_2 \le 2c_0\left(\sigma_0+K\|\theta-\theta^\star\|_2\right) \log\!\left(3\left(\frac{L_DN}{K}+1\right)\right)\sqrt{\frac dN}.

    The result combines i.i.d. concentration on the truncated sub-trajectories with the stability error between truncated and actual states. The mixing length is proportional to (1−ρ)−1(1-\rho)^{-1}, so the effective sample size is N≃T/LN\simeq T/L and the guarantee deteriorates as ρ\rho approaches one.

  7. Knowl 7 — Gradient descent identifies the dynamics up to the noise-limited statistical radius

    theoretical result

    Let L^\widehat L be the single-trajectory loss for the stable nonlinear system, and let LDL_D be its auxiliary population loss. Suppose LDL_D satisfies one-point convexity and one-point smoothness on Bd(θ⋆,r)B^d(\theta^\star,r) with parameters α\alpha and β\beta, and suppose the single-trajectory gradient obeys, uniformly on that ball,

    ∥∇L^(θ)−∇LD(θ)∥2≤ν+α2∥θ−θ⋆∥2.\|\nabla\widehat L(\theta)-\nabla L_D(\theta)\|_2 \le \nu+\frac{\alpha}{2}\|\theta-\theta^\star\|_2.

    If r≥5ν/αr\ge 5\nu/\alpha, the initialization satisfies θ0∈Bd(θ⋆,r)\theta_0\in B^d(\theta^\star,r), and gradient descent uses

    η=α16β2,\eta=\frac{\alpha}{16\beta^2},

    then every iterate remains controlled and satisfies

    ∥θτ−θ⋆∥2≤(1−α2128β2)τ∥θ0−θ⋆∥2+5να.\|\theta_\tau-\theta^\star\|_2 \le \left(1-\frac{\alpha^2}{128\beta^2}\right)^\tau \|\theta_0-\theta^\star\|_2+\frac{5\nu}{\alpha}.

    For the trajectory problem, choose N≳K2Clog⁡2d/α2N\gtrsim K^2C_{\log}^2d/\alpha^2, where

    Clog⁡:=log⁡ ⁣(3(LDNK+1)),C_{\log}:=\log\!\left(3\left(\frac{L_DN}{K}+1\right)\right),

    and choose the mixing length so that the one-trajectory gradient bound holds. If σ0≲rK\sigma_0\lesssim rK, then with probability at least

    1−2Lp0−Llog⁡ ⁣(Krσ0)e−100d,1-2Lp_0-L\log\!\left(\frac{Kr}{\sigma_0}\right)e^{-100d},

    all iterates obey

    ∥θτ−θ⋆∥2≤(1−α2128β2)τ∥θ0−θ⋆∥2+cσ0αClog⁡dN,\|\theta_\tau-\theta^\star\|_2 \le \left(1-\frac{\alpha^2}{128\beta^2}\right)^\tau \|\theta_0-\theta^\star\|_2 +\frac{c\sigma_0}{\alpha}C_{\log}\sqrt{\frac dN},

    for an absolute constant cc. Thus the method has linear computational convergence and reaches a noise-sensitive statistical radius of order σ0d/N/α\sigma_0\sqrt{d/N}/\alpha. The required number of independent-equivalent samples is linear in dd up to logarithmic factors, while the original trajectory length is approximately T≃L(N+1)T\simeq L(N+1).

  8. Knowl 8 — Separable systems reduce identification to coordinatewise problems

    model/method

    A separable nonlinear system has coordinatewise parameter blocks θk⋆∈Rdˉ\theta_k^\star\in\mathbb{R}^{\bar d}, with dˉ=d/n\bar d=d/n, and dynamics

    ht+1[k]=ϕ~k(ht,zt;θk⋆)+wt[k],1≤k≤n.h_{t+1}[k]=\widetilde\phi_k(h_t,z_t;\theta_k^\star)+w_t[k], \qquad 1\le k\le n.

    Its trajectory loss decomposes as

    L^(θ)=∑k=1nL^k(θk),\widehat L(\theta)=\sum_{k=1}^n\widehat L_k(\theta_k),

    where

    L^k(θk)=12(T−L)∑t=LT−1(ht+1[k]−ϕ~k(ht,zt;θk))2.\widehat L_k(\theta_k)=\frac{1}{2(T-L)}\sum_{t=L}^{T-1} \left(h_{t+1}[k]-\widetilde\phi_k(h_t,z_t;\theta_k)\right)^2.

    Therefore, both empirical-risk minimization and gradient descent split into nn independent parameter problems:

    θk(τ+1)=θk(τ)−η∇L^k(θk(τ)).\theta_k^{(\tau+1)}=\theta_k^{(\tau)}-\eta\nabla\widehat L_k(\theta_k^{(\tau)}).

    Assume that each coordinatewise auxiliary loss satisfies one-point convexity and smoothness with the same α,β\alpha,\beta, and that the coordinatewise gradient noise and Lipschitz conditions hold with parameters K,σ0,LDK,\sigma_0,L_D. With N=⌊(T−L)/L⌋N=\lfloor(T-L)/L\rfloor, choose

    L≥⌈1+log⁡ ⁣(CρKϕ~nN/dˉ)1−ρ⌉,N≳K2Clog⁡2dˉ/α2,L\ge \left\lceil 1+\frac{\log\!\left(C_\rho K_{\widetilde\phi}n\sqrt{N/\bar d}\right)}{1-\rho}\right\rceil, \qquad N\gtrsim K^2C_{\log}^2\bar d/\alpha^2,

    where Clog⁡=log⁡(3(LDN/K+1))C_{\log}=\log(3(L_DN/K+1)) and Kϕ~K_{\widetilde\phi} is the system-dependent truncation constant. Starting from θk(0)\theta_k^{(0)} in the local radius and using η=α/(16β2)\eta=\alpha/(16\beta^2), if σ0≲rK\sigma_0\lesssim rK, then, simultaneously for all coordinates,

    ∥θk(τ)−θk⋆∥2≤(1−α2128β2)τ∥θk(0)−θk⋆∥2+cσ0αClog⁡dˉN.\|\theta_k^{(\tau)}-\theta_k^\star\|_2 \le \left(1-\frac{\alpha^2}{128\beta^2}\right)^\tau \|\theta_k^{(0)}-\theta_k^\star\|_2 +\frac{c\sigma_0}{\alpha}C_{\log}\sqrt{\frac{\bar d}{N}}.

    The success probability is at least

    1−2Lnp0−Lnlog⁡ ⁣(Krσ0)e−100dˉ.1-2Lnp_0-Ln\log\!\left(\frac{Kr}{\sigma_0}\right)e^{-100\bar d}.

    Thus separability replaces the ambient parameter dimension dd by the per-coordinate dimension dˉ=d/n\bar d=d/n: each transition supplies nn scalar equations, yielding an O(dˉ)O(\bar d) effective sample requirement.

  9. Knowl 9 — Linear dynamical systems achieve dimension-optimal identification rates

    theoretical result

    For the stable linear system

    ht+1=A⋆ht+B⋆zt+wt,h_{t+1}=A^\star h_t+B^\star z_t+w_t,

    assume A⋆∈Rn×nA^\star\in\mathbb{R}^{n\times n} has ρ(A⋆)<1\rho(A^\star)<1, zt∼N(0,Ip)z_t\sim\mathcal N(0,I_p), and wt∼N(0,σ2In)w_t\sim\mathcal N(0,\sigma^2I_n) independently. Define xt=[ht⊤ zt⊤]⊤x_t=[h_t^\top\ z_t^\top]^\top, Θ⋆=[A⋆ B⋆]\Theta^\star=[A^\star\ B^\star], and let θk⋆\theta_k^\star be row kk of Θ⋆\Theta^\star. For

    Gt=[(A⋆)t−1B⋆ (A⋆)t−2B⋆ ⋯ B⋆],G_t=[(A^\star)^{t-1}B^\star\ (A^\star)^{t-2}B^\star\ \cdots\ B^\star], Ft=[(A⋆)t−1 (A⋆)t−2 ⋯ In],F_t=[(A^\star)^{t-1}\ (A^\star)^{t-2}\ \cdots\ I_n],

    set

    γ−:=1∧λmin⁡(GL−1GL−1⊤+σ2FL−1FL−1⊤),\gamma_-:=1\wedge\lambda_{\min}(G_{L-1}G_{L-1}^\top+\sigma^2F_{L-1}F_{L-1}^\top), γ+:=1∨λmax⁡(GL−1GL−1⊤+σ2FL−1FL−1⊤),κ:=γ+/γ−.\gamma_+:=1\vee\lambda_{\max}(G_{L-1}G_{L-1}^\top+\sigma^2F_{L-1}F_{L-1}^\top), \qquad \kappa:=\gamma_+/\gamma_-.

    With N=⌊(T−L)/L⌋N=\lfloor(T-L)/L\rfloor, choose

    L≥⌈1+log⁡ ⁣(CCρβ+N(n+p)/γ+)1−ρ⌉,L\ge \left\lceil 1+\frac{\log\!\left(CC_\rho\beta_+N(n+p)/\gamma_+\right)}{1-\rho}\right\rceil,

    where β+=1∨max⁡1≤t≤Tλmax⁡(GtGt⊤+σ2FtFt⊤)\beta_+=1\vee\max_{1\le t\le T}\lambda_{\max}(G_tG_t^\top+\sigma^2F_tF_t^\top) and CC is absolute. If

    N≳κ2log⁡2(6N+3)(n+p),N\gtrsim \kappa^2\log^2(6N+3)(n+p),

    use η=γ−/(16γ+2)\eta=\gamma_-/(16\gamma_+^2) and initialize Θ(0)=0\Theta^{(0)}=0. Provided σ≲∥Θ⋆∥Fγ+\sigma\lesssim\|\Theta^\star\|_F\sqrt{\gamma_+}, gradient descent satisfies, for every row kk,

    ∥θk(τ)−θk⋆∥2≤(1−γ−2128γ+2)τ∥θk(0)−θk⋆∥2+cσκγ−log⁡(6N+3)n+pN.\|\theta_k^{(\tau)}-\theta_k^\star\|_2 \le \left(1-\frac{\gamma_-^2}{128\gamma_+^2}\right)^\tau \|\theta_k^{(0)}-\theta_k^\star\|_2 +c\sigma\frac{\sqrt\kappa}{\sqrt{\gamma_-}}\log(6N+3)\sqrt{\frac{n+p}{N}}.

    The stated success probability is at least

    1−4Te−100n−Ln(4+log⁡ ⁣(∥Θ⋆∥Fγ+σ))e−100(n+p).1-4T e^{-100n}-Ln\left(4+\log\!\left(\frac{\|\Theta^\star\|_F\sqrt{\gamma_+}}{\sigma}\right)\right)e^{-100(n+p)}.

    The effective sample requirement is O(n+p)O(n+p) up to conditioning and logarithmic factors, matching the fact that n(n+p)n(n+p) parameters are learned from nn scalar equations per transition. The statistical error has the optimal dependence σ(n+p)/N\sigma\sqrt{(n+p)/N}, modulated by the covariance condition number.

  10. Knowl 10 — Entrywise nonlinear state equations admit explicit gradient-descent guarantees

    theoretical result

    Consider the nonlinear state equation

    ht+1=ϕ(Θ⋆ht)+zt+wt,h_{t+1}=\phi(\Theta^\star h_t)+z_t+w_t,

    where Θ⋆∈Rn×n\Theta^\star\in\mathbb{R}^{n\times n}, zt∼N(0,In)z_t\sim\mathcal N(0,I_n), wt∼N(0,σ2In)w_t\sim\mathcal N(0,\sigma^2I_n), and ϕ\phi acts entrywise. Assume the system is (Cρ,ρ)(C_\rho,\rho)-stable, ϕ(0)=0\phi(0)=0, ϕ\phi is γ\gamma-increasing with ϕ′(x)≥γ>0\phi'(x)\ge\gamma>0, and ∣ϕ′(x)∣,∣ϕ′′(x)∣≤1|\phi'(x)|,|\phi''(x)|\le1 for all xx. Gradient descent minimizes

    L^(Θ)=12(T−L)∑t=LT−1∥ht+1−ϕ(Θht)−zt∥22.\widehat L(\Theta)=\frac{1}{2(T-L)}\sum_{t=L}^{T-1} \|h_{t+1}-\phi(\Theta h_t)-z_t\|_2^2.

    Define β+:=Cρ(1+σ)/(1−ρ)\beta_+:=C_\rho(1+\sigma)/(1-\rho) and

    Dlog⁡:=log⁡ ⁣(3(1+σ)n+3Cρ(1+σ)∥Θ⋆∥Fn3/2log⁡3/2(2T)N1−ρ+3).D_{\log}:=\log\!\left(3(1+\sigma)n+\frac{3C_\rho(1+\sigma)\|\Theta^\star\|_F n^{3/2}\log^{3/2}(2T)N}{1-\rho}+3\right).

    For N=⌊(T−L)/L⌋N=\lfloor(T-L)/L\rfloor, choose

    L≥⌈1+log⁡ ⁣(CCρ(1+∥Θ⋆∥Fβ+)Nn)1−ρ⌉,L\ge\left\lceil1+\frac{\log\!\left(CC_\rho(1+\|\Theta^\star\|_F\beta_+)Nn\right)}{1-\rho}\right\rceil,

    and suppose

    N≳Cρ4γ4(1−ρ)4Dlog⁡2n.N\gtrsim \frac{C_\rho^4}{\gamma^4(1-\rho)^4}D_{\log}^2n.

    Starting from Θ(0)=0\Theta^{(0)}=0, use

    η=γ2(1−ρ)432Cρ4(1+σ)2n2.\eta=\frac{\gamma^2(1-\rho)^4}{32C_\rho^4(1+\sigma)^2n^2}.

    If σ≲∥Θ⋆∥F\sigma\lesssim\|\Theta^\star\|_F, then, for each row θk(τ)\theta_k^{(\tau)} of the iterate and corresponding true row θk⋆\theta_k^\star,

    ∥θk(τ)−θk⋆∥2≤(1−γ4(1−ρ)4512Cρ4n2)τ∥θk(0)−θk⋆∥2+cσCργ2(1−ρ)Dlog⁡nN.\|\theta_k^{(\tau)}-\theta_k^\star\|_2 \le \left(1-\frac{\gamma^4(1-\rho)^4}{512C_\rho^4n^2}\right)^\tau \|\theta_k^{(0)}-\theta_k^\star\|_2 +\frac{c\sigma C_\rho}{\gamma^2(1-\rho)}D_{\log}\sqrt{\frac nN}.

    The success probability is at least

    1−Ln(4T+log⁡ ⁣(∥Θ⋆∥FCρ(1+σ)σ(1−ρ)))e−100n.1-Ln\left(4T+\log\!\left(\frac{\|\Theta^\star\|_F C_\rho(1+\sigma)}{\sigma(1-\rho)}\right)\right)e^{-100n}.

    Because every transition supplies nn scalar equations, the effective sample requirement is O(n)O(n) up to system-dependent and logarithmic factors, and the statistical term scales as σn/N\sigma\sqrt{n/N}.

Coverage note — The numerical validation, including the page-17 state-matrix table and the page-18/page-19 experiments showing effects of nonlinearity, noise, trajectory length, and nonlinear stabilization, was omitted to prioritize the paper's ten load-bearing theoretical and methodological contributions.

References

  1. 1.Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song. On the convergence rate of training recurrent neural networks. In Advances in Neural Information Processing Systems, volume 32, pages 6676–6688, 2019.
  2. 2.Karl Johan Åström and Peter Eykhoff. System identification—a survey. Automatica, 7(2):123–162, 1971.
  3. 3.Karl Johan Åström and Tore Hägglund. PID controllers: theory, design, and tuning, volume 2. Instrument Society of America Research Triangle Park, NC, 1995.
  4. 4.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations, 2015.
  5. 5.Sohail Bahmani and Justin Romberg. Convex programming for estimation in nonlinear recurrent models. Journal of Machine Learning Research, 21(235):1–20, 2020.
  6. 6.Nicholas M Boffi, Stephen Tu, and Jean-Jacques E Slotine. Regret bounds for adaptive nonlinear control. In Learning for Dynamics and Control, pages 471–483. PMLR, 2021.
  7. 7.Sheng Chen, SA Billings, and PM Grant. Non-linear system identification using neural networks. International Journal of Control, 51(6):1191–1214, 1990.
  8. 8.Alon Cohen, Tomer Koren, and Yishay Mansour. Learning linear-quadratic regulators efficiently with only √T regret. In International Conference on Machine Learning, pages 1300–1309. PMLR, 2019.
  9. 9.Christoph Dann and Emma Brunskill. Sample complexity of episodic fixed-horizon reinforcement learning. In Advances in Neural Information Processing Systems, pages 2818–2826, 2015.
  10. 10.Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu. Regret bounds for robust adaptive control of the linear quadratic regulator. In Advances in Neural Information Processing Systems, pages 4188–4197, 2018.
  11. 11.Mohamad Kazem Shirani Faradonbeh, Ambuj Tewari, and George Michailidis. Finite time identification in unstable linear systems. Automatica, 96:342–353, 2018.
  12. 12.Mohamad Kazem Shirani Faradonbeh, Ambuj Tewari, and George Michailidis. Optimism-based adaptive regulation of linear-quadratic systems. IEEE Transactions on Automatic Control, 66(4):1802–1808, 2020.
  13. 13.Salar Fattahi, Nikolai Matni, and Somayeh Sojoudi. Learning sparse dynamical systems from a single sample trajectory. In 2019 IEEE 58th Conference on Decision and Control, pages 2682–2689. IEEE, 2019.
  14. 14.Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi. Global convergence of policy gradient methods for the linear quadratic regulator. In International Conference on Machine Learning, volume 80, pages 1467–1476. PMLR, 2018.
  15. 15.Dylan Foster, Ayush Sekhari, and Karthik Sridharan. Uniform convergence of gradients for non-convex learning and optimization. In Advances in Neural Information Processing Systems, pages 8745–8756, 2018.
  16. 16.Dylan Foster, Tuhin Sarkar, and Alexander Rakhlin. Learning nonlinear dynamical systems from a single trajectory. In Learning for Dynamics and Control, pages 851–861. PMLR, 2020.
  17. 17.Sara A Geer, Sara van de Geer, and D Williams. Empirical processes in M-estimation, volume 6. Cambridge university press, 2000.
  18. 18.Alex Graves, Abdel-Rahman Mohamed, and Geoffrey Hinton. Speech recognition with deep recurrent neural networks. In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 6645–6649. IEEE, 2013.
  19. 19.Moritz Hardt, Tengyu Ma, and Benjamin Recht. Gradient descent learns linear dynamical systems. The Journal of Machine Learning Research, 19(1):1025–1068, 2018.
  20. 20.Elad Hazan, Karan Singh, and Cyril Zhang. Learning linear dynamical systems via spectral filtering. In Advances in Neural Information Processing Systems, volume 30, pages 6702–6712, 2017.
  21. 21.Elad Hazan, Holden Lee, Karan Singh, Cyril Zhang, and Yi Zhang. Spectral filtering for general linear dynamical systems. In Advances in Neural Information Processing Systems, volume 31, pages 4639–4648, 2018.
  22. 22.BL Ho and Rudolf E Kálmán. Effective construction of linear state-variable models from input/output functions. at-Automatisierungstechnik, 14(1-12):545–548, 1966.
  23. 23.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997.
  24. 24.Prateek Jain, Suhas S Kowshik, Dheeraj Nagaraj, and Praneeth Netrapalli. Near-optimal offline and streaming algorithms for learning non-linear dynamical systems. In Advances in Neural Information Processing Systems, volume 34, pages 8518–8531, 2021.
  25. 25.Sham Kakade, Akshay Krishnamurthy, Kendall Lowrey, Motoya Ohnishi, and Wen Sun. Information theoretic regret bounds for online nonlinear control. In Advances in Neural Information Processing Systems, volume 33, pages 15312–15325, 2020.
  26. 26.Sham M Kakade, Varun Kanade, Ohad Shamir, and Adam Kalai. Efficient learning of generalized linear and single index models with isotonic regression. In Advances in Neural Information Processing Systems, volume 24, pages 927–935, 2011.
  27. 27.Seyed Mohammadreza Mousavi Kalan, Mahdi Soltanolkotabi, and A Salman Avestimehr. Fitting relus via sgd and quantized sgd. In 2019 IEEE International Symposium on Information Theory, pages 2469–2473, 2019.
  28. 28.Hamed Karimi, Julie Nutini, and Mark Schmidt. Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pages 795–811. Springer, 2016.
  29. 29.Mohammad Khosravi and Roy S Smith. Convex nonparametric formulation for identification of gradient flows. IEEE Control Systems Letters, 5(3):1097–1102, 2020a.
  30. 30.Mohammad Khosravi and Roy S Smith. Nonlinear system identification with prior knowledge on the region of attraction. IEEE Control Systems Letters, 5(3):1091–1096, 2020b.
  31. 31.Karl Krauth, Stephen Tu, and Benjamin Recht. Finite-time analysis of approximate policy iteration for the linear quadratic regulator. In Advances in Neural Information Processing Systems, volume 32, pages 8514–8524, 2019.
  32. 32.Vitaly Kuznetsov and Mehryar Mohri. Generalization bounds for non-stationary mixing processes. Machine Learning, 106(1):93–117, 2017.
  33. 33.Sahin Lale, Kamyar Azizzadenesheli, Babak Hassibi, and Anima Anandkumar. Model learning predictive control in nonlinear dynamical systems. In 2021 60th IEEE Conference on Decision and Control, pages 757–762. IEEE, 2021.
  34. 34.Michel Ledoux. The concentration of measure phenomenon. Number 89. American Mathematical Soc., 2001.
  35. 35.Shuai Li, Sanfeng Chen, and Bo Liu. Accelerating a recurrent neural network to finite-time convergence for solving time-varying sylvester equation by using a sign-bi-power activation function. Neural processing Letters, 37(2):189–205, 2013.
  36. 36.Lennart Ljung. System identification. In Signal Analysis and Prediction, pages 163–173. Springer, 1998.
  37. 37.Dhruv Malik, Ashwin Pananjady, Kush Bhatia, Koulik Khamaru, Peter Bartlett, and Martin Wainwright. Derivative-free methods for policy optimization: Guarantees for linear quadratic systems. In The 22nd International Conference on Artificial Intelligence and Statistics, pages 2916–2925, 2019.
  38. 38.Horia Mania, Stephen Tu, and Benjamin Recht. Certainty equivalence is efficient for linear quadratic control. In Advances in Neural Information Processing Systems, volume 32, pages 10154–10164, 2019.
  39. 39.Horia Mania, Michael I Jordan, and Benjamin Recht. Active learning for nonlinear system identification with guarantees. Journal of Machine Learning Research, 23(32):1–30, 2022.
  40. 40.Daniel J McDonald, Cosma Rohilla Shalizi, and Mark Schervish. Nonparametric risk bounds for time-series forecasting. The Journal of Machine Learning Research, 18(1):1044–1083, 2017.
  41. 41.Alexandre Megretski and Anders Rantzer. System analysis via integral quadratic constraints. IEEE Transactions on Automatic Control, 42(6):819–830, 1997.
  42. 42.Song Mei, Yu Bai, and Andrea Montanari. The landscape of empirical risk for nonconvex losses. The Annals of Statistics, 46(6A):2747–2774, 2018.
  43. 43.Zakaria Mhammedi, Dylan J Foster, Max Simchowitz, Dipendra Misra, Wen Sun, Akshay Krishnamurthy, Alexander Rakhlin, and John Langford. Learning the linear quadratic regulator from nonlinear observations. In Advances in Neural Information Processing Systems, volume 33, pages 14532–14543, 2020.
  44. 44.Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernock`y, and Sanjeev Khudanpur. Recurrent neural network based language model. In Interspeech, volume 2, pages 1045–1048. Makuhari, 2010.
  45. 45.John Miller and Moritz Hardt. Stable recurrent models. In International Conference on Learning Representations, 2019.
  46. 46.Mehryar Mohri and Afshin Rostamizadeh. Stability bounds for non-iid processes. In Advances in Neural Information Processing Systems, volume 20, pages 1025–1032, 2007.
  47. 47.Mehryar Mohri and Afshin Rostamizadeh. Rademacher complexity bounds for non-iid processes. In Advances in Neural Information Processing Systems, volume 21, pages 1097–1104, 2008.
  48. 48.Yurii Nesterov. Introductory lectures on convex optimization: A basic course, volume 87. Springer Science & Business Media, 2003.
  49. 49.Samet Oymak. Learning compact neural networks with regularization. In International Conference on Machine Learning, pages 3966–3975. PMLR, 2018.
  50. 50.Samet Oymak. Stochastic gradient descent learns state equations with nonlinear activations. In Conference on Learning Theory, pages 2551–2579. PMLR, 2019.
  51. 51.Samet Oymak and Necmiye Ozay. Non-asymptotic identification of lti systems from a single trajectory. In 2019 American control conference, pages 5655–5661. IEEE, 2019.
  52. 52.Rik Pintelon and Johan Schoukens. System identification: a frequency domain approach. John Wiley & Sons, 2012.
  53. 53.Stephen Prajna, Antonis Papachristodoulou, and Pablo A Parrilo. Introducing sostools: A general purpose sum of squares programming solver. In 2002 41st IEEE Conference on Decision and Control, volume 1, pages 741–746. IEEE, 2002.
  54. 54.Benjamin Recht. A tour of reinforcement learning: The view from continuous control. Annual Review of Control, Robotics, and Autonomous Systems, 2:253–279, 2019.
  55. 55.Haşim Sak, Andrew Senior, and Françoise Beaufays. Long short-term memory recurrent neural network architectures for large scale acoustic modeling. In Fifteenth Annual Conference of the International Speech Communication Association, 2014.
  56. 56.Tuhin Sarkar and Alexander Rakhlin. Near optimal finite time identification of arbitrary linear dynamical systems. In International Conference on Machine Learning, pages 5610–5618. PMLR, 2019.
  57. 57.Tuhin Sarkar, Alexander Rakhlin, and Munther Dahleh. Nonparametric system identification of stochastic switched linear systems. In 2019 IEEE 58th Conference on Decision and Control, pages 3623–3628. IEEE, 2019.
  58. 58.Tuhin Sarkar, Alexander Rakhlin, and Munther A Dahleh. Finite time lti system identification. Journal of Machine Learning Research, 22:1–61, 2021.
  59. 59.Yahya Sattar and Samet Oymak. A simple framework for learning stabilizable systems. In 2019 IEEE 8th International Workshop on Computational Advances in Multi-Sensor Adaptive Processing, pages 116–120. IEEE, 2019.
  60. 60.Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht. Learning without mixing: Towards a sharp analysis of linear system identification. In Conference On Learning Theory, pages 439–473. PMLR, 2018.
  61. 61.Max Simchowitz, Ross Boczar, and Benjamin Recht. Learning linear dynamical systems with semi-parametric least squares. In Conference on Learning Theory, pages 2714–2802. PMLR, 2019.
  62. 62.Sumeet Singh, Spencer M Richards, Vikas Sindhwani, Jean-Jacques E Slotine, and Marco Pavone. Learning stabilizable nonlinear dynamics with contraction-based regularization. The International Journal of Robotics Research, 40(10-11):1123–1150, 2021.
  63. 63.Anastasios Tsiamis and George J Pappas. Finite sample analysis of stochastic system identification. In 2019 IEEE 58th Conference on Decision and Control, pages 3648–3654. IEEE, 2019.
  64. 64.Anastasios Tsiamis, Nikolai Matni, and George Pappas. Sample complexity of kalman filtering for unknown systems. In Learning for Dynamics and Control, pages 435–444. PMLR, 2020.
  65. 65.Roman Vershynin. Introduction to the non-asymptotic analysis of random matrices. In Compressed Sensing: Theory and Applications, page 210–268. Cambridge University Press, 2012.
  66. 66.Andrew Wagenmaker and Kevin Jamieson. Active learning for identification of linear dynamical systems. In Conference on Learning Theory, pages 3487–3582. PMLR, 2020.
  67. 67.Greg Welch and Gary Bishop. An Introduction to the Kalman Filter. University of North Carolina at Chapel Hill, USA, 1995.
  68. 68.Bin Yu. Rates of convergence for empirical processes of stationary mixing sequences. The Annals of Probability, pages 94–116, 1994.
  69. 69.Shaofeng Zou, Tengyu Xu, and Yingbin Liang. Finite-sample analysis for sarsa with linear function approximation. In Advances in Neural Information Processing Systems, volume 32, pages 8668–8678, 2019.

Citation

MLA
Sattar, Y., and S. Oymak. “Non-asymptotic and Accurate Learning of Nonlinear Dynamical Systems”. Journal of Machine Learning Research, vol. 23, no. 140, 2022, pp. 1–9, https://www.jmlr.org/papers/v23/20-617.html.
APA
Sattar, Y., & Oymak, S. (2022). Non-asymptotic and Accurate Learning of Nonlinear Dynamical Systems. Journal of Machine Learning Research, 23(140), 1–49. https://www.jmlr.org/papers/v23/20-617.html
Chicago
Sattar, Y., and S. Oymak. 2022. “Non-asymptotic and Accurate Learning of Nonlinear Dynamical Systems”. Journal of Machine Learning Research 23 (140): 1–49. https://www.jmlr.org/papers/v23/20-617.html.
Harvard
Sattar, Y. and Oymak, S. (2022) “Non-asymptotic and Accurate Learning of Nonlinear Dynamical Systems”, Journal of Machine Learning Research, 23(140), pp. 1–49. Available at: https://www.jmlr.org/papers/v23/20-617.html.
Vancouver
1. Sattar Y, Oymak S (2022) Non-asymptotic and Accurate Learning of Nonlinear Dynamical Systems. Journal of Machine Learning Research 23:1–49

BibTeX

@article{JMLR:v23:20-617,
  author  = {Yahya Sattar and Samet Oymak},
  title   = {Non-asymptotic and Accurate Learning of Nonlinear Dynamical Systems},
  journal = {Journal of Machine Learning Research},
  year    = {2022},
  volume  = {23},
  number  = {140},
  pages   = {1--49},
  url     = {http://jmlr.org/papers/v23/20-617.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/