Generalized Teacher Forcing for Learning Chaotic Dynamics

Florian HessZahra MonfaredManuel BrennerDaniel Durstewitz

article2023ICML85 citations

Proves that a generalized teacher forcing scheme strictly bounds loss gradients during training on chaotic systems, enabling piecewise-linear recurrent neural networks to achieve state-of-the-art dynamical system reconstruction in minimal state dimensions.

Listen

Complex physical, biological, and societal systems—ranging from climate patterns and epidemiological spread to cardiac activity and brain function—are inherently chaotic. Scientists and decision-makers often seek to reconstruct the underlying dynamical systems from observed time series to understand core mechanisms and anticipate future behavior. However, training recurrent neural networks to capture these systems using standard gradient-based optimization frequently fails because trajectories in chaotic systems diverge exponentially fast, causing mathematical loss gradients to explode. Furthermore, existing models that attempt to stabilize training often suppress chaotic behavior entirely, require impractical prior knowledge of the system, or rely on high-dimensional architectures that obscure scientific interpretability.

The article demonstrates that combining a modified training strategy termed Generalized Teacher Forcing with a compact, one-hidden-layer piecewise-linear recurrent neural network solves the exploding gradient problem and enables accurate, low-dimensional reconstructions of chaotic systems. The primary objective is to prove mathematically that this method keeps loss gradients strictly bounded for arbitrary time horizons and to benchmark its empirical performance against state-of-the-art reconstruction algorithms on simulated and real-world datasets.

To evaluate this approach, the researchers proved the gradient-bounding properties analytically and tested the combined model across simulated benchmarks (such as the Lorenz-63 and Lorenz-96 atmospheric models) and challenging empirical datasets, including a 5-dimensional human electrocardiogram (ECG) and a 64-channel electroencephalogram (EEG). The method was compared across 20 independent runs against established architectures, including Long Short-Term Memory networks, Reservoir Computing, Sparse Identification of Nonlinear Dynamical Systems, and Neural Ordinary Differential Equations, evaluating both short-term prediction accuracy and long-term preservation of the system's geometric and temporal properties.

The results show that the proposed framework delivers superior reconstruction fidelity, particularly on noisy, real-world data where existing methods struggle. On the 64-channel EEG dataset, the shallow piecewise-linear model trained with Generalized Teacher Forcing captured true long-term dynamics using only 16 latent variables, whereas competing models failed to sustain realistic chaotic behavior, frequently collapsing into artificial static states or diverging. Additionally, the adaptive implementation automatically tunes the forcing strength during training without requiring prior estimates of system divergence rates, achieving results comparable to exhaustive manual parameter searches.

These findings indicate that organizations modeling complex or chaotic systems no longer need to compromise between model stability and interpretability. The proposed architecture retains exact mathematical tractability, allowing analysts to extract fixed points and periodic cycles semi-analytically while operating in much lower dimensions than conventional models. This significantly reduces computational overhead and provides faithful generative simulations for safety, forecasting, and policy analysis in complex domains.

Teams working on complex time series modeling should consider piloting the open-source shallow piecewise-linear recurrent neural network alongside the adaptive forcing protocol. Future work should investigate how to adapt Generalized Teacher Forcing to other recurrent network architectures and non-rectified activation functions, as the current performance benefits appear closely linked to the specific one-hidden-layer linear structure.

Confidence in the mathematical proofs and benchmark findings is high across the evaluated domains. However, users should exercise caution when applying the method to architectures with different activation functions or complex nonlinear observation models, as the tight bounding mechanisms and structural advantages may not transfer directly without further algorithmic tuning.

Cover for Generalized Teacher Forcing for Learning Chaotic Dynamics

Abstract

Chaotic dynamical systems (DS) are ubiquitous in nature and society. Often we are interested in reconstructing such systems from observed time series for prediction or mechanistic insight, where by reconstruction we mean learning geometrical and invariant temporal properties of the system in question (like attractors). However, training reconstruction algorithms like recurrent neural networks (RNNs) on such systems by gradient-descent based techniques faces severe challenges. This is mainly due to exploding gradients caused by the exponential divergence of trajectories in chaotic systems. Moreover, for (scientific) interpretability we wish to have as low dimensional reconstructions as possible, preferably in a model which is mathematically tractable. Here we report that a surprisingly simple modification of teacher forcing leads to provably strictly all-time bounded gradients in training on chaotic systems, and, when paired with a simple architectural rearrangement of a tractable RNN design, piecewise-linear RNNs (PLRNNs), allows for faithful reconstruction in spaces of at most the dimensionality of the observed system. We show on several DS that with these amendments we can reconstruct DS better than current SOTA algorithms, in much lower dimensions. Performance differences were particularly compelling on real world data with which most other methods severely struggled. This work thus led to a simple yet powerful DS reconstruction algorithm which is highly interpretable at the same time.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Theoretical Analysis & Methods
  • 3.1. Problem setting: Loss gradients and chaotic dynamics
  • 3.2. Generalized Teacher Forcing
  • 3.3. Reconstruction models
  • 3.4. Model training
  • 4. Results
  • 4.1. DS data sets
  • 4.2. Evaluation measures
  • 4.3. Experimental evaluation
  • 5. Conclusions
  • Acknowledgements
  • References
  • 6. Appendix
  • 6.1. Theorems: Proofs
  • 6.1.1. PROOF OF PROPOSITION 1
  • 6.1.2. PROOF OF PROPOSITION 2
  • 6.1.3. PROOF OF PROPOSITION 3
  • 6.2. Training Protocol
  • 6.3. Benchmark Systems and Real-World Data
  • 6.4. Details on Evaluation Measures
  • 6.5. GTF and RNN architecture
  • 6.6. More Details on Comparison Methods

Knowls

  1. Knowl 1 — Generalized teacher forcing controls recurrent Jacobians

    model/method

    For a recurrent map zt=Fθ(zt−1)z_t=F_\theta(z_{t-1}), where zt∈RMz_t\in\mathbb{R}^M is the model state and θ\theta denotes its parameters, generalized teacher forcing (GTF) interpolates between the model-generated state ztz_t and a data-inferred target state zˉt\bar z_t:

    z~t=(1−α)zt+αzˉt,0≤α≤1.\tilde z_t=(1-\alpha)z_t+\alpha\bar z_t,\qquad 0\leq\alpha\leq 1.

    The next recurrent update uses the forced state, zt=Fθ(z~t−1)z_t=F_\theta(\tilde z_{t-1}). If J~t=∂Fθ(z~t−1)/∂z~t−1\tilde J_t=\partial F_\theta(\tilde z_{t-1})/\partial\tilde z_{t-1} is the Jacobian evaluated at the forced state, then the Jacobian with respect to the preceding unforced state is

    ∂zt∂zt−1=(1−α)J~t.\frac{\partial z_t}{\partial z_{t-1}}=(1-\alpha)\tilde J_t.

    Consequently, gradients propagated across t−rt-r recurrent steps contain the factor (1−α)t−r(1-\alpha)^{t-r}:

    ∂zt∂zr=(1−α)t−r∏k=0t−r−1J~t−k.\frac{\partial z_t}{\partial z_r}=(1-\alpha)^{t-r}\prod_{k=0}^{t-r-1}\tilde J_{t-k}.

    Thus α=0\alpha=0 gives ordinary backpropagation through time, α=1\alpha=1 cuts all temporal gradient propagation, and intermediate values attenuate trajectory divergence without requiring the recurrent architecture itself to be non-chaotic.

  2. Knowl 2 — GTF yields a nondivergent gradient regime for chaotic RNNs

    theoretical result

    Let J={J~κ}κ∈K\mathcal{J}=\{\tilde J_\kappa\}_{\kappa\in K} be the set of all Jacobians encountered by a recurrent map, and define

    σ~max⁡=sup⁡J~κ∈Jσmax⁡(J~κ),λ~min⁡=inf⁡J~κ∈Jλmin⁡(J~κ),\tilde\sigma_{\max}=\sup_{\tilde J_\kappa\in\mathcal{J}}\sigma_{\max}(\tilde J_\kappa), \qquad \tilde\lambda_{\min}=\inf_{\tilde J_\kappa\in\mathcal{J}}\lambda_{\min}(\tilde J_\kappa),

    where σmax⁡(J)\sigma_{\max}(J) is the largest singular value of JJ and λmin⁡(J)\lambda_{\min}(J) is the smallest absolute eigenvalue of JJ. A recurrent system that is chaotic must satisfy σ~max⁡>1\tilde\sigma_{\max}>1, because a positive maximum Lyapunov exponent requires some local Jacobian expansion.

    For σ~max⁡≥1\tilde\sigma_{\max}\geq 1, choosing

    α∗=1−1σ~max⁡\alpha^*=1-\frac{1}{\tilde\sigma_{\max}}

    makes every GTF Jacobian product nondivergent as the temporal separation t−rt-r tends to infinity. More specifically, the norm is bounded above by

    ∥∂zt∂zr∥2≤((1−α)σ~max⁡)t−r,\left\|\frac{\partial z_t}{\partial z_r}\right\|_2 \leq \big((1-\alpha)\tilde\sigma_{\max}\big)^{t-r},

    which is at most 11 when α=α∗\alpha=\alpha^*. If γ~=λ~min⁡/σ~max⁡=1\tilde\gamma=\tilde\lambda_{\min}/\tilde\sigma_{\max}=1, the gradient product neither explodes nor vanishes asymptotically. If γ~≠1\tilde\gamma\neq1, it still cannot explode, although it may vanish; when it vanishes, a value of γ~\tilde\gamma closer to 11 gives slower decay. The result applies over arbitrarily long training horizons while leaving the recurrent map free to retain positive Lyapunov exponents and hence chaotic dynamics.

  3. Knowl 3 — Shallow PLRNN provides a low-dimensional tractable reconstruction model

    model/method

    The reconstruction model is a shallow piecewise-linear recurrent neural network (shPLRNN) with latent state zt∈RMz_t\in\mathbb{R}^M:

    zt=Azt−1+W1ϕ(W2zt−1+h2)+h1,z_t=Az_{t-1}+W_1\phi(W_2z_{t-1}+h_2)+h_1,

    where A∈RM×MA\in\mathbb{R}^{M\times M} is diagonal, W1∈RM×LW_1\in\mathbb{R}^{M\times L}, W2∈RL×MW_2\in\mathbb{R}^{L\times M}, h1∈RMh_1\in\mathbb{R}^M, h2∈RLh_2\in\mathbb{R}^L, L>ML>M, and ϕ(u)=max⁡(0,u)\phi(u)=\max(0,u) is applied componentwise. A linear observation map produces NN observed variables through x^t=Bozt\hat x_t=B_o z_t, with Bo∈RN×MB_o\in\mathbb{R}^{N\times M}.

    The version used to prevent unbounded growth from the ReLU units is

    zt=Azt−1+W1[ϕ(W2zt−1+h2)−ϕ(W2zt−1)]+h1.z_t=Az_{t-1}+W_1\left[\phi(W_2z_{t-1}+h_2)-\phi(W_2z_{t-1})\right]+h_1.

    For diagonal AA, every orbit of this clipped model is bounded when ∥A∥2<1\|A\|_2<1. The shPLRNN is equivalent to a dendritic PLRNN with the same latent dimension, and an MM-dimensional dendritic PLRNN with BB spline bases can be represented by an shPLRNN with hidden-layer size L=MBL=MB. Therefore, fixed points and cycles of the shPLRNN can be computed by the same piecewise-linear, semi-analytic procedures used for dendritic PLRNNs, while the shallow arrangement permits reconstructions with no more latent variables than the observed system and sometimes fewer.

  4. Knowl 4 — Adaptive GTF estimates forcing from teacher-signal Jacobians

    algorithm

    Adaptive GTF (aGTF) selects the forcing strength from Jacobians evaluated at data-inferred teacher states rather than enumerating all recurrent linear regions. For a batch of teacher sequences {z^1:T(p)}p=1S\{\hat z^{(p)}_{1:T}\}_{p=1}^S, let J^t(p)\hat J^{(p)}_t be the shPLRNN Jacobian at z^t(p)\hat z^{(p)}_t. The method estimates the geometric mean of the Jacobians using the arithmetic approximation

    G(J^T:2(p))≈1T−1∑t=2TJ^t(p),\mathcal{G}(\hat J^{(p)}_{T:2})\approx\frac{1}{T-1}\sum_{t=2}^{T}\hat J^{(p)}_t,

    and sets the batch forcing proposal to α=max⁡(0,1−1/κ)\alpha=\max\left(0,1-1/\kappa\right), where κ=max⁡p∥G(J^T:2(p))∥2\kappa=\max_p\|\mathcal{G}(\hat J^{(p)}_{T:2})\|_2. The paper also considers an exp-log approximation to this geometric mean and a norm upper bound, but finds the arithmetic approximation substantially faster with similar estimates.

    Input: Previous forcing estimate αn−1\alpha_{n-1}, update interval kk, decay rate γ\gamma, teacher-state batch, recurrent map FθnF_{\theta_n}
    Output: Current forcing estimate αn\alpha_n
    If nn is divisible by kk:
        Compute J^t(p)=Jacobian⁡(Fθn,z^t(p))\hat J^{(p)}_t=\operatorname{Jacobian}(F_{\theta_n},\hat z^{(p)}_t) for every sequence and time step
        For each sequence, approximate G\mathcal{G} by (T−1)−1∑t=2TJ^t(p)(T-1)^{-1}\sum_{t=2}^{T}\hat J^{(p)}_t
        Set κ\kappa to the largest spectral norm of these approximations
        Set proposal α=max⁡(0,1−1/κ)\alpha=\max(0,1-1/\kappa)
        If proposal α>αn−1\alpha>\alpha_{n-1}, set αn=α\alpha_n=\alpha
        Otherwise set αn=(1−γ)α+γαn−1\alpha_n=(1-\gamma)\alpha+\gamma\alpha_{n-1}
    Else:
        Set αn=αn−1\alpha_n=\alpha_{n-1}
    Return αn\alpha_n

    The experiments use α0=1\alpha_0=1, k=5k=5, and γ=0.999\gamma=0.999. The exponential moving average makes forcing decay from strong initial teacher control while allowing immediate increases if a parameter update produces a sudden increase in local expansion. GTF is used only during training; autonomous test-time trajectories receive no teacher states.

  5. Knowl 5 — BPTT training uses inverted observations as teacher signals

    experimental setup

    Given observed sequences x1:T∈RN×Tx_{1:T}\in\mathbb{R}^{N\times T}, the linear observation model x^t=Bozt\hat x_t=B_o z_t supplies latent teacher signals through the Moore–Penrose inverse:

    z^t=Bo+xt.\hat z_t=B_o^+x_t.

    For each optimization update, the method samples a batch of S=16S=16 random subsequences of length T~\tilde T, initializes each recurrent trajectory at its first teacher state z^1(p)\hat z_1^{(p)}, and propagates with GTF. Predictions are mapped back to observation space and optimized with

    LMSE=1S(T~−1)∑p=1S∑t=2T~∥x~t(p)−x^t(p)∥22.\mathcal{L}_{\mathrm{MSE}}=\frac{1}{S(\tilde T-1)}\sum_{p=1}^{S}\sum_{t=2}^{\tilde T}\left\|\tilde x_t^{(p)}-\hat x_t^{(p)}\right\|_2^2.

    The optimizer is RAdam; the learning rate decays exponentially from 10−310^{-3} to 10−610^{-6}. Each run lasts 5000 epochs, with 50 batches per epoch, totaling 250,000 parameter updates. The sequence length is T~=200\tilde T=200 for Lorenz-63, Lorenz-96, and ECG, and T~=50\tilde T=50 for EEG. The experiments use noisy simulated Lorenz systems, a five-dimensional delay embedding of ECG, a 64-channel EEG recording, and a partially observed multiscale Lorenz-96 system. At test time, models are initialized from data only and then evolved freely.

  6. Knowl 6 — Reconstruction is evaluated by attractor geometry, spectra, and short-term prediction

    equation

    The paper evaluates whether a trained generative model reproduces invariant dynamics rather than relying only on pointwise forecasts. The state-space divergence is the Kullback–Leibler divergence between the probability distributions p(x)p(x) and q(x)q(x) induced by long true and generated trajectories:

    Dstsp=DKL(p∥q)=∫p(x)log⁡p(x)q(x) dx.D_{\mathrm{stsp}}=D_{\mathrm{KL}}(p\|q)=\int p(x)\log\frac{p(x)}{q(x)}\,dx.

    For low-dimensional systems, the distributions are estimated by binning the observed and generated attractors; for high-dimensional systems, they are represented as Gaussian mixtures centered on trajectory points with covariance Σ=σ2I\Sigma=\sigma^2I, using σ2=1\sigma^2=1.

    For each observed variable ii, let f^i\hat f_i and g^i\hat g_i be normalized, smoothed power spectra of the true and generated signals. Their Hellinger distance is

    H(f^i,g^i)=12∥f^i−g^i∥2,H(\hat f_i,\hat g_i)=\frac{1}{\sqrt{2}}\left\|\sqrt{\hat f_i}-\sqrt{\hat g_i}\right\|_2,

    and the reported temporal error is the average

    DH=1N∑i=1NH(f^i,g^i).D_H=\frac{1}{N}\sum_{i=1}^{N}H(\hat f_i,\hat g_i).

    The nn-step prediction error is

    PE(n)=1N(T−n)∑t=1T−n∥xt+n−x^t+n∥22.\mathrm{PE}(n)=\frac{1}{N(T-n)}\sum_{t=1}^{T-n}\left\|x_{t+n}-\hat x_{t+n}\right\|_2^2.

    Lower values are better for all three measures. For invariant statistics, each model is freely simulated for 1.25T1.25T steps, the first 0.25T0.25T transient steps are discarded, and the remaining TT steps are compared with the data. Because chaotic trajectories diverge exponentially, the paper treats DstspD_{\mathrm{stsp}} and DHD_H as more informative for reconstruction than prediction error alone.

  7. Knowl 7 — GTF substantially improves ECG and EEG reconstruction over competing methods

    data/table

    The page-8 comparison reports medians ± median absolute deviations over 20 independent training runs. DstspD_{\mathrm{stsp}}, DHD_H, and PE(20)\mathrm{PE}(20) are state-space, temporal-spectral, and 20-step prediction errors; lower is better. The shPLRNN trained with either fixed GTF or adaptive GTF uses far fewer latent variables than competing models while achieving the strongest results on both empirical datasets.

    Could not parse LaTeX table

    On ECG, fixed GTF attains the best state-space and prediction scores among the listed methods, while on EEG it is substantially better in attractor and spectral reconstruction than all alternatives. Adaptive GTF performs comparably to grid-selected fixed GTF without requiring a forcing-strength grid search.

  8. Knowl 8 — Performance remains competitive on simulated and partially observed systems

    empirical result

    On Lorenz-63, Lorenz-96, and multiscale Lorenz-96, the shPLRNN trained with GTF performs at least about as well as the comparison methods, although the simulated benchmarks exhibit a ceiling effect because many reconstruction algorithms solve them reasonably well.

    For the partially observed multiscale Lorenz-96 benchmark, the underlying system has I=J=K=8I=J=K=8 variables at three scales, totaling 584 dynamical variables, but only the K=8K=8 slow variables are observed. Across 20 runs, the reported results are:

    Could not parse LaTeX table

    The three methods produce statistically indistinguishable long-term reconstructions on this partially observed system: the paper reports a Kruskal–Wallis test with p=0.84p=0.84, with differences within a 10% median-absolute-deviation and 34% standard-error margin. The GTF shPLRNN nevertheless uses only the observed eight-dimensional state, compared with 25 latent variables for the dendritic PLRNN.

  9. Knowl 9 — Long-term attractor structure, not just short forecasts, separates the methods

    empirical result

    The freely generated traces shown on page 9 demonstrate that the shPLRNN trained with GTF preserves the temporal structure of the empirical EEG signal over long autonomous runs, and the dendPLRNN with identity teacher forcing does so somewhat less consistently. In contrast, most LSTM, reservoir, Neural-ODE, and LEM reconstructions either collapse toward an equilibrium-like behavior or diverge, despite often producing reasonable short-term predictions.

    This distinction is important because a chaotic system can have useful short-term prediction error while failing to reproduce the invariant attractor geometry and power-spectrum structure. The GTF shPLRNN reconstructs the 64-channel EEG with only 16 latent dynamical variables, whereas the dendPLRNN requires 105 and several competing methods use 64 or more. Across approximately 90% of the selected EEG runs, the GTF shPLRNN and dendPLRNN yielded similarly good qualitative traces, while the other methods had highly variable and generally poor long-term behavior.

  10. Knowl 10 — GTF performance depends on the recurrent architecture and activation

    limitation

    Although the Jacobian argument for GTF is architecture-independent in principle, the empirical benefit is not. A direct GTF implementation for dendPLRNNs did not produce the same improvement over sparse teacher forcing as the shPLRNN implementation. The dendPLRNN latent state is higher-dimensional than the observations, so projecting an observation into it is underdetermined; forcing only the observed coordinates or using a trainable linear inverse did not remove this difficulty.

    Replacing the shPLRNN ReLU activation with tanh⁡\tanh also degraded performance. On ECG, shPLRNN+GTF with tanh⁡\tanh reached Dstsp=7.8±2.1D_{\mathrm{stsp}}=7.8\pm2.1, DH=0.37±0.03D_H=0.37\pm0.03, and PE(20)=(8.6±0.3)⋅10−3\mathrm{PE}(20)=(8.6\pm0.3)\cdot10^{-3}, compared with 4.3±0.64.3\pm0.6, 0.34±0.020.34\pm0.02, and (2.4±0.1)⋅10−3(2.4\pm0.1)\cdot10^{-3} for ReLU GTF. On EEG, the corresponding tanh⁡\tanh values were 17±317\pm3, 0.37±0.070.37\pm0.07, and (6.6±0.1)⋅10−1(6.6\pm0.1)\cdot10^{-1}, versus 2.1±0.22.1\pm0.2, 0.11±0.010.11\pm0.01, and (5.5±0.1)⋅10−1(5.5\pm0.1)\cdot10^{-1} for ReLU.

    The authors also note that sparse teacher forcing requires a predictability-time estimate, and their STF comparison did not systematically optimize the forcing interval. How to implement GTF most effectively for other architectures, and how it compares with STF under fully optimized settings, remains unresolved.

Coverage note — Detailed proofs, comparator-specific hyperparameter grids, preprocessing minutiae, auxiliary regularizers, and the full Lorenz-63/Lorenz-96 comparison tables were omitted because they support the main theory and empirical claims without adding separate load-bearing contributions.

References

  1. 1.Abarbanel, H. Predicting the future: completing models of observed complex systems. Springer, 2013.
  2. 2.Ardizzone, L., Kruse, J., Rother, C., and Kothe, U. Analyzing inverse problems with invertible neural networks. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=rJed6j0cKX.
  3. 3.Arjovsky, M., Shah, A., and Bengio, Y. Unitary evolution recurrent neural networks. In International conference on machine learning, pp. 1120–1128. PMLR, 2016.
  4. 4.Bakarji, J., Champion, K., Kutz, J. N., and Brunton, S. L. Discovering governing equations from partial measurements with deep delay autoencoders. arXiv preprint arXiv:2201.05136, 2022.
  5. 5.Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N. Scheduled sampling for sequence prediction with recurrent neural networks. Advances in neural information processing systems, 28, 2015.
  6. 6.Bengio, Y., Simard, P., and Frasconi, P. Learning long-term dependencies with gradient descent is difficult. IEEE transactions on neural networks, 5(2):157–166, 1994.
  7. 7.Bottcher, A. and Wenzel, D. The frobenius norm and the commutator. Linear algebra and its applications, 429(8-9):1864–1885, 2008.
  8. 8.Brenner, M., Hess, F., Mikhaeil, J. M., Bereska, L. F., Monfared, Z., Kuo, P.-C., and Durstewitz, D. Tractable dendritic RNNs for reconstructing nonlinear dynamical systems. In Proceedings of the 39th International Conference on Machine Learning, volume 162, pp. 2292–2320. PMLR, 2022.
  9. 9.Brunton, S. L., Proctor, J. L., and Kutz, J. N. Discovering governing equations from data by sparse identification of nonlinear dynamical systems. Proceedings of the national academy of sciences, 113(15):3932–3937, 2016.
  10. 10.Brunton, S. L., Budisic, M., Kaiser, E., and Kutz, J. N. Modern koopman theory for dynamical systems. SIAM Review, 64(2):229–340, 2022.
  11. 11.Bury, T. M., Sujith, R. I., Pavithran, I., Scheffer, M., Lenton, T. M., Anand, M., and Bauch, C. T. Deep learning for early warning signals of tipping points. Proceedings of the National Academy of Sciences, 118(39):e2106140118, 2021. doi: 10.1073/pnas.2106140118. URL https://www.pnas.org/doi/10.1073/pnas.2106140118.
  12. 12.Champion, K., Lusch, B., Kutz, J. N., and Brunton, S. L. Data-driven discovery of coordinates and governing equations. Proceedings of the National Academy of Sciences, 116(45):22445–22451, 2019.
  13. 13.Chang, B., Chen, M., Haber, E., and Chi, E. H. AntisymmetricRNN: A dynamical system view on recurrent neural networks. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=ryxepo0cFX.
  14. 14.Chattopadhyay, A., Hassanzadeh, P., and Subramanian, D. Data-driven predictions of a multiscale lorenz 96 chaotic system using machine-learning methods: reservoir computing, artificial neural network, and long short-term memory network. Nonlinear Processes in Geophysics, 27(3):373–389, 2020.
  15. 15.Chen, R. T., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K. Neural ordinary differential equations. Advances in neural information processing systems, 31, 2018.
  16. 16.Cho, K., Van Merrienboer, B., Bahdanau, D., and Bengio, Y. On the properties of neural machine translation: Encoder-decoder approaches. arXiv preprint arXiv:1409.1259, 2014.
  17. 17.Cooley, J. W. and Tukey, J. W. An algorithm for the machine calculation of complex fourier series. Mathematics of Computation, 19(90):297–301, 1965. ISSN 0025-5718. doi: 10.2307/2003354. URL https://www.jstor.org/stable/2003354. Publisher: American Mathematical Society.
  18. 18.Datseris, G. Dynamicalsystems.jl: A julia software library for chaos and nonlinear dynamics. Journal of Open Source Software, 3(23):598, mar 2018. doi: 10.21105/joss.00598. URL https://doi.org/10.21105/joss.00598.
  19. 19.de Silva, B., Champion, K., Quade, M., Loiseau, J.-C., Kutz, J., and Brunton, S. Pysindy: A python package for the sparse identification of nonlinear dynamical systems from data. Journal of Open Source Software, 5(49):2104, 2020. doi: 10.21105/joss.02104. URL https://doi.org/10.21105/joss.02104.
  20. 20.Dinh, L., Sohl-Dickstein, J., and Bengio, S. Density estimation using real NVP. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=HkpbnH9lx.
  21. 21.Doya, K. et al. Bifurcations in the learning of recurrent neural networks 3. learning (RTRL), 3:17, 1992.
  22. 22.Duarte, J., Januario, C., Martins, N., and Sardanyés, J. Quantifying chaos for ecological stoichiometry. Chaos: An Interdisciplinary Journal of Nonlinear Science, 20(3):033105, September 2010.
  23. 23.Durstewitz, D. A state space approach for piecewise-linear recurrent neural networks for identifying computational dynamics from neural measurements. PLoS computational biology, 13(6):e1005542, 2017.
  24. 24.Durstewitz, D. and Gabriel, T. Dynamical basis of irregular spiking in nmda-driven prefrontal cortex neurons. Cerebral cortex, 17(4):894–908, 2007a.
  25. 25.Durstewitz, D. and Gabriel, T. Dynamical basis of irregular spiking in NMDA-driven prefrontal cortex neurons. Cerebral Cortex (New York, N.Y.: 1991), 17(4):894–908, 2007b. ISSN 1047-3211. doi: 10.1093/cercor/bhk044.
  26. 26.Engelken, R., Wolf, F., and Abbott, L. F. Lyapunov spectra of chaotic recurrent neural networks. arXiv preprint arXiv:2006.02427, 2020.
  27. 27.Erichson, N. B., Azencot, O., Queiruga, A., Hodgkinson, L., and Mahoney, M. W. Lipschitz recurrent neural networks. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=-N7PBXqOUJZ.
  28. 28.Faggini, M. Chaotic time series analysis in economics: Balance and perspectives. Chaos: An Interdisciplinary Journal of Nonlinear Science, 24(4):042101, 2014. Publisher: American Institute of Physics.
  29. 29.Funahashi, K.-i. and Nakamura, Y. Approximation of dynamical systems by continuous time recurrent neural networks. Neural networks, 6(6):801–806, 1993.
  30. 30.Goldberger, A. L., Amaral, L. A., Glass, L., Hausdorff, J. M., Ivanov, P. C., Mark, R. G., Mietus, J. E., Moody, G. B., Peng, C.-K., and Stanley, H. E. Physiobank, physiotoolkit, and physionet: components of a new research resource for complex physiologic signals. circulation, 101(23):e215–e220, 2000.
  31. 31.Govindan, R., Narayanan, K., and Gopinathan, M. On the evidence of deterministic chaos in ecg: Surrogate and predictability analysis. Chaos: An Interdisciplinary Journal of Nonlinear Science, 8(2):495–502, 1998.
  32. 32.Hanson, J. and Raginsky, M. Universal simulation of stable dynamical systems by recurrent neural nets. In Learning for Dynamics and Control, pp. 384–392. PMLR, 2020.
  33. 33.Hernandez, D., Moretti, A. K., Saxena, S., Wei, Z., Cunningham, J., and Paninski, L. Nonlinear evolution via spatially-dependent linear dynamics for electrophysiology and calcium data. Neurons, Behavior, Data analysis, and Theory, 3(3):13476, 2020.
  34. 34.Hershey, J. R. and Olsen, P. A. Approximating the kullback leibler divergence between gaussian mixture models. In 2007 IEEE International Conference on Acoustics, Speech and Signal Processing-ICASSP’07, volume 4, pp. IV–317. IEEE, 2007.
  35. 35.Higham, N. J. and Lin, L. A schur-pade algorithm for fractional powers of a matrix. SIAM Journal on Matrix Analysis and Applications, 32(3):1056–1078, 2011.
  36. 36.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  37. 37.Hochreiter, S., Bengio, Y., Frasconi, P., Schmidhuber, J., et al. Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001.
  38. 38.Inoue, K., Ohara, S., Kuniyoshi, Y., and Nakajima, K. Transient chaos in bert. arXiv:2106.03181, 2021.
  39. 39.Jordan, M. I. Attractor dynamics and parallelism in a connectionist sequential machine. In Artificial neural networks: concept learning, pp. 112–127. IEEE Press, 1990.
  40. 40.Kamdjeu Kengne, L., Mboupda Pone, J. R., and Fotsin, H. B. On the dynamics of chaotic circuits based on memristive diode-bridge with variable symmetry: A case study. Chaos, Solitons & Fractals, 145:110795, 2021.
  41. 41.Kantz, H. A robust method to estimate the maximal lyapunov exponent of a time series. Physics letters A, 185(1):77–87, 1994.
  42. 42.Kaptanoglu, A. A., de Silva, B. M., Fasel, U., Kaheman, K., Goldschmidt, A. J., Callaham, J., Delahunt, C. B., Nicolaou, Z. G., Champion, K., Loiseau, J.-C., Kutz, J. N., and Brunton, S. L. Pysindy: A comprehensive python package for robust sparse system identification. Journal of Open Source Software, 7(69):3994, 2022. doi: 10.21105/joss.03994. URL https://doi.org/10.21105/joss.03994.
  43. 43.Kesmia, M., Boughaba, S., and Jacquir, S. Control of continuous dynamical systems modeling physiological states. Chaos, Solitons & Fractals, 136:109805, 2020.
  44. 44.Kimura, M. and Nakano, R. Learning dynamical systems by recurrent neural networks from orbits. Neural Networks, 11(9):1589–1599, 1998.
  45. 45.Kirchmeyer, M., Yin, Y., Dona, J., Baskiotis, N., Rakotomamonjy, A., and Gallinari, P. Generalizing to new physical systems via context-informed dynamics model. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 11283–11301. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/kirchmeyer22a.html.
  46. 46.Koppe, G., Toutounji, H., Kirsch, P., Lis, S., and Durstewitz, D. Identifying nonlinear dynamical systems via generative recurrent neural networks with applications to fmri. PLoS computational biology, 15(8):e1007263, 2019.
  47. 47.Kramer, D., Bommer, P. L., Durstewitz, D., Tombolini, C., and Koppe, G. Reconstructing nonlinear dynamical systems from multi-modal time series. In International Conference on Machine Learning, pp. 11613–11633. PMLR, 2022.
  48. 48.Kramer, K.-H., Datseris, G., Kurths, J., Kiss, I. Z., Ocampo-Espindola, J. L., and Marwan, N. A unified and automated approach to attractor reconstruction. New Journal of Physics, 23(3):033017, 2021.
  49. 49.Lejarza, F. and Baldea, M. Data-driven discovery of the governing equations of dynamical systems via moving horizon optimization. Scientific Reports, 12(1):1–15, 2022.
  50. 50.Li, Z., Kovachki, N. B., Azizzadenesheli, K., liu, B., Bhattacharya, K., Stuart, A., and Anandkumar, A. Fourier neural operator for parametric partial differential equations. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=c8P9NQVtmnO.
  51. 51.Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J. On the variance of the adaptive learning rate and beyond. In Proceedings of the Eighth International Conference on Learning Representations (ICLR 2020), April 2020.
  52. 52.Lorenz, E. N. Deterministic nonperiodic flow. Journal of atmospheric sciences, 20(2):130–141, 1963.
  53. 53.Lorenz, E. N. Predictability: A problem partly solved. In Proc. Seminar on predictability, volume 1, 1996.
  54. 54.Lusch, B., Kutz, J. N., and Brunton, S. L. Deep learning for universal linear embeddings of nonlinear dynamics. Nature communications, 9(1):1–10, 2018.
  55. 55.Mangiarotti, S., Peyre, M., Zhang, Y., Huc, M., Roger, F., and Kerr, Y. Chaos theory applied to the outbreak of COVID-19: an ancillary approach to decision making in pandemic context. Epidemiology and Infection, 148:e95, 2020.
  56. 56.Mikhaeil, J. M., Monfared, Z., and Durstewitz, D. On the difficulty of learning chaotic dynamics with RNNs. In Proceedings of the 36th Conference on Neural Information Processing Systems. NeurIPS, 2022.
  57. 57.Monfared, Z. and Durstewitz, D. Transformation of relu-based recurrent neural networks from discrete-time to continuous-time. In International Conference on Machine Learning, pp. 6999–7009. PMLR, 2020.
  58. 58.Olsen, L. F., Truty, G. L., and Schaffer, W. M. Oscillations and chaos in epidemics: a nonlinear dynamic study of six childhood diseases in Copenhagen, Denmark. Theoretical Population Biology, 33(3):344–370, 1988.
  59. 59.Otto, S. E. and Rowley, C. W. Linearly recurrent autoencoder networks for learning dynamics. SIAM Journal on Applied Dynamical Systems, 18(1):558–593, 2019. doi: 10.1137/18M1177846. URL https://epubs.siam.org/doi/abs/10.1137/18M1177846. Publisher: Society for Industrial and Applied Mathematics.
  60. 60.Patel, D. and Ott, E. Using machine learning to anticipate tipping points and extrapolate to post-tipping dynamics of non-stationary dynamical systems. arXiv preprint arXiv:2207.00521, 2022.
  61. 61.Pathak, J., Hunt, B., Girvan, M., Lu, Z., and Ott, E. Model-free prediction of large spatiotemporally chaotic systems from data: A reservoir computing approach. Physical review letters, 120(2):024102, 2018.
  62. 62.Pearlmutter, B. A. Learning state space trajectories in recurrent neural networks. Neural Computation, 1(2):263–269, 1989.
  63. 63.Perko, L. Differential Equations and Dynamical Systems, volume 7. Springer, New York, NY, 2001.
  64. 64.Pineda, F. J. Dynamics and architecture for neural computation. Journal of Complexity, 4(3):216–245, 1988.
  65. 65.Raissi, M. Deep hidden physics models: Deep learning of nonlinear partial differential equations. The Journal of Machine Learning Research, 19(1):932–955, 2018.
  66. 66.Reiss, A., Indlekofer, I., Schmidt, P., and Van Laerhoven, K. Deep ppg: Large-scale heart rate estimation with convolutional neural networks. Sensors, 19(14), 2019. ISSN 1424-8220. doi: 10.3390/s19143079. URL https://www.mdpi.com/1424-8220/19/14/3079.
  67. 67.Ricci, M., Moriel, N., Piran, Z., and Nitzan, M. Phase2vec: Dynamical systems embedding with a physics-informed convolutional network. arXiv preprint arXiv:2212.03857, 2022.
  68. 68.Rubanova, Y., Chen, R. T., and Duvenaud, D. K. Latent ordinary differential equations for irregularly-sampled time series. Advances in neural information processing systems, 32, 2019.
  69. 69.Rumelhart, D. E., Hinton, G. E., and Williams, R. J. Learning representations by back-propagating errors. nature, 323(6088):533–536, 1986.
  70. 70.Rusch, T. K. and Mishra, S. Coupled oscillatory recurrent neural network (cornn): An accurate and (gradient) stable architecture for learning long time dependencies. arXiv preprint arXiv:2010.00951, 2020.
  71. 71.Rusch, T. K. and Mishra, S. Coupled oscillatory recurrent neural network (co{rnn}): An accurate and (gradient) stable architecture for learning long time dependencies. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=F3s69XzWOia.
  72. 72.Rusch, T. K., Mishra, S., Erichson, N. B., and Mahoney, M. W. Long expressive memory for sequence modeling. In International Conference on Learning Representations, 2022.
  73. 73.Sauer, T., Yorke, J. A., and Casdagli, M. Embedology. Journal of statistical Physics, 65(3):579–616, 1991.
  74. 74.Schalk, G., McFarland, D. J., Hinterberger, T., Birbaumer, N., and Wolpaw, J. R. BCI2000: a general-purpose brain-computer interface (BCI) system. IEEE transactions on bio-medical engineering, 51(6):1034–1043, 2000. ISSN 0018-9294. doi: 10.1109/TBME.2004.827072.
  75. 75.Schmidt, D., Koppe, G., Monfared, Z., Beutelspacher, M., and Durstewitz, D. Identifying nonlinear dynamical systems with multiple time scales and long-range dependencies. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=_XYzwxPIQu6.
  76. 76.Skokos, C. H., Gottwald, G. A., and Laskar, J. Chaos Detection and Predictability, volume 1. Springer, 2016.
  77. 77.Smith, J., Linderman, S., and Sussillo, D. Reverse engineering recurrent neural networks with jacobian switching linear dynamical systems. Advances in Neural Information Processing Systems, 34:16700–16713, 2021.
  78. 78.Takens, F. Detecting strange attractors in turbulence. In Dynamical systems and turbulence, Warwick 1980, pp. 366–381. Springer, 1981.
  79. 79.Thornes, T., Duben, P., and Palmer, T. On the use of scale-dependent precision in earth system modelling. Quarterly Journal of the Royal Meteorological Society, 143(703):897–908, 2017.
  80. 80.Van Vreeswijk, C. and Sompolinsky, H. Chaos in neuronal networks with balanced excitatory and inhibitory activity. Science (New York, N.Y.), 274(5293):1724–1726, 1996. ISSN 0036-8075. doi: 10.1126/science.274.5293.1724.
  81. 81.Vlachas, P. R. and Koumoutsakos, P. Learning from predictions: Fusing training and autoregressive inference for long-term spatiotemporal forecasts. arXiv preprint arXiv:2302.11101, 2023.
  82. 82.Vlachas, P. R., Byeon, W., Wan, Z. Y., Sapsis, T. P., and Koumoutsakos, P. Data-driven forecasting of high-dimensional chaotic systems with long short-term memory networks. Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences, 474(2213):20170844, 2018.
  83. 83.Vlachas, P. R., Pathak, J., Hunt, B. R., Sapsis, T. P., Girvan, M., Ott, E., and Koumoutsakos, P. Backpropagation algorithms and reservoir computing in recurrent neural networks for the forecasting of complex spatiotemporal dynamics. Neural Networks, 126:191–217, 2020.
  84. 84.Voss, H. U., Timmer, J., and Kurths, J. Nonlinear dynamical system identification from uncertain and indirect measurements. International Journal of Bifurcation and Chaos, 14(06):1905–1933, 2004.
  85. 85.Werbos, P. J. Backpropagation through time: what it does and how to do it. Proceedings of the IEEE, 78(10):1550–1560, 1990.
  86. 86.Williams, R. J. and Zipser, D. A learning algorithm for continually running fully recurrent neural networks. Neural computation, 1(2):270–280, 1989.
  87. 87.Xu, X., Chen, Z., Si, G., Hu, X., Jiang, Y., and Xu, X. The chaotic dynamics of the social behavior selection networks in crowd simulation. Nonlinear Dynamics, 64(1):117–126, 2011. ISSN 1573-269X. doi: 10.1007/s11071-010-9850-z. URL https://doi.org/10.1007/s11071-010-9850-z.
  88. 88.Zhao, Y. and Park, I. M. Variational online learning of neural dynamics. Frontiers in computational neuroscience, 14:71, 2020.

Citation

MLA
Hess, F., et al. “Generalized Teacher Forcing for Learning Chaotic Dynamics”. International Conference on Machine Learning, vol. 202, 2023, pp. 13017–49, https://proceedings.mlr.press/v202/hess23a.html.
APA
Hess, F., Monfared, Z., Brenner, M., & Durstewitz, D. (2023). Generalized Teacher Forcing for Learning Chaotic Dynamics. International Conference on Machine Learning, 202, 13017–13049. https://proceedings.mlr.press/v202/hess23a.html
Chicago
Hess, F., Z. Monfared, M. Brenner, and D. Durstewitz. 2023. “Generalized Teacher Forcing for Learning Chaotic Dynamics”. International Conference on Machine Learning 202: 13017–49. https://proceedings.mlr.press/v202/hess23a.html.
Harvard
Hess, F. et al. (2023) “Generalized Teacher Forcing for Learning Chaotic Dynamics”, International Conference on Machine Learning. PMLR, pp. 13017–13049. Available at: https://proceedings.mlr.press/v202/hess23a.html.
Vancouver
1. Hess F, Monfared Z, Brenner M, Durstewitz D (2023) Generalized Teacher Forcing for Learning Chaotic Dynamics. In: International Conference on Machine Learning. PMLR, pp 13017–13049

BibTeX

@InProceedings{pmlr-v202-hess23a,
  title = 	 {Generalized Teacher Forcing for Learning Chaotic Dynamics},
  author =       {Hess, Florian and Monfared, Zahra and Brenner, Manuel and Durstewitz, Daniel},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {13017--13049},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/hess23a/hess23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/hess23a.html},
  abstract = 	 {Chaotic dynamical systems (DS) are ubiquitous in nature and society. Often we are interested in reconstructing such systems from observed time series for prediction or mechanistic insight, where by reconstruction we mean learning geometrical and invariant temporal properties of the system in question (like attractors). However, training reconstruction algorithms like recurrent neural networks (RNNs) on such systems by gradient-descent based techniques faces severe challenges. This is mainly due to exploding gradients caused by the exponential divergence of trajectories in chaotic systems. Moreover, for (scientific) interpretability we wish to have as low dimensional reconstructions as possible, preferably in a model which is mathematically tractable. Here we report that a surprisingly simple modification of teacher forcing leads to provably strictly all-time bounded gradients in training on chaotic systems, and, when paired with a simple architectural rearrangement of a tractable RNN design, piecewise-linear RNNs (PLRNNs), allows for faithful reconstruction in spaces of at most the dimensionality of the observed system. We show on several DS that with these amendments we can reconstruct DS better than current SOTA algorithms, in much lower dimensions. Performance differences were particularly compelling on real world data with which most other methods severely struggled. This work thus led to a simple yet powerful DS reconstruction algorithm which is highly interpretable at the same time.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/