Generalizing to New Physical Systems via Context-Informed Dynamics Model

Matthieu KirchmeyerYuan YinJérémie DonàNicolas BaskiotisAlain RakotomamonjyPatrick Gallinari

article2022ICML54 citationsOutstanding Paper Award

Proposes a context-informed dynamics adaptation framework that uses low-rank hypernetworks to rapidly generalize physical models to unseen environments from only a few trajectory observations.

Listen

Machine learning models are increasingly used to emulate complex physical dynamics in domains like climate forecasting, fluid mechanics, and biomedicine. However, standard machine learning methods assume that training and testing data come from the identical underlying distribution. In real-world physical systems, conditions change: physical coefficients, boundary conditions, and external forces vary across different environments. When conventional models encounter an unseen physical system governed by different parameters, they fail to generalize or require substantial retraining data that is rarely available.

The article introduces and evaluates a machine learning framework called Context-Informed Dynamics Adaptation (CoDA). The objective of the article is to demonstrate how a dynamics model can rapidly and accurately adapt to new, unseen physical environments using only a single observed trajectory, while maintaining high forecasting accuracy and inferring underlying physical parameters with minimal supervision.

The proposed framework addresses multi-environment adaptation through a combination of shared foundational parameters and low-dimensional context conditioning. Specifically, a linear hypernetwork maps low-dimensional context vectors—which represent environment-specific traits—to shift the parameters of the base dynamics model. The training objective incorporates a locality constraint that ensures environment-specific solutions stay near the shared base model, simplifying optimization into a well-behaved problem. The authors evaluated the approach across four representative nonlinear benchmark systems spanning ordinary and partial differential equations: predator-prey dynamics (Lotka-Volterra), biochemical glycolysis (Glycolytic-Oscillator), reaction-diffusion pattern formation (Gray-Scott), and fluid flow (Navier-Stokes). Baseline comparisons included standard gradient-based meta-learning algorithms and multi-task learning models.

The evaluation yielded several key findings. First, CoDA achieved state-of-the-art forecasting accuracy across all tested systems, outperforming gradient-based meta-learning baselines by orders of magnitude in one-shot adaptation settings where only a single trajectory is observed. Second, unlike conventional meta-learning methods that suffered severe performance drops when transitioning from training environments to unseen environments, CoDA maintained virtually identical error rates across both in-domain and out-of-domain tests. Third, the sparsity-inducing formulation (CoDA-L1) delivered the strongest generalization by automatically selecting the most relevant parameter subspaces for adaptation. Fourth, the method exhibited high sample efficiency, achieving top-tier adaptation accuracy with just one sample trajectory where competing methods required substantially more data to reduce error. Finally, the learned context vectors established a direct correspondence with actual physical parameters, allowing the model to infer true system parameters (such as viscosity and reaction rates) with very low error rates, even outside the immediate training envelope.

These findings indicate that data-driven physical emulators do not require expensive end-to-end retraining or massive datasets to adapt to novel operating conditions. By constraining the adaptation search space to a low-dimensional context, CoDA mitigates meta-overfitting and significantly reduces computational cost and latency. This capability has strong practical implications for digital twins, real-time control, and simulation workflows where environmental parameters fluctuate dynamically and fast adaptation is critical for operational performance and safety.

Organizations developing or deploying neural network-based physical simulations should consider integrating low-rank, context-informed adaptation architectures into their pipelines rather than relying on standard invariant models or unconstrained meta-learning. For practical deployment, practitioners should cross-validate context vector dimensionality against expected physical degrees of freedom and apply sparsity constraints to optimize performance. Before deploying the approach in safety-critical operations, pilot studies should assess performance on partially observed states, noisy sensor data, and higher-dimensional physical domains.

Confidence in these findings is supported by theoretical proofs for linearly parametrized systems and consistent experimental results across diverse physical equations. However, users should note key limitations: generalization degrades when new environments deviate substantially from the neighborhood of the training distributions, and full state observability was assumed across the experimental evaluations. Further validation is warranted for nonlinearly parametrized systems under extreme distributional shifts.

arXiv: 2202.01889yuan-yin/CoDA

No sufficiently relevant recommendations were found.

Cover for Generalizing to New Physical Systems via Context-Informed Dynamics Model

Abstract

Data-driven approaches to modeling physical systems fail to generalize to unseen systems that share the same general dynamics with the learning domain, but correspond to different physical contexts. We propose a new framework for this key problem, context-informed dynamics adaptation (CoDA), which takes into account the distributional shift across systems for fast and efficient adaptation to new dynamics. CoDA leverages multiple environments, each associated to a different dynamic, and learns to condition the dynamics model on contextual parameters, specific to each environment. The conditioning is performed via a hypernetwork, learned jointly with a context vector from observed data. The proposed formulation constrains the search hypothesis space for fast adaptation and better generalization across environments with few samples. We theoretically motivate our approach and show state-of-the-art generalization results on a set of nonlinear dynamics, representative of a variety of application domains. We also show, on these systems, that new system parameters can be inferred from context vectors with minimal supervision.

Table of Contents

  • 1. Introduction
  • 2. Generalization for Dynamical Systems
  • 2.1. Problem Setting
  • 2.2. Multi-Environment Learning Problem
  • 3. The CoDA Learning Framework
  • 3.1. Adaptation Rule
  • 3.2. Constrained Optimization Problem
  • 3.3. Context-Informed Hypernetwork
  • 3.4. Validity for Dynamical Systems
  • 3.5. Benefits of CoDA
  • 4. Framework Implementation
  • 5. Experiments
  • 5.1. Dynamical Systems
  • 5.2. Experimental Setting
  • 5.3. Implementation of CoDA
  • 5.4. Baselines
  • 5.5. Generalization Results
  • 5.6. Ablation Studies
  • 5.7. Sample Efficiency
  • 5.8. Parameter Estimation
  • 5.8.1. Empirical Observations
  • 5.8.2. Theoretical Motivation
  • 6. Related Work
  • 7. Conclusion
  • References
  • A. Discussion
  • A.1. Adaptation Rule
  • A.2. Decoding for Context-Informed Adaptation
  • A.2.1. Conditioning via Concatenation
  • A.2.2. Conditioning via Feature Modulation
  • B. Proofs
  • C. System Parameter Estimation
  • D. Low-Rank Assumption
  • E. Locality Constraint
  • F. Experimental Settings
  • F.1. Dynamical Systems
  • F.2. Implementation and Hyperparameters
  • G. Trajectory Prediction Visualization

Knowls

  1. Knowl 1 — Generalization across physical environments

    definition

    The paper formulates dynamics generalization as learning from multiple environments, where each environment ee has a vector field fef^e and trajectories satisfying dxdt=fe(x)\frac{dx}{dt}=f^e(x). The environments share the parameterized form of the underlying differential equation but can differ in physical parameters, external forcing, or initial conditions. Training supplies trajectories from a set of environments EtrE_{\mathrm{tr}}; at test time, an unseen environment in EadE_{\mathrm{ad}} supplies a small amount of trajectory data, and the model must adapt to predict its dynamics. Thus the target is not one environment-invariant vector field, but a family of related dynamics that can be adapted to from observations.

  2. Knowl 2 — CoDA learns low-dimensional context-conditioned dynamics

    model/method

    Context-Informed Dynamics Adaptation (CoDA) represents the dynamics model for environment ee by adapting shared parameters θc∈Rdθ\theta^c\in\mathbb{R}^{d_\theta} with a context vector ξe∈Rdξ\xi^e\in\mathbb{R}^{d_\xi}. A shared linear hypernetwork with weights W∈Rdθ×dξW\in\mathbb{R}^{d_\theta\times d_\xi} produces the environment-specific parameter update:

    θe=θc+Wξe.\theta^e=\theta^c+W\xi^e.

    Here, gθeg_{\theta^e} maps system states to their time derivatives, and the adaptation subspace is the column span of WW, whose dimension is at most dξd_\xi. CoDA trains θc\theta^c, WW, and training-environment contexts by minimizing trajectory-fit loss plus a locality penalty, ∑e∈Etr[L(θc+Wξe,De)+λ∥Wξe∥22]\sum_{e\in E_{\mathrm{tr}}}\big[\mathcal{L}(\theta^c+W\xi^e,D^e)+\lambda\|W\xi^e\|_2^2\big], where DeD^e is the observed trajectory data for environment ee and λ\lambda is a regularization weight. At adaptation time, θc\theta^c and WW are fixed and only the new environment's context is optimized with the same objective. Since dξd_\xi can be much smaller than dθd_\theta, adaptation estimates a small context rather than a full set of model parameters.

  3. Knowl 3 — Locality regularization and sparse adaptation

    empirical result

    CoDA uses locality regularization to keep environment-specific parameter changes near the shared parameters. Instead of directly penalizing ∥Wξe∥\|W\xi^e\|, its implementation uses R(W,ξe)=λξ∥ξe∥22+λΩΩ(W)R(W,\xi^e)=\lambda_\xi\|\xi^e\|_2^2+\lambda_\Omega\Omega(W), with weights λξ,λΩ\lambda_\xi,\lambda_\Omega. CoDA-ℓ2\ell_2 uses an ℓ2\ell_2 parameter-change constraint and Ω(W)=∥W∥22\Omega(W)=\|W\|_2^2. CoDA-ℓ1\ell_1 uses an ℓ1\ell_1 parameter-change constraint and Ω(W)=∑r=1dθ∥Wr,:∥2\Omega(W)=\sum_{r=1}^{d_\theta}\|W_{r,:}\|_2, which encourages row sparsity and selects a restricted set of model parameters for adaptation.

    The in-domain ablation below reports test MSE; LV values are multiplied by 10−510^{-5} and GO values by 10−410^{-4}. It shows that locality improves the full-model results on both systems, while restricting adaptation to a particular layer can help or hurt depending on the system.

    LV (×10−5\times 10^{-5}) GO (×10−4\times 10^{-4})
    Adapted parameters Without ℓ2\ell_2 With ℓ2\ell_2 Without ℓ2\ell_2 With ℓ2\ell_2
    Full model 2.28±0.292.28\pm0.29 1.52±0.081.52\pm0.08 2.98±0.712.98\pm0.71 2.45±0.382.45\pm0.38
    First layer 2.25±0.292.25\pm0.29 2.41±0.232.41\pm0.23 2.38±0.712.38\pm0.71 2.12±0.552.12\pm0.55
    Last layer 1.86±0.241.86\pm0.24 1.27±0.031.27\pm0.03 28.4±0.6028.4\pm0.60 28.4±0.6428.4\pm0.64
  4. Knowl 4 — Trajectory-based training from state observations

    equation

    Because trajectory datasets do not directly provide the vector-field values needed for a derivative-matching loss, CoDA trains by comparing observed trajectories with model rollouts. For environment ee, let xe,i(t,sj)x^{e,i}(t,s_j) be the observed state of trajectory ii at time tt and spatial grid point sjs_j, and let x0e,ix_0^{e,i} be its initial condition. The predicted trajectory is obtained by integrating the learned vector field:

    x~e,i(tk)=x0e,i+∫0tkgθ(x~e,i(τ)) dτ.\tilde{x}^{e,i}(t_k)=x_0^{e,i}+\int_0^{t_k}g_\theta\big(\tilde{x}^{e,i}(\tau)\big)\,d\tau.

    The training loss sums squared errors over trajectories, observation times, and spatial grid points:

    L(θ,De)=∑i=1N∑k=1K∑j=1M∥xe,i(tk,sj)−x~e,i(tk,sj)∥22.\mathcal{L}(\theta,D^e)=\sum_{i=1}^{N}\sum_{k=1}^{K}\sum_{j=1}^{M}\left\|x^{e,i}(t_k,s_j)-\tilde{x}^{e,i}(t_k,s_j)\right\|_2^2.

    Here NN is the number of trajectories, KK the number of sampled times, and MM the number of spatial grid points; for an ODE there is no spatial grid and the spatial sum is omitted. The integral is evaluated with a numerical ODE solver, allowing the same continuous-time formulation to be trained on discretely sampled trajectories.

  5. Knowl 5 — Environment gradients occupy a low-dimensional subspace for linear systems

    theoretical result

    Suppose the family of vector fields is linearly parameterized by dpd_p varying physical parameters, with dp≪dθd_p\ll d_\theta, where dθd_\theta is the number of dynamics-model parameters. For any shared model parameter vector θc\theta^c, the span of the environment-specific gradients {∇θL(θc,De):e∈E}\{\nabla_\theta\mathcal{L}(\theta^c,D^e):e\in E\} has dimension at most dpd_p. This result motivates restricting CoDA's updates to a low-dimensional learned subspace. It is a theoretical guarantee for linearly parameterized dynamics; for nonlinear parameterizations, the paper reports empirical evidence of low-dimensional gradient structure rather than the same guarantee.

  6. Knowl 6 — A closed-form context solution under quadratic adaptation loss

    theoretical result

    With shared parameters θc\theta^c and decoder WW fixed, consider adaptation of one environment's context ξe\xi^e under an ℓ2\ell_2 locality penalty of weight λ\lambda. In the quadratic-loss setting considered by the paper, define He=W⊤∇θ2L(θc,De)WH^e=W^\top\nabla_\theta^2\mathcal{L}(\theta^c,D^e)W and λ′=2λ\lambda'=2\lambda. When He+λ′W⊤WH^e+\lambda'W^\top W is invertible, the objective is convex and has a unique minimizer:

    ξe∗=−(He+λ′W⊤W)−1W⊤∇θL(θc,De).\xi^{e*}=-\left(H^e+\lambda'W^\top W\right)^{-1}W^\top\nabla_\theta\mathcal{L}(\theta^c,D^e).

    Here DeD^e is the observed data for the adapted environment, and the gradient and Hessian are with respect to the full model parameters θ\theta. The paper states that, if either HeH^e or λ′W⊤W\lambda'W^\top W is invertible, the matrix in this solution is invertible for all but finitely many values of λ′\lambda'. The result supplies a unique context estimate under its stated quadratic and invertibility conditions.

  7. Knowl 7 — Benchmark systems and evaluation protocol

    experimental setup

    CoDA is evaluated on four nonlinear systems: Lotka–Volterra predator–prey dynamics (LV), a glycolytic oscillator (GO), the Gray–Scott reaction–diffusion PDE (GS), and the two-dimensional incompressible Navier–Stokes system (NS). The number of training/adaptation environments is respectively 9/4 for LV, 9/4 for GO, 4/4 for GS, and 5/4 for NS. The environment-varying parameter dimensions are dp=2d_p=2 for LV, GO, and GS, and dp=1d_p=1 for NS. Training trajectories per environment are 4, 32, 1, and 16, respectively. Adaptation uses one trajectory per new environment; evaluation uses 32 test trajectories per environment.

    The task is forecasting: the initial condition is used to predict the trajectory, while the adaptation trajectory is used to infer the new environment's context. The dynamics approximators are four-layer width-64 MLPs for LV and GO, a four-layer 64-channel ConvNet for GS, and a Fourier Neural Operator with four spectral convolution layers, 12 frequency modes, and width 10 for NS. All models use the same dynamics architecture within each system and a trajectory-based loss. The reported benchmark compares CoDA-ℓ1\ell_1 and CoDA-ℓ2\ell_2 with MAML, ANIL, Meta-SGD, LEADS, and CAVIA with concatenation or FiLM conditioning. Results are means and standard deviations across four seeds.

  8. Knowl 8 — CoDA improves one-shot generalization across four systems

    data/table

    The table reports test MSE in training environments (in-domain) and in new environments after one-trajectory adaptation. Entries are mean ±\pm standard deviation across four seeds; each system has its own MSE scale, shown in the system label. Both CoDA variants attain the lowest reported MSE in every system and evaluation setting. The comparison indicates that CoDA retains similar performance between in-domain evaluation and one-shot adaptation, whereas several baselines have much larger adaptation errors. The accompanying LV sweep evaluates 2,601 adaptation environments on a 51×5151\times51 grid over (β,δ)∈[0.25,1.25]2(\beta,\delta)\in[0.25,1.25]^2: MAPE is very low inside the convex hull of training environments, remains low nearby, and rises farther away. In an LV sample-efficiency comparison, CoDA-ℓ1\ell_1 remains nearly flat as the number of adaptation trajectories increases from one to ten, while MAML improves substantially and LEADS more moderately.

    System (MSE scale) Method In-domain Adaptation
    LV (×10−5\times10^{-5}) MAML 60.3±1.360.3\pm1.3 3150±9403150\pm940
    ANIL 381±76381\pm76 4570±23904570\pm2390
    Meta-SGD 32.7±12.632.7\pm12.6 7220±45807220\pm4580
    LEADS 3.70±0.273.70\pm0.27 47.61±12.4747.61\pm12.47
    CAVIA-FILM 4.38±1.154.38\pm1.15 8.41±3.208.41\pm3.20
    CAVIA-CONCAT 2.43±0.662.43\pm0.66 6.26±0.776.26\pm0.77
    CoDA-ℓ2\ell_2 1.52±0.081.52\pm0.08 1.82±0.241.82\pm0.24
    CoDA-ℓ1\ell_1 1.35±0.221.35\pm0.22 1.24±0.201.24\pm0.20
    GO (×10−4\times10^{-4}) MAML 57.3±2.157.3\pm2.1 1081±621081\pm62
    ANIL 74.5±11.574.5\pm11.5 1688±2261688\pm226
    Meta-SGD 42.3±6.942.3\pm6.9 1573±4131573\pm413
    LEADS 31.4±3.331.4\pm3.3 113.8±41.5113.8\pm41.5
    CAVIA-FILM 4.44±1.464.44\pm1.46 3.87±1.283.87\pm1.28
    CAVIA-CONCAT 5.09±0.355.09\pm0.35 2.37±0.232.37\pm0.23
    CoDA-ℓ2\ell_2 2.45±0.382.45\pm0.38 1.98±0.061.98\pm0.06
    CoDA-ℓ1\ell_1 2.20±0.262.20\pm0.26 1.86±0.291.86\pm0.29
    GS (×10−3\times10^{-3}) MAML 3.67±0.533.67\pm0.53 2.25±0.392.25\pm0.39
    ANIL 5.01±0.805.01\pm0.80 3.95±0.113.95\pm0.11
    Meta-SGD 2.85±0.542.85\pm0.54 2.68±0.202.68\pm0.20
    LEADS 2.90±0.762.90\pm0.76 1.36±0.431.36\pm0.43
    CAVIA-FILM 2.81±1.152.81\pm1.15 1.43±1.071.43\pm1.07
    CAVIA-CONCAT 2.67±0.482.67\pm0.48 1.62±0.851.62\pm0.85
    CoDA-ℓ2\ell_2 1.01±0.151.01\pm0.15 0.77±0.100.77\pm0.10
    CoDA-ℓ1\ell_1 0.90±0.0570.90\pm0.057 0.74±0.100.74\pm0.10
    NS (×10−4\times10^{-4}) MAML 68.0±8.068.0\pm8.0 51.1±4.051.1\pm4.0
    ANIL 61.7±4.361.7\pm4.3 48.6±3.248.6\pm3.2
    Meta-SGD 53.9±28.153.9\pm28.1 44.3±27.144.3\pm27.1
    LEADS 14.0±1.5514.0\pm1.55 28.6±7.2328.6\pm7.23
    CAVIA-FILM 23.2±12.123.2\pm12.1 22.6±9.8822.6\pm9.88
    CAVIA-CONCAT 25.5±6.3125.5\pm6.31 26.0±8.2426.0\pm8.24
    CoDA-ℓ2\ell_2 9.40±1.139.40\pm1.13 10.3±1.4810.3\pm1.48
    CoDA-ℓ1\ell_1 8.35±1.718.35\pm1.71 9.65±1.379.65\pm1.37

    For LV, test MSE as the adaptation-trajectory count NadN_{\mathrm{ad}} varies is:

    Method Nad=1N_{\mathrm{ad}}=1 Nad=5N_{\mathrm{ad}}=5 Nad=10N_{\mathrm{ad}}=10
    MAML 3150±9403150\pm940 239±16239\pm16 173±10173\pm10
    LEADS 47.61±12.4747.61\pm12.47 19.89±7.2319.89\pm7.23 19.42±3.5219.42\pm3.52
    CoDA-ℓ1\ell_1 1.24±0.201.24\pm0.20 1.21±0.181.21\pm0.18 1.20±0.171.20\pm0.17

    The LV sample-efficiency values are test MSE multiplied by 10−510^{-5}.

  9. Knowl 9 — Contexts can encode and identify physical system parameters

    theoretical result

    CoDA's learned context can be used to estimate environment parameters when the context-to-dynamics map preserves the physical parameterization. The paper's exact identification result assumes: (1) dynamics are linear in the inputs and system parameters; (2) the dynamics model and hypernetwork are linear; (3) each system has a unique parameter vector; (4) context dimension equals the number of varying parameters, dξ=dpd_\xi=d_p; and (5) the physical parameters are known for dynamics forming a basis of the family. If the model and hypernetwork represent that basis correctly, the context-to-dynamics mapping identifies the physical parameters of new environments. The paper also gives a local extension for systems and dynamics models nonlinear in the inputs: with a linear hypernetwork, exact identification is guaranteed for sufficiently small context norms when the basis systems are represented under the stated parameter-rescaling condition.

    Empirically, the authors find a linear correspondence between learned contexts and true parameters for LV, and use learned contexts to estimate parameters in LV, GS, and NS. The table reports parameter-estimation MAPE as mean ±\pm standard deviation, with the number of evaluated environments shown for the in-hull and out-of-hull groups. The relatively low errors in both groups show that estimation worked beyond the convex hull of training environments in these experiments.

    System In convex hull: MAPE (%), environments Out of convex hull: MAPE (%), environments Overall MAPE (%)
    LV 0.15±0.110.15\pm0.11, 625 0.73±1.330.73\pm1.33, 1976 0.59±1.330.59\pm1.33
    GS 0.37±0.250.37\pm0.25, 625 0.74±0.670.74\pm0.67, 1976 0.65±0.620.65\pm0.62
    NS 0.10±0.080.10\pm0.08, 40 0.51±0.350.51\pm0.35, 41 0.30±0.330.30\pm0.33

Coverage note — The supplementary material's full governing equations and initial-condition recipes for the four established benchmark systems, plus detailed context-dimension and gradient-spectrum diagnostics, are omitted because they provide testbed specifications or supporting analyses rather than additional core methods or results.

References

  1. 1.Antoniou, A., Edwards, H., and Storkey, A. J. How to train your MAML. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=HJGven05Y7. (p. 5)
  2. 2.Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. Invariant risk minimization. CoRR, abs/1907.02893, 2019. URL http://arxiv.org/abs/1907.02893. (p. 9)
  3. 3.Bertinetto, L., Henriques, J. F., Torr, P., and Vedaldi, A. Meta-learning with differentiable closed-form solvers. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=HyxnZh0ct7. (p. 14)
  4. 4.Caruana, R. Multitask learning. Machine Learning, 28(1):41–75, 1997. (pp. 9 and 14)
  5. 5.Chen, R. T. Q. torchdiffeq, 2021. URL https://github.com/rtqichen/torchdiffeq. (p. 5)
  6. 6.Chen, Y., Friesen, A. L., Behbahani, F., Doucet, A., Budden, D., Hoffman, M., and de Freitas, N. Modular meta-learning with shrinkage. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M. F., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 2858–2869. Curran Associates, Inc., 2020a. (p. 7)
  7. 7.Chen, Z., Zhang, J., Arjovsky, M., and Bottou, L. Symplectic recurrent neural networks. In International Conference on Learning Representations, 2020b. URL https://openreview.net/forum?id=BkgYPREtPr. (p. 1)
  8. 8.Clavera, I., Nagabandi, A., Liu, S., Fearing, R. S., Abbeel, P., Levine, S., and Finn, C. Learning to adapt in dynamic, real-world environments through meta-reinforcement learning. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=HyztsoC5Y7. (p. 2)
  9. 9.Courtier, P., Thepaut, J.-N., and Hollingsworth, A. A strategy for operational implementation of 4D-Var, using an incremental approach. Quarterly Journal of the Royal Meteorological Society, 120(519):1367–1387, 1994. (p. 1)
  10. 10.Daniels, B. C. and Nemenman, I. Efficient inference of parsimonious phenomenological models of cellular dynamics using s-systems and alternating regression. PLOS ONE, 10(3):1–14, 03 2015. (pp. 6, 17, and 18)
  11. 11.de Bezenac, E., Pajot, A., and Gallinari, P. Deep learning for physical processes: Incorporating prior scientific knowledge. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=By4HsfWAZ. (p. 1)
  12. 12.Duraisamy, K., Iaccarino, G., and Xiao, H. Turbulence modeling in the age of data. Annual Review of Fluid Mechanics, 51:357–377, 2019. (p. 1)
  13. 13.Finn, C., Abbeel, P., and Levine, S. Model-agnostic meta-learning for fast adaptation of deep networks. In Precup, D. and Teh, Y. W. (eds.), Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pp. 1126–1135. PMLR, 06–11 Aug 2017. URL http://proceedings.mlr.press/v70/finn17a.html. (pp. 2, 6, 9, and 14)
  14. 14.Flennerhag, S., Rusu, A. A., Pascanu, R., Visin, F., Yin, H., and Hadsell, R. Meta-learning with warped gradient descent. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rkeiQlBFPB. (p. 9)
  15. 15.Frankle, J. and Carbin, M. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=rJl-b3RcF7. (p. 5)
  16. 16.Fresca, S., Manzoni, A., Dede, L., and Quarteroni, A. Deep learning-based reduced order models in cardiac electrophysiology. PloS one, 15(10):e0239416–e0239416, 10 2020. doi: 10.1371/journal.pone.0239416. URL https://pubmed.ncbi.nlm.nih.gov/33002014. (p. 1)
  17. 17.Garnelo, M., Rosenbaum, D., Maddison, C., Ramalho, T., Saxton, D., Shanahan, M., Teh, Y. W., Rezende, D. J., and Eslami, S. M. A. Conditional neural processes. In Dy, J. G. and Krause, A. (eds.), Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmassan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pp. 1690–1699. PMLR, 2018. URL http://proceedings.mlr.press/v80/garnelo18a.html. (pp. 5 and 9)
  18. 18.Goyal, A., Lamb, A., Zhang, Y., Zhang, S., Courville, A., and Bengio, Y. Professor forcing: A new algorithm for training recurrent networks. In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS’16, pp. 4608–4616, Red Hook, NY, USA, 2016. Curran Associates Inc. ISBN 9781510838819. (p. 5)
  19. 19.Greydanus, S., Dzamba, M., and Yosinski, J. Hamiltonian neural networks. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alche-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 15353–15363, 2019. URL https://proceedings.neurips.cc/paper/2019/file/26cd8ecadce0d4efd6cc8a8725cbd1f8-Paper.pdf. (p. 1)
  20. 20.Gur-Ari, G., Roberts, D. A., and Dyer, E. Gradient descent happens in a tiny subspace, 2019. URL https://openreview.net/forum?id=ByeTHsAqtX. (p. 5)
  21. 21.Ha, D., Dai, A. M., and Le, Q. V. Hypernetworks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. URL https://openreview.net/forum?id=rkpACe1lx. (p. 4)
  22. 22.Hairer, E., Nørsett, S. P., and Wanner, G. Solving Ordinary Differential Equations I: Nonstiff problems. Springer, Berlin, second edition, 2000. (p. 5)
  23. 23.Kalman, R. E. A New Approach to Linear Filtering and Prediction Problems. Journal of Basic Engineering, 82 (1):35–45, 03 1960. ISSN 0021-9223. doi: 10.1115/1.3662552. URL https://doi.org/10.1115/1.3662552. (p. 1)
  24. 24.Kim, H., Mnih, A., Schwarz, J., Garnelo, M., Eslami, A., Rosenbaum, D., Vinyals, O., and Teh, Y. W. Attentive neural processes. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=SkE6PjC9KX. (p. 5)
  25. 25.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In Bengio, Y. and LeCun, Y. (eds.), 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015. URL http://arxiv.org/abs/1412.6980. (pp. 6 and 18)
  26. 26.Kochkov, D., Smith, J. A., Alieva, A., Wang, Q., Brenner, M. P., and Hoyer, S. Machine learning accelerated computational fluid dynamics. Proceedings of the National Academy of Sciences, 118, 2021. URL https://www.pnas.org/content/118/21/e2101784118. (p. 1)
  27. 27.Krueger, D., Caballero, E., Jacobsen, J.-H., Zhang, A., Binas, J., Zhang, D., Priol, R. L., and Courville, A. Out-of-distribution generalization via risk extrapolation (REx), 2021. URL https://arxiv.org/pdf/2003.00688.pdf. (p. 9)
  28. 28.Lee, K., Maji, S., Ravichandran, A., and Soatto, S. Meta-learning with differentiable convex optimization. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pp. 10657–10665. Computer Vision Foundation / IEEE, 2019. doi: 10.1109/CVPR.2019.01091. URL http://openaccess.thecvf.com/content_CVPR_2019/html/Lee_Meta-Learning_With_Differentiable_Convex_Optimization_CVPR_2019_paper.html. (p. 14)
  29. 29.Lee, K., Seo, Y., Lee, S., Lee, H., and Shin, J. Context-aware dynamics model for generalization in model-based reinforcement learning. In III, H. D. and Singh, A. (eds.), Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pp. 5757–5766. PMLR, 13–18 Jul 2020. URL http://proceedings.mlr.press/v119/lee20g.html. (p. 2)
  30. 30.Lee, Y. and Choi, S. Gradient-based meta-learning with learned layerwise metric and subspace. In International Conference on Machine Learning, pp. 2933–2942, 2018. (p. 9)
  31. 31.Li, C., Farkhoor, H., Liu, R., and Yosinski, J. Measuring the intrinsic dimension of objective landscapes. In International Conference on Learning Representations, 2018a. URL https://openreview.net/forum?id=ryup8-WCW. (p. 5)
  32. 32.Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T. Visualizing the loss landscape of neural nets. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018b. (pp. 4 and 5)
  33. 33.Li, Z., Zhou, F., Chen, F., and Li, H. Meta-SGD: Learning to learn quickly for few shot learning. CoRR, abs/1707.09835, 2017. URL http://arxiv.org/abs/1707.09835. (pp. 7 and 9)
  34. 34.Li, Z., Kovachki, N. B., Azizzadenesheli, K., liu, B., Bhattacharya, K., Stuart, A., and Anandkumar, A. Fourier neural operator for parametric partial differential equations. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=c8P9NQVtmnO. (pp. 1, 6, and 18)
  35. 35.Long, Z., Lu, Y., Ma, X., and Dong, B. PDE-Net: Learning PDEs from data. In Dy, J. G. and Krause, A. (eds.), Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmassan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pp. 3214–3222. PMLR, 2018. URL http://proceedings.mlr.press/v80/long18a.html. (p. 5)
  36. 36.Lotka, A. Elements of physical biology. Nature, 116 (2917):461–461, 1925. (pp. 6 and 17)
  37. 37.Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P. A simple neural attentive meta-learner. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=B1DmUzWAW. (pp. 2 and 8)
  38. 38.Neic, A., Campos, F. O., Prassl, A. J., Niederer, S. A., Bishop, M. J., Vigmond, E. J., and Plank, G. Efficient computation of electrograms and ecgs in human whole heart simulations using a reaction-eikonal model. J. Comput. Phys., 346:191–211, 2017. doi: 10.1016/j.jcp.2017.06.020. URL https://doi.org/10.1016/j.jcp.2017.06.020. (p. 1)
  39. 39.Park, J. J., Florence, P., Straub, J., Newcombe, R. A., and Lovegrove, S. DeepSDF: Learning continuous signed distance functions for shape representation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pp. 165–174. Computer Vision Foundation / IEEE, 2019. doi: 10.1109/CVPR.2019.00025. URL http://openaccess.thecvf.com/content_CVPR_2019/html/Park_DeepSDF_Learning_Continuous_Signed_Distance_Functions_for_Shape_Representation_CVPR_2019_paper.html. (p. 9)
  40. 40.Pearson, J. E. Complex patterns in a simple system. Science, 261(5118):189–192, 1993. doi: 10.1126/science.261.5118.189. URL https://www.science.org/doi/abs/10.1126/science.261.5118.189. (pp. 6 and 18)
  41. 41.Perez, E., Strub, F., de Vries, H., Dumoulin, V., and Courville, A. C. FiLM: Visual reasoning with a general conditioning layer. In McIlraith, S. A. and Weinberger, K. Q. (eds.), Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2-7, 2018, pp. 3942–3951. AAAI Press, 2018. URL https://www.aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/16528. (pp. 7 and 14)
  42. 42.Raghu, A., Raghu, M., Bengio, S., and Vinyals, O. Rapid learning or feature reuse? towards understanding the effectiveness of maml. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rkgMkCEtPB. (pp. 7, 9, and 14)
  43. 43.Raissi, M., Perdikaris, P., and Karniadakis, G. E. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378:686–707, 2019. ISSN 0021-9991. doi: https://doi.org/10.1016/j.jcp.2018.10.045. URL https://www.sciencedirect.com/science/article/pii/S0021999118307125. (p. 5)
  44. 44.Ramachandran, P., Zoph, B., and Le, Q. V. Searching for activation functions, 2018. URL https://openreview.net/forum?id=SkBYYyZRZ. (p. 18)
  45. 45.Rebuffi, S., Bilen, H., and Vedaldi, A. Efficient parametrization of multi-domain deep neural networks. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pp. 8119–8127. Computer Vision Foundation / IEEE Computer Society, 2018. doi: 10.1109/CVPR.2018.00847. URL http://openaccess.thecvf.com/content_cvpr_2018/html/Rebuffi_Efficient_Parametrization_of_CVPR_2018_paper.html. (p. 9)
  46. 46.Rebuffi, S.-A., Bilen, H., and Vedaldi, A. Learning multiple visual domains with residual adapters. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. (p. 9)
  47. 47.Reichstein, M., Camps-Valls, G., Stevens, B., Jung, M., Denzler, J., Carvalhais, N., and Prabhat. Deep learning and process understanding for data-driven earth system science. Nature, 566(7743):195–204, 2019. (p. 1)
  48. 48.Requeima, J., Gordon, J., Bronskill, J., Nowozin, S., and Turner, R. E. Fast and flexible multi-task classification using conditional neural adaptive processes. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alche-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 7957–7968. 2019. URL https://proceedings.neurips.cc/paper/2019/file/1138d90ef0a0848a542e57d1595f58ea-Paper.pdf. (p. 9)
  49. 49.Ruder, S. An overview of multi-task learning in deep neural networks. CoRR, abs/1706.05098, 2017. URL http://arxiv.org/abs/1706.05098. (p. 14)
  50. 50.Rusu, A. A., Rao, D., Sygnowski, J., Vinyals, O., Pascanu, R., Osindero, S., and Hadsell, R. Meta-learning with latent embedding optimization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=BJgklhAcK7. (p. 7)
  51. 51.Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P. Distributionally robust neural networks. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=ryxGuJrFvS. (p. 9)
  52. 52.Shaier, S., Raissi, M., and Seshaiyer, P. Data-driven approaches for predicting spread of infectious diseases through DINNs: Disease informed neural networks. 2021. (p. 1)
  53. 53.Sirignano, J. and Spiliopoulos, K. DGM: A deep learning algorithm for solving partial differential equations. Journal of Computational Physics, 375:1339–1364, 2018. ISSN 10902716. doi: 10.1016/j.jcp.2018.08.029. (p. 1)
  54. 54.Stokes, G. G. On the Effect of the Internal Friction of Fluids on the Motion of Pendulums. Transactions of the Cambridge Philosophical Society, 9:8, January 1851. (pp. 6 and 18)
  55. 55.Thrun, S. and Pratt, L. Y. Learning to learn: Introduction and overview. In Thrun, S. and Pratt, L. Y. (eds.), Learning to Learn, pp. 3–17. Springer, 1998. ISBN 978-1-4613-7527-2. doi: 10.1007/978-1-4615-5529-2 1. URL https://doi.org/10.1007/978-1-4615-5529-2_1. (p. 2)
  56. 56.Vogels, T., Karimireddy, S. P., and Jaggi, M. Powersgd: Practical low-rank gradient compression for distributed optimization. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alche-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/d9fbed9da256e344c1fa46bb46c34c5f-Paper.pdf. (p. 5)
  57. 57.Wandel, N., Weinmann, M., and Klein, R. Learning incompressible fluid dynamics from scratch - towards fast, differentiable fluid models that generalize. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=KUDUoRsEphu. (p. 1)
  58. 58.Wang, H., Zhao, H., and Li, B. Bridging multi-task learning and meta-learning: Towards efficient training and effective adaptation. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pp. 10991–11002. PMLR, 2021a. URL http://proceedings.mlr.press/v139/wang21ad.html. (p. 9)
  59. 59.Wang, J., Lan, C., Liu, C., Ouyang, Y., and Qin, T. Generalizing to unseen domains: A survey on domain generalization. In Zhou, Z. (ed.), Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI 2021, Virtual Event / Montreal, Canada, 19-27 August 2021, pp. 4627–4635. ijcai.org, 2021b. doi: 10.24963/ijcai.2021/628. URL https://doi.org/10.24963/ijcai.2021/628. (p. 1)
  60. 60.Wang, R., Walters, R., and Yu, R. Meta-learning dynamics forecasting using task inference. CoRR, abs/2102.10271, 2021c. URL https://arxiv.org/abs/2102.10271. (pp. 2 and 9)
  61. 61.Yin, Y., Ayed, I., de Bezenac, E., Baskiotis, N., and Gallinari, P. LEADS: Learning dynamical systems that generalize across environments. In Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, 2021a. URL https://openreview.net/forum?id=HD6CxZtbmIx. (pp. 2, 7, 9, 14, and 18)
  62. 62.Yin, Y., Le Guen, V., Dona, J., de Bezenac, E., Ayed, I., Thome, N., and Gallinari, P. Augmenting physical models with deep networks for complex dynamics forecasting. Journal of Statistical Mechanics: Theory and Experiment, 2021(12):124012, dec 2021b. doi: 10.1088/1742-5468/ac3ae5. URL https://doi.org/10.1088/1742-5468/ac3ae5. (pp. 1 and 5)
  63. 63.Zintgraf, L., Shiarli, K., Kurin, V., Hofmann, K., and Whiteson, S. Fast context adaptation via meta-learning. In Chaudhuri, K. and Salakhutdinov, R. (eds.), Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pp. 7693–7702. PMLR, 09–15 Jun 2019. URL http://proceedings.mlr.press/v97/zintgraf19a.html. (pp. 5, 7, 9, and 14)

Citation

MLA
Kirchmeyer, M., et al. “Generalizing to New Physical Systems via Context-Informed Dynamics Model”. International Conference on Machine Learning, vol. 162, 2022, pp. 11283–301, https://proceedings.mlr.press/v162/kirchmeyer22a.html.
APA
Kirchmeyer, M., Yin, Y., Dona, J., Baskiotis, N., Rakotomamonjy, A., & Gallinari, P. (2022). Generalizing to New Physical Systems via Context-Informed Dynamics Model. International Conference on Machine Learning, 162, 11283–11301. https://proceedings.mlr.press/v162/kirchmeyer22a.html
Chicago
Kirchmeyer, M., Y. Yin, J. Dona, N. Baskiotis, A. Rakotomamonjy, and P. Gallinari. 2022. “Generalizing to New Physical Systems via Context-Informed Dynamics Model”. International Conference on Machine Learning 162: 11283–301. https://proceedings.mlr.press/v162/kirchmeyer22a.html.
Harvard
Kirchmeyer, M. et al. (2022) “Generalizing to New Physical Systems via Context-Informed Dynamics Model”, International Conference on Machine Learning. PMLR, pp. 11283–11301. Available at: https://proceedings.mlr.press/v162/kirchmeyer22a.html.
Vancouver
1. Kirchmeyer M, Yin Y, Dona J, Baskiotis N, Rakotomamonjy A, Gallinari P (2022) Generalizing to New Physical Systems via Context-Informed Dynamics Model. In: International Conference on Machine Learning. PMLR, pp 11283–11301

BibTeX

@InProceedings{pmlr-v162-kirchmeyer22a,
  title = 	 {Generalizing to New Physical Systems via Context-Informed Dynamics Model},
  author =       {Kirchmeyer, Matthieu and Yin, Yuan and Dona, Jeremie and Baskiotis, Nicolas and Rakotomamonjy, Alain and Gallinari, Patrick},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {11283--11301},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/kirchmeyer22a/kirchmeyer22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/kirchmeyer22a.html},
  abstract = 	 {Data-driven approaches to modeling physical systems fail to generalize to unseen systems that share the same general dynamics with the learning domain, but correspond to different physical contexts. We propose a new framework for this key problem, context-informed dynamics adaptation (CoDA), which takes into account the distributional shift across systems for fast and efficient adaptation to new dynamics. CoDA leverages multiple environments, each associated to a different dynamic, and learns to condition the dynamics model on contextual parameters, specific to each environment. The conditioning is performed via a hypernetwork, learned jointly with a context vector from observed data. The proposed formulation constrains the search hypothesis space for fast adaptation and better generalization across environments with few samples. We theoretically motivate our approach and show state-of-the-art generalization results on a set of nonlinear dynamics, representative of a variety of application domains. We also show, on these systems, that new system parameters can be inferred from context vectors with minimal supervision.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/