Adaptive Conformal Predictions for Time Series

Margaux ZaffranOlivier FéronYannig GoudeJulie JosseAymeric Dieuleveut

article2022ICML201 citations

Develops a parameter-free online expert aggregation algorithm that adapts conformal inference to dependent time series data, providing theoretical analysis of its efficiency and validated performance on electricity price forecasting.

Listen

Modern industrial and financial operations increasingly rely on complex machine learning models to forecast dynamic variables such as energy prices and market demand. However, deploying automated forecasts in high-stakes environments requires dependable uncertainty quantification to manage operational and financial risk. Traditional conformal prediction methods provide mathematically sound, distribution-free prediction intervals, but they rely on the assumption of exchangeable data—a condition routinely violated by real-world time series due to temporal dependencies, trends, and regime shifts. The article evaluates how Adaptive Conformal Inference (ACI), a framework designed to update coverage parameters sequentially, can be adapted to provide reliable and informative prediction intervals for dependent time series data.

To address this challenge, the authors mathematically analyzed how ACI's learning rate parameter affects interval width across both independent and autoregressive noise processes. Recognizing that manually tuning this parameter in production is difficult and error-prone, the authors introduced an automated method called Aggregated Adaptive Conformal Inference (AgACI). This parameter-free approach uses online expert aggregation to dynamically weight multiple learning rates based on past performance. The authors validated the framework through extensive synthetic benchmarks across varying degrees of temporal correlation and demonstrated its practical utility on a real-world case study forecasting hourly French electricity spot prices across four years (2016–2019) using random forest regression models.

The analysis produced several key findings. First, while adaptive updates slightly widen intervals when data points are completely independent, they substantially tighten interval widths in strongly autocorrelated time series when paired with an effective learning rate. Second, selecting an inappropriate static learning rate creates severe trade-offs: learning rates that are too small result in under-coverage during persistent forecast errors, while rates that are too large generate uninformative, infinite-width intervals. Third, the proposed AgACI algorithm consistently achieved the target 90% coverage rate (maintaining above 89.8% across all synthetic benchmarks) while achieving the narrowest intervals among valid competing methods. Finally, in the electricity market application, adaptive conformal methods adapted effectively to price volatility and sudden market spikes while maintaining tighter median interval bounds than non-adaptive baselines.

These findings mean that organizations can implement rigorous, distribution-free risk boundaries around their predictive models without assuming idealized statistical conditions. Adopting adaptive conformal methods reduces exposure to unexpected prediction errors in volatile markets, supporting better capital allocation, hedging, and operational planning. Decision-makers should consider adopting online adaptive conformal algorithms like AgACI over static split conformal methods for time series pipelines, as automated aggregation eliminates the need for manual parameter tuning while preserving coverage validity.

While the article establishes strong theoretical and empirical confidence in overall marginal coverage, leaders should exercise caution regarding localized sub-group performance. The electricity market experiments revealed that coverage varied systematically across calendar regimes, achieving approximately 93% coverage on midweek days but dropping to roughly 88% on weekends and Mondays. Organizations should conduct pilot evaluations to verify conditional coverage across operational sub-periods before full-scale deployment, and future analytical efforts should focus on refining conditional calibration for recurring temporal cycles.

arXiv: 2202.07282
  • Paper: A tutorial on conformal prediction, Glenn Shafer et al. (2007). Introduces the foundational framework and sequential error guarantees of conformal prediction that the source directly adapts for non-exchangeable time series.
  • Paper: Distribution-Free Predictive Inference for Regression, Jing Lei et al. (2016). Establishes distribution-free split conformal inference for regression models, which serves as the core static baseline compared and modified in the source.
  • Paper: Introduction to Online Convex Optimization, Elad Hazan (2016). Provides the essential theory of online convex optimization and gradient descent updates that underpin Adaptive Conformal Inference and online expert aggregation.
  • Paper: Quantile Regression Forests, Nicolai Meinshausen (2006). Details quantile regression forests, a foundational non-parametric methodology for conditional interval estimation applied in the source's empirical benchmarks.
Cover for Adaptive Conformal Predictions for Time Series

Abstract

Uncertainty quantification of predictive models is crucial in decision-making problems. Conformal prediction is a general and theoretically sound answer. However, it requires exchangeable data, excluding time series. While recent works tackled this issue, we argue that Adaptive Conformal Inference (ACI, Gibbs & Candès, 2021), developed for distribution-shift time series, is a good procedure for time series with general dependency. We theoretically analyse the impact of the learning rate on its efficiency in the exchangeable and auto-regressive case. We propose a parameter-free method, AgACI, that adaptively builds upon ACI based on online expert aggregation. We lead extensive fair simulations against competing methods that advocate for ACI’s use in time series. We conduct a real case study: electricity price forecasting. The proposed aggregation algorithm provides efficient prediction intervals for day-ahead forecasting. All the code and data to reproduce the experiments are made available on GitHub.

Table of Contents

  • 1. Introduction
  • 2. Setting: ACI for time series
  • 3. Impact of γ on ACI efficiency
  • 3.1. Exchangeable case
  • 3.2. AR(1) case
  • 4. Adaptive strategies based on ACI
  • 5. Numerical evaluation on synthetic data sets
  • 5.1. Data generation process and settings
  • 5.2. Impact of γ, performance of AgACI
  • 5.3. Description of baseline methods
  • 5.4. Experimental results: impact of φ, θ
  • 6. Forecasting French electricity spot prices
  • 6.1. Presentation of the price data
  • 6.2. Price prediction with predictive intervals in 2019
  • 7. Conclusion
  • Acknowledgements
  • References
  • Appendices
  • A. Details on Split Conformal Prediction
  • A.1. Split Conformal Prediction Algorithm
  • A.2. Illustration of the SCP procedure
  • A.3. Theoretical guarantees of Split Conformal Prediction
  • A.4. Examples of dependent scores when data noise is exchangeable
  • A.5. How should we visualise CP predicted intervals?
  • B. Proof of the results presented in Section 3 and additional numerical experiments
  • B.1. Proof of Theorem 3.1
  • B.2. Proof of Lemmas B.1 and B.2
  • B.3. Control of the first four moments: Lemmas B.3 to B.6
  • B.4. Proof of Theorem 3.3
  • B.5. Numerical study of ACI efficiency with AR(1) residuals, with respect to the median length
  • C. Experimental details.
  • C.1. Details on the BOA procedure
  • C.2. Details ARMA(1,1) processes
  • C.3. Random forest parameters
  • C.4. Details about the baselines and comparison
  • D. Additional experiments on synthetic data sets
  • D.1. Additional experimental results of ACI sensitivity to γ, presented in Section 5.2
  • D.2. Comparison to baselines, extension of Section 5.4
  • D.3. Closer look at infinite intervals
  • D.4. Randomised, sequential and other splits.
  • E. Forecasting French electricity spot prices
  • E.1. Details about the data set
  • E.2. Forecasting year 2019

Knowls

  1. Knowl 1 — Adaptive Conformal Inference for Time Series

    model/method

    Adaptive Conformal Inference (ACI) adapts split conformal prediction to time series settings subject to non-exchangeability or distribution shifts. Given historical observations ((x1,y1),…,(xT0,yT0))∈(Rd×R)T0((x_1, y_1), \dots, (x_{T_0}, y_{T_0})) \in (\mathbb{R}^d \times \mathbb{R})^{T_0}, predictions and predictive intervals are constructed sequentially for t∈{T0+1,…,T0+T1}t \in \{T_0+1, \dots, T_0+T_1\}. At step tt, the historical window ((xt−T0,yt−T0),…,(xt−1,yt−1))((x_{t-T_0}, y_{t-T_0}), \dots, (x_{t-1}, y_{t-1})) is partitioned sequentially into a proper training set Trt\mathrm{Tr}_t and a calibration set Calt\mathrm{Cal}_t, where Trt\mathrm{Tr}_t contains earlier indices and Calt\mathrm{Cal}_t contains the most recent indices.

    A regression model μ^t\hat{\mu}_t is fitted on Trt\mathrm{Tr}_t, and conformity scores on the calibration set are computed as absolute residuals SCalt={si=∣yi−μ^t(xi)∣:i∈Calt}S_{\mathrm{Cal}_t} = \{s_i = |y_i - \hat{\mu}_t(x_i)| : i \in \mathrm{Cal}_t\}. A target miscoverage level α∈[0,1]\alpha \in [0, 1] is dynamically adjusted through an effective miscoverage parameter αt\alpha_t, initialized at α1=α\alpha_1 = \alpha. At each time step t≥1t \ge 1:

    C^αt(xt)=[μ^t(xt)±Q^1−αt(SCalt)]\hat{C}_{\alpha_t}(x_t) = \left[ \hat{\mu}_t(x_t) \pm \hat{Q}_{1-\alpha_t}(S_{\mathrm{Cal}_t}) \right] αt+1=αt+γ(α−1{yt∉C^αt(xt)})\alpha_{t+1} = \alpha_t + \gamma \left( \alpha - \mathbf{1}\{y_t \notin \hat{C}_{\alpha_t}(x_t)\} \right)

    where γ≥0\gamma \ge 0 is a learning rate step size, Q^1−αt(SCalt)\hat{Q}_{1-\alpha_t}(S_{\mathrm{Cal}_t}) denotes the (1−αt)(1-\alpha_t)-empirical quantile of the calibration conformity scores, and 1{⋅}\mathbf{1}\{\cdot\} is the indicator function.

    If αt≤0\alpha_t \le 0, by convention Q^1−αt=+∞\hat{Q}_{1-\alpha_t} = +\infty, yielding an infinite prediction interval C^αt(xt)=R\hat{C}_{\alpha_t}(x_t) = \mathbb{R}. If αt≥1\alpha_t \ge 1, by convention Q^1−αt=0\hat{Q}_{1-\alpha_t} = 0, collapsing the interval to the point prediction C^αt(xt)={μ^t(xt)}\hat{C}_{\alpha_t}(x_t) = \{\hat{\mu}_t(x_t)\}. If the interval covers yty_t, αt+1\alpha_{t+1} increases, tightening the next interval; if it fails to cover, αt+1\alpha_{t+1} decreases, widening the next interval.

  2. Knowl 2 — Asymptotic Efficiency Degradation of ACI in Exchangeable Settings

    theoretical result

    In the ideal setting where data errors are exchangeable and the score quantile function QQ is perfectly known, Adaptive Conformal Inference (ACI) strictly degrades interval efficiency (increases expected interval length) compared to standard non-adaptive Split Conformal Prediction (SCP).

    Let α∈Q∩(0,1)\alpha \in \mathbb{Q} \cap (0, 1) be a rational target miscoverage rate. Assume the conformity scores St=∣Yt−μ^(Xt)∣S_t = |Y_t - \hat{\mu}(X_t)| are exchangeable with a bounded and 4-times continuously differentiable quantile function Q∈C4([0,1])Q \in \mathcal{C}^4([0, 1]), and the empirical quantile estimator is exact: Q^1−α^(SCalt)=Q(1−α^)\hat{Q}_{1-\hat{\alpha}}(S_{\mathrm{Cal}_t}) = Q(1 - \hat{\alpha}) for all α^\hat{\alpha}. Let L(αt)=2Q(1−αt)L(\alpha_t) = 2 Q(1 - \alpha_t) denote the length of the ACI prediction interval at step tt, with Q(1−α^)Q(1-\hat{\alpha}) extended by Q(1)Q(1) for α^<0\hat{\alpha} < 0, and let L0=2Q(1−α)L_0 = 2 Q(1 - \alpha) denote the non-adaptive conformal interval length (corresponding to γ=0\gamma = 0).

    For any learning rate γ>0\gamma > 0, the sequence of effective miscoverage levels (αt)t>0(\alpha_t)_{t > 0} forms an irreducible Markov chain on a finite state space:

    A={α+γgcd⁡(q−p,p)qZ}∩]γ(α−1),1+γα[\mathcal{A} = \left\{ \alpha + \gamma \frac{\gcd(q-p, p)}{q} \mathbb{Z} \right\} \cap \left] \gamma(\alpha - 1), 1 + \gamma\alpha \right[

    where α=p/q\alpha = p/q with (p,q)∈N×N∗(p, q) \in \mathbb{N} \times \mathbb{N}^*, p<qp < q, and gcd⁡\gcd denotes the greatest common divisor. The Markov chain admits a unique stationary distribution πγ\pi_\gamma, and the time-averaged interval length converges almost surely:

    1T∑t=1TL(αt)→T→+∞a.s.Eα~∼πγ[L(α~)].\frac{1}{T} \sum_{t=1}^T L(\alpha_t) \xrightarrow[T \to +\infty]{\text{a.s.}} \mathbb{E}_{\tilde{\alpha} \sim \pi_\gamma} [L(\tilde{\alpha})].

    Furthermore, as the learning rate γ→0\gamma \to 0, the stationary expected interval length satisfies:

    Eπγ[L]=L0+Q′′(1−α)2γα(1−α)+O(γ3/2).\mathbb{E}_{\pi_\gamma}[L] = L_0 + \frac{Q''(1 - \alpha)}{2} \gamma \alpha (1 - \alpha) + O(\gamma^{3/2}).

    For standard error distributions where the probability density function is decreasing in the upper tail, Q′′(1−α)>0Q''(1 - \alpha) > 0. Thus, when data are exchangeable, running ACI increases the expected interval width linearly with γ\gamma relative to standard conformal prediction.

  3. Knowl 3 — Stationarity and Asymptotic Efficiency of ACI under AR(1) Residuals

    theoretical result

    When model residuals follow a stationary autoregressive process of order 1, the joint state of the Adaptive Conformal Inference (ACI) miscoverage level and the lagged residual forms an ergodic Markov chain that admits a unique stationary distribution.

    Let α=p/q∈Q∩(0,1)\alpha = p/q \in \mathbb{Q} \cap (0, 1) with (p,q)∈N×N∗(p, q) \in \mathbb{N} \times \mathbb{N}^*. Assume the residuals εt=Yt−μ^(Xt)\varepsilon_t = Y_t - \hat{\mu}(X_t) evolve as a clipped AR(1) process:

    εt+1=−R∨(ϕεt+ξt+1)∧R\varepsilon_{t+1} = -R \vee (\phi \varepsilon_t + \xi_{t+1}) \wedge R

    where ∣ϕ∣<1|\phi| < 1, R>0R > 0 is a large clipping bound, and (ξt)t≥1(\xi_t)_{t \ge 1} are i.i.d. innovations possessing a continuous probability density with respect to the Lebesgue measure on support S⊇[−R,R]\mathcal{S} \supseteq [-R, R]. Assume the quantile function QQ of the stationary distribution of ∣εt∣|\varepsilon_t| is known.

    The joint process Zt=(αt,εt−1)Z_t = (\alpha_t, \varepsilon_{t-1}) defined on the state space Z=A×[−R,R]\mathcal{Z} = \mathcal{A} \times [-R, R] (where A={α+γgcd⁡(q−p,p)qZ}∩]γ(α−1),1+γα[\mathcal{A} = \{ \alpha + \gamma \frac{\gcd(q-p, p)}{q} \mathbb{Z} \} \cap ]\gamma(\alpha-1), 1+\gamma\alpha[) is a homogeneous, Harris-recurrent Markov chain. It admits a unique stationary distribution πγ,ϕ\pi_{\gamma, \phi}. Consequently, the time-averaged interval length L(αt)=2Q(1−αt)L(\alpha_t) = 2 Q(1 - \alpha_t) converges almost surely:

    1T∑t=1TL(αt)→T→+∞a.s.E(α~,ε~)∼πγ,ϕ[L(α~)].\frac{1}{T} \sum_{t=1}^T L(\alpha_t) \xrightarrow[T \to +\infty]{\text{a.s.}} \mathbb{E}_{(\tilde{\alpha}, \tilde{\varepsilon}) \sim \pi_{\gamma, \phi}} [L(\tilde{\alpha})].
  4. Knowl 4 — Optimal Learning Rate Dynamics of ACI under Autoregressive Dependencies

    theoretical result

    For dependent time series with autoregressive residuals εt+1=ϕεt+ξt+1\varepsilon_{t+1} = \phi \varepsilon_t + \xi_{t+1}, the optimal learning rate γϕ∗=arg⁡min⁡γEπγ,ϕ[L]\gamma^*_\phi = \arg\min_{\gamma} \mathbb{E}_{\pi_{\gamma, \phi}}[L] that minimizes the expected interval length exhibits three key dynamics:

    1. Efficiency Gain for Positive γ\gamma: Unlike the exchangeable case (ϕ=0\phi = 0) where γ∗=0\gamma^* = 0, for correlated residuals (ϕ>0\phi > 0), there exists a strictly positive γϕ∗>0\gamma^*_\phi > 0 such that Eπγϕ∗,ϕ[L]<Eπ0,ϕ[L]\mathbb{E}_{\pi_{\gamma^*_\phi, \phi}}[L] < \mathbb{E}_{\pi_{0, \phi}}[L], proving that ACI strictly reduces interval length in the presence of temporal dependency.
    2. Monotonicity in Correlation Strength: For any fixed learning rate γ\gamma, the expected interval length ϕ↦Eπγ,ϕ[L]\phi \mapsto \mathbb{E}_{\pi_{\gamma, \phi}}[L] is monotonically decreasing in ϕ\phi. Stronger positive autocorrelation between consecutive residuals allows ACI to construct narrower valid prediction intervals.
    3. Non-Monotonicity of Optimal γϕ∗\gamma^*_\phi: The mapping ϕ↦γϕ∗\phi \mapsto \gamma^*_\phi is non-monotonic. Starting from γ0∗=0\gamma^*_0 = 0, the optimal rate first increases for moderate values of ϕ\phi (e.g., ϕ∈[0.6,0.95]\phi \in [0.6, 0.95]) to quickly adapt to short-term temporal dependency. However, as ϕ→1\phi \to 1 (e.g., ϕ≥0.98\phi \ge 0.98), the optimal learning rate decreases because long-memory dynamics require a smaller, more stable step size to prevent producing excessive uninformative infinite intervals (where αt≤0\alpha_t \le 0).
  5. Knowl 5 — Online Expert Aggregation on Adaptive Conformal Inference (AgACI)

    algorithm

    AgACI is a parameter-free conformal prediction algorithm for time series that adaptively combines predictions from a grid of KK parallel ACI instances running with different learning rates {γ1,…,γK}\{\gamma_1, \dots, \gamma_K\}. It uses online expert aggregation (Bernstein Online Aggregation, BOA) with the pinball loss to predict the lower and upper interval bounds separately, eliminating the need to tune γ\gamma.

    Input: Target miscoverage rate α∈(0,1)\alpha \in (0, 1), learning rate grid {γk}k=1K\{\gamma_k\}_{k=1}^K, aggregation rule Φ\Phi (BOA), threshold values M(ℓ),M(u)M^{(\ell)}, M^{(u)}, prediction indices t∈{T0+1,…,T0+T1}t \in \{T_0+1, \dots, T_0+T_1\}.
    Output: Prediction intervals C~t(xt)=[b~t(ℓ)(xt),b~t(u)(xt)]\tilde{C}_t(x_t) = [\tilde{b}_t^{(\ell)}(x_t), \tilde{b}_t^{(u)}(x_t)].
    Let β(ℓ)=α/2\beta^{(\ell)} = \alpha/2 and β(u)=1−α/2\beta^{(u)} = 1 - \alpha/2
    for t=T0+1,…,T0+T1t = T_0+1, \dots, T_0+T_1 do
        for k=1,…,Kk = 1, \dots, K do
            Compute bounds b^t,k(ℓ)(xt)\hat{b}_{t,k}^{(\ell)}(x_t) and b^t,k(u)(xt)\hat{b}_{t,k}^{(u)}(x_t) using ACI with learning rate γk\gamma_k
            if b^t,k(ℓ)(xt)∉R\hat{b}_{t,k}^{(\ell)}(x_t) \notin \mathbb{R} then
                Set b^t,k(ℓ)(xt)=M(ℓ)\hat{b}_{t,k}^{(\ell)}(x_t) = M^{(\ell)}
            end if
            if b^t,k(u)(xt)∉R\hat{b}_{t,k}^{(u)}(x_t) \notin \mathbb{R} then
                Set b^t,k(u)(xt)=M(u)\hat{b}_{t,k}^{(u)}(x_t) = M^{(u)}
            end if
        end for
        for bound type ⋅∈{ℓ,u}\cdot \in \{\ell, u\} do
            for k=1,…,Kk = 1, \dots, K do
                Compute weight ωt,k(⋅)=Φ({ρβ(⋅)(ys,b^s,l(⋅)(xs)):s∈{T0+1,…,t−1},l∈{1,…,K}})\omega_{t,k}^{(\cdot)} = \Phi\left(\{\rho_{\beta^{(\cdot)}}(y_s, \hat{b}_{s,l}^{(\cdot)}(x_s)) : s \in \{T_0+1, \dots, t-1\}, l \in \{1, \dots, K\}\}\right)
            end for
            Compute aggregated bound:
            b~t(⋅)(xt)=∑k=1Kωt,k(⋅)b^t,k(⋅)(xt)∑l=1Kωt,l(⋅)\tilde{b}_t^{(\cdot)}(x_t) = \frac{\sum_{k=1}^K \omega_{t,k}^{(\cdot)} \hat{b}_{t,k}^{(\cdot)}(x_t)}{\sum_{l=1}^K \omega_{t,l}^{(\cdot)}}
        end for
        Set prediction interval C~t(xt)=[b~t(ℓ)(xt),b~t(u)(xt)]\tilde{C}_t(x_t) = [\tilde{b}_t^{(\ell)}(x_t), \tilde{b}_t^{(u)}(x_t)]
        Observe true value yty_t
    end for

    The pinball loss is defined as ρβ(y,y^)=(y−y^)(β−1{y<y^})\rho_\beta(y, \hat{y}) = (y - \hat{y})(\beta - \mathbf{1}\{y < \hat{y}\}). All KK experts share the same underlying regression model μ^t\hat{\mu}_t and calibration set Calt\mathrm{Cal}_t, so running KK parallel experts only adds the minor computational cost of evaluating KK different empirical quantiles.

  6. Knowl 6 — Online Sequential Split Conformal Prediction (OSSCP)

    model/method

    Online Sequential Split Conformal Prediction (OSSCP) adapts Split Conformal Prediction to streaming time series data by combining periodic model refitting with a strictly temporal, sequential train/calibration split.

    At each time step t∈{T0+1,…,T0+T1}t \in \{T_0+1, \dots, T_0+T_1\}, given a sliding window of the most recent T0T_0 historical observations ((xt−T0,yt−T0),…,(xt−1,yt−1))((x_{t-T_0}, y_{t-T_0}), \dots, (x_{t-1}, y_{t-1})):

    1. Sequential Partitioning: The window is partitioned into a proper training set Trt={t−T0,…,t−T0+∣Tr∣−1}\mathrm{Tr}_t = \{t - T_0, \dots, t - T_0 + |\mathrm{Tr}| - 1\} and a calibration set Calt={t−T0+∣Tr∣,…,t−1}\mathrm{Cal}_t = \{t - T_0 + |\mathrm{Tr}|, \dots, t - 1\}. All time indices in Trt\mathrm{Tr}_t strictly precede those in Calt\mathrm{Cal}_t. This deterministic ordering avoids temporal data leakage; in contrast, randomized splitting places future observations into the training set and past ones into calibration, leading to an underestimation of errors on Calt\mathrm{Cal}_t and consequent undercoverage on test data.
    2. Model Refitting: A regression model μ^t\hat{\mu}_t is fitted on Trt\mathrm{Tr}_t.
    3. Score Calibration: Absolute residual conformity scores SCalt={∣yj−μ^t(xj)∣:j∈Calt}S_{\mathrm{Cal}_t} = \{ |y_j - \hat{\mu}_t(x_j)| : j \in \mathrm{Cal}_t \} are evaluated on Calt\mathrm{Cal}_t.
    4. Interval Generation: The prediction interval at xtx_t is constructed as:
    C^α(xt)=[μ^t(xt)±Q^1−αSCP(SCalt)]\hat{C}_\alpha(x_t) = \left[ \hat{\mu}_t(x_t) \pm \hat{Q}_{1-\alpha_{\mathrm{SCP}}}(S_{\mathrm{Cal}_t}) \right]

    with finite-sample quantile level 1−αSCP=(1−α)(1+1/∣Calt∣)1 - \alpha_{\mathrm{SCP}} = (1 - \alpha)(1 + 1/|\mathrm{Cal}_t|).

  7. Knowl 7 — EnbPI V2: Consistent Mean Aggregation for Ensemble Prediction Intervals

    model/method

    Ensemble Prediction Interval (EnbPI) constructs sequential prediction intervals for time series without refitting the base regression model by maintaining a rolling buffer of out-of-sample calibration residuals computed from BB bootstrap models.

    In the original EnbPI formulation, out-of-bag ensemble predictions on training points were aggregated using the mean:

    f^−i(xi)=1∣{b:i∉Sb}∣∑b:i∉Sbf^b(xi),\hat{f}_{-i}(x_i) = \frac{1}{|\{b : i \notin S_b\}|} \sum_{b : i \notin S_b} \hat{f}^b(x_i),

    whereas predictions for test points f^−t(xt)\hat{f}_{-t}(x_t) were computed using the (1−α)(1-\alpha)-quantile of the ensemble predictors {f^b(xt)}b=1B\{\hat{f}^b(x_t)\}_{b=1}^B.

    EnbPI V2 replaces the test-time quantile aggregation with the exact same aggregation function used during calibration (the sample mean):

    f^−t(xt)=1B∑b=1Bf^b(xt).\hat{f}_{-t}(x_t) = \frac{1}{B} \sum_{b=1}^B \hat{f}^b(x_t).

    Using consistent mean aggregation across both training and test phases eliminates an architectural discrepancy that impairs coverage, bringing the empirical coverage of the ensemble prediction intervals substantially closer to the nominal 1−α1-\alpha target.

  8. Knowl 8 — Synthetic Experimental Setup for Time Series Conformal Prediction

    experimental setup

    Conformal prediction methods for time series are benchmarked on synthetic data generated from a non-linear multivariate regression function with autocorrelated noise:

    Yt=10sin⁡(πXt,1Xt,2)+20(Xt,3−0.5)2+10Xt,4+5Xt,5+0Xt,6+εtY_t = 10 \sin(\pi X_{t,1} X_{t,2}) + 20 (X_{t,3} - 0.5)^2 + 10 X_{t,4} + 5 X_{t,5} + 0 X_{t,6} + \varepsilon_t

    where covariates Xt=(Xt,1,…,Xt,6)∼U([0,1]6)X_t = (X_{t,1}, \dots, X_{t,6}) \sim \mathcal{U}([0, 1]^6) are i.i.d., with Xt,6X_{t,6} being an uninformative variable. The noise εt\varepsilon_t follows a Gaussian ARMA(1,1)\mathrm{ARMA}(1, 1) process:

    εt+1=ϕεt+ξt+1+θξt,ξt∼N(0,σ2)\varepsilon_{t+1} = \phi \varepsilon_t + \xi_{t+1} + \theta \xi_t, \quad \xi_t \sim \mathcal{N}(0, \sigma^2)

    where the dependence parameters are varied over ϕ,θ∈{0.1,0.8,0.9,0.95,0.99}\phi, \theta \in \{0.1, 0.8, 0.9, 0.95, 0.99\}. To isolate the effect of temporal dependence structure from variance scaling, the innovation variance is set to maintain a constant asymptotic noise variance lim⁡t→∞Var(εt)=v\lim_{t\to\infty} \mathrm{Var}(\varepsilon_t) = v (tested at v=10v = 10 and v=1v = 1):

    σ2=v1−ϕ21−2ϕθ+θ2.\sigma^2 = v \frac{1 - \phi^2}{1 - 2\phi\theta + \theta^2}.

    The regression model μ^\hat{\mu} is a Random Forest with 1000 trees, a minimum leaf size of 1, and maximum features d=6d = 6. For each parameter configuration, n=500n = 500 independent sequences of length T=T0+T1=300T = T_0 + T_1 = 300 (T0=200T_0 = 200 initial window, T1=100T_1 = 100 sequential evaluation steps) are evaluated at target miscoverage α=0.1\alpha = 0.1. The methods compared include Offline SCP, OSSCP, EnbPI, EnbPI V2, ACI (with γ∈{0.01,0.05}\gamma \in \{0.01, 0.05\}), and AgACI (K=30K = 30 experts).

  9. Knowl 9 — Comparative Validity and Efficiency Across Autoregressive Error Regimes

    empirical result

    Across 500 independent simulations of non-linear time series with ARMA(1,1)\mathrm{ARMA}(1, 1) and AR(1)\mathrm{AR}(1) noise at asymptotic variance 10 and nominal miscoverage α=0.10\alpha = 0.10 (target coverage 90%90\%):

    1. Online Adaptation: Online sequential split conformal prediction (OSSCP) significantly outperforms offline SCP, which exhibits severe validity degradation over time.
    2. OSSCP Coverage Loss: OSSCP achieves valid coverage (≥90%\ge 90\%) when autocorrelation is weak (ϕ,θ≤0.8\phi, \theta \le 0.8), but degrades as temporal dependence increases, falling below 84%84\% empirical coverage at ϕ=θ=0.99\phi = \theta = 0.99.
    3. EnbPI Vulnerability to Autoregressive Noise: Both EnbPI and EnbPI V2 lose coverage as autocorrelation increases in AR(1)\mathrm{AR}(1) and ARMA(1,1)\mathrm{ARMA}(1,1) noise (dropping toward 83%83\% coverage at ϕ=0.99\phi = 0.99). EnbPI achieves nominal coverage only for MA(1)\mathrm{MA}(1) noise, where temporal memory is strictly 1-lag.
    4. ACI Robustness: ACI with γ=0.05\gamma = 0.05 maintains empirical coverage near the nominal 90%90\% target across all dependence regimes ϕ,θ∈[0.1,0.99]\phi, \theta \in [0.1, 0.99], showing strong robustness to autoregressive structure.
    5. AgACI Performance: AgACI achieves valid coverage (≥89.8%\ge 89.8\%) across all values of ϕ\phi and θ\theta, while producing the shortest median prediction interval lengths among all methods that maintain validity.
  10. Knowl 10 — Infinite Interval Proportions in ACI under Varying Learning Rates

    data/table

    In Adaptive Conformal Inference (ACI), when the effective miscoverage parameter αt≤0\alpha_t \le 0, the algorithm outputs an uninformative infinite prediction interval C^αt(x)=R\hat{C}_{\alpha_t}(x) = \mathbb{R}. The table below reports the percentage of infinite intervals produced by ACI with γ=0.01\gamma = 0.01 versus γ=0.05\gamma = 0.05 across 500 simulation runs (T0=200,T1=100T_0 = 200, T_1 = 100, target α=0.1\alpha = 0.1, asymptotic variance 10), alongside the intersection rate (the number of points where γ=0.05\gamma = 0.05 predicted R\mathbb{R} and γ=0.01\gamma = 0.01 failed to cover):

    Noise parameters γ=0.01\gamma = 0.01 (%) γ=0.05\gamma = 0.05 (%) Intersection
    ϕ=θ=0.1\phi = \theta = 0.1 0.00 1.12 53 out of 562 (9.43%)
    ϕ=θ=0.8\phi = \theta = 0.8 0.00 2.76 263 out of 1381 (19.04%)
    ϕ=θ=0.9\phi = \theta = 0.9 0.00 3.72 425 out of 1862 (22.83%)
    ϕ=θ=0.95\phi = \theta = 0.95 0.03 4.45 514 out of 2224 (23.11%)
    ϕ=θ=0.99\phi = \theta = 0.99 0.04 6.22 554 out of 3109 (17.82%)
    ϕ=0.1,θ=0\phi = 0.1, \theta = 0 0.00 1.00 37 out of 500 (7.40%)
    ϕ=0.8,θ=0\phi = 0.8, \theta = 0 0.00 2.75 212 out of 1373 (15.44%)
    ϕ=0.9,θ=0\phi = 0.9, \theta = 0 0.00 3.24 359 out of 1622 (22.13%)
    ϕ=0.95,θ=0\phi = 0.95, \theta = 0 0.03 4.32 488 out of 2160 (22.59%)
    ϕ=0.99,θ=0\phi = 0.99, \theta = 0 0.06 6.15 560 out of 3073 (18.22%)
    θ=0.1,ϕ=0\theta = 0.1, \phi = 0 0.00 1.03 38 out of 516 (7.36%)
    θ=0.8,ϕ=0\theta = 0.8, \phi = 0 0.00 1.42 49 out of 710 (6.90%)
    θ=0.9,ϕ=0\theta = 0.9, \phi = 0 0.00 1.54 47 out of 772 (6.09%)
    θ=0.95,ϕ=0\theta = 0.95, \phi = 0 0.00 1.54 45 out of 770 (5.84%)
    θ=0.99,ϕ=0\theta = 0.99, \phi = 0 0.00 1.56 53 out of 781 (6.79%)

    Setting γ=0.01\gamma = 0.01 generates virtually zero infinite intervals (≤0.06%\le 0.06\% across all settings). In contrast, γ=0.05\gamma = 0.05 produces between 1.00%1.00\% and 6.22%6.22\% infinite intervals. Less than 25%25\% of these infinite intervals occur at points where γ=0.01\gamma = 0.01 miscovered, indicating that larger learning rates frequently produce uninformative infinite intervals to compensate for easy errors.

  11. Knowl 11 — Day-Ahead French Electricity Spot Price Forecasting Application

    empirical result

    Conformal prediction methods were evaluated on hourly French electricity spot prices from 2016 to 2019 (35,064 hourly observations). The objective is predicting the 24 hourly prices of day D+1D+1 before 12:00 PM on day DD, using day-ahead forecast consumption, day of week, 24 hourly prices from day D−1D-1, and 24 hourly prices from day D−7D-7.

    Individual Random Forest models (1000 trees) were trained for each of the 24 hours on a 3-year rolling window (1.5 years training, 1.5 years calibration), generating sequential prediction intervals across all 8,760 hours of 2019 at target coverage 1−α=0.901-\alpha = 0.90:

    1. Offline SCP vs. OSSCP: Offline SCP over-covers (>93.5%>93.5\%) with a large median interval width (≈26.5\approx 26.5 €/MWh) because extreme price spikes in 2016 (reaching 800 €/MWh) permanently inflate the calibration quantile. OSSCP recalibrates over the sliding window, achieving 93%93\% coverage with a narrower median width (≈24.8\approx 24.8 €/MWh).
    2. EnbPI V2: Achieves valid coverage (≈92.7%\approx 92.7\%) with median width ≈24.5\approx 24.5 €/MWh.
    3. ACI and AgACI: ACI with γ=0.01\gamma = 0.01 and γ=0.05\gamma = 0.05 yields the most efficient valid intervals, with median lengths of ≈21.8\approx 21.8 €/MWh and ≈23.0\approx 23.0 €/MWh, respectively. AgACI achieves valid empirical coverage (≈91.5%\approx 91.5\%) with a median length of ≈23.8\approx 23.8 €/MWh, improving efficiency over the non-adaptive baseline γ=0\gamma = 0 (median length ≈25.0\approx 25.0 €/MWh) without requiring parameter tuning.
  12. Knowl 12 — Day-of-the-Week Conditional Undercoverage in Time Series Conformal Prediction

    limitation

    Although online conformal methods (OSSCP, EnbPI V2, ACI, and AgACI) achieve valid marginal coverage (≥90%\ge 90\%) over the full 2019 French electricity spot price test period, all methods suffer from conditional coverage discrepancies across days of the week:

    • Tuesdays through Fridays: All algorithms systematically over-cover, attaining empirical coverage rates between 92%92\% and 95%95\%.
    • Mondays and Weekends (Saturdays and Sundays): All algorithms systematically under-cover, dropping to empirical coverage rates of approximately 88%88\% (below the nominal 90%90\% target).

    This discrepancy arises because electricity consumption patterns and price dynamics differ substantially between weekdays and weekends. Standard conformal inference guarantees apply only marginally over time, failing to guarantee conditional validity across cyclic calendar sub-periods.

Coverage note — Deliberately omitted the toy 1D synthetic visualization in Model A.5 and the generic mathematical descriptions of BOA/OPERA aggregation rules from prior literature, as they represent illustrative examples or standard background material rather than original contributions.

References

  1. 1.Angelopoulos, A. N. and Bates, S. A gentle introduction to conformal prediction and distribution-free uncertainty quantification. arXiv preprint arXiv:2107.07511, 2021.
  2. 2.Berrisch, J. and Ziel, F. CRPS learning. Journal of Econometrics, 2021.
  3. 3.Cai, Y. and Davies, N. A simple bootstrap method for time series. Communications in Statistics-Simulation and Computation, 41(5):621–631, 2012.
  4. 4.Cauchois, M., Gupta, S., Ali, A., and Duchi, J. C. Robust Validation: Confident Predictions Even When Distributions Shift. arXiv preprint arXiv:2008.04267, 2020.
  5. 5.Cesa-Bianchi, N. and Lugosi, G. Prediction, learning, and games. Cambridge University Press, 2006.
  6. 6.Chernozhukov, V., Wuthrich, K., and Yinchu, Z. Exact and Robust Conformal Inference Methods for Predictive Machine Learning with Dependent Data. In Conference On Learning Theory, pp. 732–749. PMLR, 2018.
  7. 7.Dashevskiy, M. and Luo, Z. Network traffic demand prediction with confidence. In IEEE Global Telecommunications Conference. IEEE, 2008.
  8. 8.Dashevskiy, M. and Luo, Z. Time series prediction with performance guarantee. IET communications, 5(8):1044–1051, 2011.
  9. 9.EUPHEMIA. Euphemia public description, single price coupling algorithm, April 2019. URL https://www.nemo-committee.eu/assets/files/190410_Euphemia%20Public%20Description%20version%20NEMO%20Committee.pdf.
  10. 10.Friedman, J. H., Grosse, E., and Stuetzle, W. Multidimensional additive spline approximation. SIAM J. Sci. Stat. Comput., 1983.
  11. 11.Gaillard, P. and Goude, Y. OPERA, 2021. URL https://cran.r-project.org/package=opera. R package version 1.2.0.
  12. 12.Gaillard, P., Stoltz, G., and Van Erven, T. A second-order bound with excess losses. In Conference on Learning Theory, pp. 176–196. PMLR, 2014.
  13. 13.Gaillard, P., Goude, Y., and Nedellec, R. Additive models and robust aggregation for GEFCom2014 probabilistic electric load and electricity price forecasting. International Journal of Forecasting, 32(3):1038–1050, 2016.
  14. 14.Gibbs, I. and Cand`es, E. Adaptive conformal inference under distribution shift. In Advances in Neural Information Processing Systems, 2021.
  15. 15.Goehry, B. Random forests for time-dependent processes. ESAIM: Probability and Statistics, 24:801–826, 2020.
  16. 16.Goehry, B., Yan, H., Goude, Y., Massart, P., and Poggi, J.-M. Random forests for time series. HAL hal-03129751, 2021.
  17. 17.H¨ardle, W., Horowitz, J., and Kreiss, J.-P. Bootstrap methods for time series. International Statistical Review, 71(2):435–459, 2003.
  18. 18.Hazan, E. Introduction to online convex optimization. arXiv preprint arXiv:1909.05207, 2019.
  19. 19.Kath, C. and Ziel, F. Conformal prediction interval estimation and applications to day-ahead and intraday power markets. International Journal of Forecasting, 37(2):777–799, 2021.
  20. 20.Kreiss, J.-P. and Paparoditis, E. The hybrid wild bootstrap for time series. Journal of the American Statistical Association, 107(499):1073–1084, 2012.
  21. 21.Lago, J., De Ridder, F., and De Schutter, B. Forecasting spot electricity prices: Deep learning approaches and empirical comparison of traditional algorithms. Applied Energy, 221:386–405, 2018.
  22. 22.Lago, J., Marcjasz, G., De Schutter, B., and Weron, R. Forecasting day-ahead electricity prices: A review of state-of-the-art algorithms, best practices and an openaccess benchmark. Applied Energy, 293:116983, 2021.
  23. 23.Lei, J., G’Sell, M., Rinaldo, A., Tibshirani, R. J., and Wasserman, L. Distribution-Free Predictive Inference for Regression. Journal of the American Statistical Association, 113(523):1094–1111, 2018.
  24. 24.Maciejowska, K., Nowotarski, J., and Weron, R. Probabilistic forecasting of electricity spot prices using Factor Quantile Regression Averaging. International Journal of Forecasting, 32(3):957–965, 2016.
  25. 25.Meyn, S. P. and Tweedie, R. L. Markov chains and stochastic stability. Springer Science & Business Media, 2012.
  26. 26.Nowotarski, J. and Weron, R. Recent advances in electricity price forecasting: A review of probabilistic forecasting. Renewable and Sustainable Energy Reviews, 81:1548–1568, 2018.
  27. 27.Papadopoulos, H., Proedrou, K., Vovk, V., and Gammerman, A. Inductive Confidence Machines for Regression. In Machine Learning: ECML 2002, pp. 345–356. Springer, 2002.
  28. 28.Remlinger, C., Bri`ere, M., Alasseur, C., and Mikael, J. Expert aggregation for financial forecasting. arXiv preprint arXiv:2111.15365, 2021.
  29. 29.Romano, Y., Patterson, E., and Cand`es, E. Conformalized Quantile Regression. Advances in Neural Information Processing Systems, 32, 2019.
  30. 30.Saha, A., Basu, S., and Datta, A. Random forests for spatially dependent data. Journal of the American Statistical Association, 0(0):1–19, 2021.
  31. 31.Shafer, G. and Vovk, V. A Tutorial on Conformal Prediction. JMLR, 9:51, 2008.
  32. 32.Shalev-Shwartz, S. and Singer, Y. A primal-dual perspective of online learning algorithms. Machine Learning, 69(2-3):115–142, 2007.
  33. 33.Tibshirani, R. J., Barber, R. F., Cand`es, E., and Ramdas, A. Conformal Prediction Under Covariate Shift. Advances in Neural Information Processing Systems, 32:11, 2019.
  34. 34.Uniejewski, B. and Weron, R. Regularized quantile regression averaging for probabilistic electricity price forecasting. Energy Economics, pp. 105121, 2021.
  35. 35.Vovk, V., Gammerman, A., and Saunders, C. Machine-Learning Applications of Algorithmic Randomness. In Proceedings of the Sixteenth International Conference on Machine Learning, pp. 444–453. Morgan Kaufmann Publishers Inc., 1999.
  36. 36.Vovk, V., Gammerman, A., and Shafer, G. Algorithmic Learning in a Random World. Springer US, 2005.
  37. 37.Vovk, V. G. Aggregating strategies. Proc. of Computational Learning Theory, 1990.
  38. 38.Weron, R. Electricity price forecasting: A review of the state-of-the-art with a look into the future. International Journal of Forecasting, 30(4):1030–1081, 2014.
  39. 39.Wintenberger, O. Optimal learning with bernstein online aggregation. Machine Learning, 106(1):119–141, 2017.
  40. 40.Wisniewski, W., Lindsay, D., and Lindsay, S. Application of conformal prediction interval estimations to market makers’ net positions. In Proceedings of the Ninth Symposium on Conformal and Probabilistic Prediction and Applications, volume 128 of Proceedings of Machine Learning Research, pp. 285–301. PMLR, 2020.
  41. 41.Xu, C. and Xie, Y. Conformal prediction interval for dynamic time-series. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 11559–11569. PMLR, 2021a.
  42. 42.Xu, C. and Xie, Y. Conformal prediction for dynamic timeseries. arXiv preprint arXiv:2010.09107, 2021b.

Citation

MLA
Zaffran, M., et al. “Adaptive Conformal Predictions for Time Series”. International Conference on Machine Learning, vol. 162, 2022, pp. 25834–66, https://proceedings.mlr.press/v162/zaffran22a.html.
APA
Zaffran, M., Feron, O., Goude, Y., Josse, J., & Dieuleveut, A. (2022). Adaptive Conformal Predictions for Time Series. International Conference on Machine Learning, 162, 25834–25866. https://proceedings.mlr.press/v162/zaffran22a.html
Chicago
Zaffran, M., O. Feron, Y. Goude, J. Josse, and A. Dieuleveut. 2022. “Adaptive Conformal Predictions for Time Series”. International Conference on Machine Learning 162: 25834–66. https://proceedings.mlr.press/v162/zaffran22a.html.
Harvard
Zaffran, M. et al. (2022) “Adaptive Conformal Predictions for Time Series”, International Conference on Machine Learning. PMLR, pp. 25834–25866. Available at: https://proceedings.mlr.press/v162/zaffran22a.html.
Vancouver
1. Zaffran M, Feron O, Goude Y, Josse J, Dieuleveut A (2022) Adaptive Conformal Predictions for Time Series. In: International Conference on Machine Learning. PMLR, pp 25834–25866

BibTeX

@InProceedings{pmlr-v162-zaffran22a,
  title = 	 {Adaptive Conformal Predictions for Time Series},
  author =       {Zaffran, Margaux and Feron, Olivier and Goude, Yannig and Josse, Julie and Dieuleveut, Aymeric},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {25834--25866},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/zaffran22a/zaffran22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/zaffran22a.html},
  abstract = 	 {Uncertainty quantification of predictive models is crucial in decision-making problems. Conformal prediction is a general and theoretically sound answer. However, it requires exchangeable data, excluding time series. While recent works tackled this issue, we argue that Adaptive Conformal Inference (ACI, Gibbs & Cand{è}s, 2021), developed for distribution-shift time series, is a good procedure for time series with general dependency. We theoretically analyse the impact of the learning rate on its efficiency in the exchangeable and auto-regressive case. We propose a parameter-free method, AgACI, that adaptively builds upon ACI based on online expert aggregation. We lead extensive fair simulations against competing methods that advocate for ACI’s use in time series. We conduct a real case study: electricity price forecasting. The proposed aggregation algorithm provides efficient prediction intervals for day-ahead forecasting. All the code and data to reproduce the experiments are made available on GitHub.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/