Conformal Inference for Online Prediction with Arbitrary Distribution Shifts

Isaac GibbsEmmanuel J. Candès

article2024JMLR147 citations

Develops an adaptive conformal inference method that dynamically tunes its update step size to construct valid prediction sets under arbitrary, unknown distribution shifts without relying on heavy historical weighting.

Listen

Modern predictive machine learning models are widely deployed in critical real-time systems, yet their reliability often breaks down when real-world conditions evolve over time. Existing uncertainty quantification techniques, such as conformal prediction sets, typically rely on the assumption that incoming data behave identically to historical data. While adaptive conformal inference methods were developed to adjust prediction bands dynamically, earlier approaches require advance knowledge of how rapidly the environment changes or over-weight older historical data, causing them to lag dangerously behind sudden shifts.

The article demonstrates and evaluates a new framework called dynamically-tuned adaptive conformal inference (DtACI). Its primary objective is to automatically generate valid prediction intervals in real time under arbitrary, unknown distribution shifts without requiring distributional assumptions or pre-tuned update rates.

To achieve this, the article frames interval tuning as an online optimization problem and deploys an expert-aggregation scheme that runs multiple step sizes in parallel. By continuously re-weighting these candidate parameters using recent loss performance, the system dynamically selects the optimal adjustment speed. The researchers established theoretical guarantees that bound prediction errors across local time windows and validated the method through numerical simulations and two real-world case studies: forecasting daily stock market volatility and predicting county-level weekly COVID-19 case counts across multiple US regions.

The analysis reveals several key findings. First, DtACI achieves provably small regret and tightly maintains the target coverage level (such as 90% accuracy) over any local time interval, scaling directly with the rate of environmental drift. Second, under sudden jump shifts, DtACI rapidly adjusts its update rate, whereas competing methods lag significantly because they assign excessive weight to obsolete historical records. Third, in real-world market and epidemiological testing, the method matched the theoretical error rate of an idealized process and successfully adapted whether point models were well-calibrated or highly unstable. Fourth, testing confirmed that DtACI accurately learns underlying target values rather than merely oscillating reactively between extreme over-coverage and under-coverage.

These findings mean organizations deploying automated forecasting can reliably quantify uncertainty without risking systematic failure during regime changes, financial disruptions, or public health emergencies. Practitioners do not need prior knowledge of market volatility speeds or disease transmission shifts to maintain reliable statistical guarantees.

For operational deployment, the article recommends implementing DtACI alongside existing predictive point or quantile models using fixed hyperparameter defaults, which perform robustly across diverse tasks. If an environment is known to remain perfectly static over long horizons, traditional methods may offer slightly narrower calibration; however, for dynamic real-world settings, DtACI provides the best trade-off by preventing delayed reactions to abrupt disruptions.

A primary boundary condition is that when hyperparameters are configured for maximum local adaptivity, long-term average coverage may exhibit a minor statistical bias, though empirical evidence shows this effect is negligible in practice. Confidence in the approach is high, supported by mathematical proofs and validated performance across real-world data streams.

  • Paper: Replicable Conformal Prediction, Marios Papamichalis et al. (2026). It addresses deployment-level replicability and stability bottlenecks in conformal prediction sets calibrated across independent analysts or shifting operational audits.
Cover for Conformal Inference for Online Prediction with Arbitrary Distribution Shifts

Abstract

We consider the problem of forming prediction sets in an online setting where the distribution generating the data is allowed to vary over time. Previous approaches to this problem suffer from over-weighting historical data and thus may fail to quickly react to the underlying dynamics. Here, we correct this issue and develop a novel procedure with provably small regret over all local time intervals of a given width. We achieve this by modifying the adaptive conformal inference (ACI) algorithm of Gibbs and Candès (2021) to contain an additional step in which the step-size parameter of ACI’s gradient descent update is tuned over time. Crucially, this means that unlike ACI, which requires knowledge of the rate of change of the data-generating mechanism, our new procedure is adaptive to both the size and type of the distribution shift. Our methods are highly flexible and can be used in combination with any baseline predictive algorithm that produces point estimates or estimated quantiles of the target without the need for distributional assumptions. We test our techniques on two real-world datasets aimed at predicting stock market volatility and COVID-19 case counts and find that they are robust and adaptive to real-world distribution shifts.

Table of Contents

  • 1. Introduction
  • 2. Methodology
  • 2.1 Conformal Inference
  • 2.2 Adaptive Conformal Inference
  • 2.3 Dynamically-Tuned Adaptive Conformal Inference
  • 2.4 Comparison to Existing Methods
  • 3. Coverage Properties of DtACI
  • 3.1 Dynamic Regret of DtACI
  • 3.2 Bounds on the Short-Term Coverage
  • 3.3 Bounds on the Long-Term Coverage
  • 3.4 Removing Randomness in the Choice of α t α t α t
  • 4. Empirical Results
  • 4.1 Simulated Examples
  • 4.2 Real Data Examples
  • 4.2.1 Online Prediction in the Stock Market
  • 4.2.2 Evaluating the Reactivity of DtACI
  • 4.2.3 Predicting Covid-19 Case Counts
  • Acknowledgments
  • Appendix A. Detailed Description of Multivalid Conformal Prediction
  • Appendix B. Details of the Block Bootstrap for Section 4.2.2
  • Appendix C. Proofs for Section 3
  • C.1 Proof of Lemma 2
  • C.2 Proof of Lemma 3
  • C.3 Proof of Theorem 4
  • C.4 Results for Variable η
  • C.5 Proof of Proposition 5
  • C.6 Proof of Theorem 6
  • Appendix D. Additional Figures
  • References

Knowls

  1. Knowl 1 — Dynamically-Tuned Adaptive Conformal Inference

    algorithm

    Dynamically-Tuned Adaptive Conformal Inference (DtACI) adaptively tunes the nominal error rate αt\alpha_t in online conformal prediction by aggregating multiple parallel instances (experts) of Adaptive Conformal Inference (ACI), each running with a different step size γi\gamma_i. At each time step tt, the algorithm computes a convex combination αˉt\bar{\alpha}_t of the expert parameters, forms a prediction set C^t(αˉt)\hat{C}_t(\bar{\alpha}_t), observes the ground-truth outcome YtY_t, updates exponential tracking weights wtiw_t^i via the pinball loss ℓ(βt,αti)\ell(\beta_t, \alpha_t^i), injects a mixing probability σ\sigma to enable fast adaptation to abrupt shifts, and updates each expert's parameter αti\alpha_t^i via gradient descent.

    Input: Sequence of observed minimum-containing threshold values {βt\beta_t}t=1T_{t=1}^T, candidate step sizes {γi\gamma_i}i=1k_{i=1}^k, initial points {α1i\alpha_1^i}i=1k_{i=1}^k, mixing parameter σ∈(0,1/2]\sigma \in (0, 1/2], learning rate η>0\eta > 0, nominal coverage level 1−α∈(0,1)1 - \alpha \in (0, 1).
    Initialize expert weights w1i←1w_1^i \leftarrow 1 for all i∈{1,…,k}i \in \{1, \dots, k\}
    for t=1,2,…,Tt = 1, 2, \dots, T do
        Compute probabilities pti←wti/∑j=1kwtjp_t^i \leftarrow w_t^i / \sum_{j=1}^k w_t^j for all i∈{1,…,k}i \in \{1, \dots, k\}
        Output deterministic estimate αˉt←∑i=1kptiαti\bar{\alpha}_t \leftarrow \sum_{i=1}^k p_t^i \alpha_t^i
        Form prediction set C^t(αˉt)\hat{C}_t(\bar{\alpha}_t) and observe label YtY_t
        Compute smallest containing quantile $\beta_t \leftarrow \sup\{\beta : Y_t \in \hat{C}_t(\beta)\}
        for i=1,…,ki = 1, \dots, k do
            Compute pinball loss $\ell(\beta_t, \alpha_t^i) \leftarrow \alpha(\beta_t - \alpha_t^i) - \min\{0, \beta_t - \alpha_t^i\}
            Update intermediate weight $\bar{w}_t^i \leftarrow w_t^i \exp(-\eta \ell(\beta_t, \alpha_t^i))
        end for
        Compute total intermediate weight Wˉt←∑i=1kwˉti\bar{W}_t \leftarrow \sum_{i=1}^k \bar{w}_t^i
        for i=1,…,ki = 1, \dots, k do
            Update expert weight $w_{t+1}^i \leftarrow (1 - \sigma) \bar{w}_t^i + \bar{W}_t \frac{\sigma}{k}
            Compute expert error indicator $\text{err}_t^i \leftarrow \mathbf{1}\{Y_t \notin \hat{C}_t(\alpha_t^i)\}
            Update expert parameter $\alpha_{t+1}^i \leftarrow \alpha_t^i + \gamma_i (\alpha - \text{err}_t^i)
        end for
    end for
  2. Knowl 2 — Dynamic Regret Bound for DtACI

    theoretical result

    Let candidate step sizes satisfy 0<γ1<γ2<⋯<γk0 < \gamma_1 < \gamma_2 < \dots < \gamma_k with γi+1/γi≤2\gamma_{i+1}/\gamma_i \le 2 for all 1≤i<k1 \le i < k, and let γmax⁡:=max⁡1≤i≤kγi\gamma_{\max} := \max_{1 \le i \le k} \gamma_i. Consider any local time interval I=[r,s]⊆[T]I = [r, s] \subseteq [T] of length ∣I∣=s−r+1|I| = s - r + 1, and let γk≥1+1/∣I∣\gamma_k \ge \sqrt{1 + 1/|I|} and σ≤1/2\sigma \le 1/2. For any reference sequence of target parameters αr∗,…,αs∗∈[0,1]\alpha_r^*, \dots, \alpha_s^* \in [0, 1], the dynamic regret of DtACI (under either randomized expert selection αt\alpha_t or deterministic averaging αˉt\bar{\alpha}_t) with respect to the pinball loss ℓ(βt,θ)=α(βt−θ)−min⁡{0,βt−θ}\ell(\beta_t, \theta) = \alpha(\beta_t - \theta) - \min\{0, \beta_t - \theta\} satisfies:

    1∣I∣∑t=rsE[ℓ(βt,αt)]−1∣I∣∑t=rsℓ(βt,αt∗)≤log⁡(k/σ)+2σ∣I∣η∣I∣+η∣I∣∑t=rsE[ℓ(βt,αt)2]+4(1+γmax⁡)2max⁡{∑t=r+1s∣αt∗−αt−1∗∣+1∣I∣,γ1}\frac{1}{|I|}\sum_{t=r}^s \mathbb{E}[\ell(\beta_t, \alpha_t)] - \frac{1}{|I|}\sum_{t=r}^s \ell(\beta_t, \alpha_t^*) \le \frac{\log(k/\sigma) + 2\sigma|I|}{\eta |I|} + \frac{\eta}{|I|}\sum_{t=r}^s \mathbb{E}[\ell(\beta_t, \alpha_t)^2] + 4(1 + \gamma_{\max})^2 \max\left\{ \frac{\sum_{t=r+1}^s |\alpha_t^* - \alpha_{t-1}^*| + 1}{|I|}, \gamma_1 \right\}

    where the expectation is over the internal randomness of the algorithm and the sequence {βt}\{\beta_t\} is treated as fixed.

    Setting σ=12∣I∣\sigma = \frac{1}{2|I|} and η=log⁡(2k∣I∣)+1∑t=rsE[ℓ(βt,αt)2]\eta = \sqrt{\frac{\log(2k|I|) + 1}{\sum_{t=r}^s \mathbb{E}[\ell(\beta_t, \alpha_t)^2]}} (assuming γ1≤∑t=r+1s∣αt∗−αt−1∗∣+1∣I∣\gamma_1 \le \sqrt{\frac{\sum_{t=r+1}^s |\alpha_t^* - \alpha_{t-1}^*| + 1}{|I|}}) yields the simplified rate:

    1∣I∣∑t=rsE[ℓ(βt,αt)]−1∣I∣∑t=rsℓ(βt,αt∗)≤O(log⁡(∣I∣)∣I∣)+O(∑t=r+1s∣αt∗−αt−1∗∣∣I∣).\frac{1}{|I|}\sum_{t=r}^s \mathbb{E}[\ell(\beta_t, \alpha_t)] - \frac{1}{|I|}\sum_{t=r}^s \ell(\beta_t, \alpha_t^*) \le \mathcal{O}\left(\sqrt{\frac{\log(|I|)}{|I|}}\right) + \mathcal{O}\left(\sqrt{\frac{\sum_{t=r+1}^s |\alpha_t^* - \alpha_{t-1}^*|}{|I|}}\right).

  3. Knowl 3 — Local Estimation and Coverage Gap Bounds via Excess Pinball Loss

    theoretical result

    For any random variable β∈[0,1]\beta \in [0, 1] satisfying P(β<α∗)=α\mathbb{P}(\beta < \alpha^*) = \alpha, the excess pinball loss ℓ(β,τ)=α(β−τ)−min⁡{0,β−τ}\ell(\beta, \tau) = \alpha(\beta - \tau) - \min\{0, \beta - \tau\} relative to α∗\alpha^* satisfies the identity:

    E[ℓ(β,τ)]−E[ℓ(β,α∗)]={E[(τ−β)1α∗<β≤τ],if τ≥α∗,E[(β−τ)1τ<β≤α∗],if τ<α∗.\mathbb{E}[\ell(\beta, \tau)] - \mathbb{E}[\ell(\beta, \alpha^*)] = \begin{cases} \mathbb{E}[(\tau - \beta)\mathbf{1}_{\alpha^* < \beta \le \tau}], & \text{if } \tau \ge \alpha^*, \\ \mathbb{E}[(\beta - \tau)\mathbf{1}_{\tau < \beta \le \alpha^*}], & \text{if } \tau < \alpha^*. \end{cases}

    If β\beta has a density function p(⋅)p(\cdot) on [0,1][0, 1] with p(x)≥p‾>0p(x) \ge \underline{p} > 0 for all x∈[0,1]x \in [0, 1], then:

    E[ℓ(β,τ)]−E[ℓ(β,α∗)]≥p‾(τ−α∗)22.\mathbb{E}[\ell(\beta, \tau)] - \mathbb{E}[\ell(\beta, \alpha^*)] \ge \frac{\underline{p}(\tau - \alpha^*)^2}{2}.

    Consequently, letting αt∗\alpha_t^* denote the value satisfying P(Yt∈C^t(αt∗)∣{βs}s<t)=1−α\mathbb{P}(Y_t \in \hat{C}_t(\alpha_t^*) \mid \{\beta_s\}_{s<t}) = 1 - \alpha, the local root-mean-squared estimation error of DtACI over any interval I=[r,s]I = [r, s] satisfies:

    1∣I∣∑t=rsp‾E[(αt−αt∗)2]2≤O(log⁡(∣I∣)∣I∣)+O(∑t=r+1sE[∣αt∗−αt−1∗∣]∣I∣)\frac{1}{|I|}\sum_{t=r}^s \frac{\underline{p}\mathbb{E}[(\alpha_t - \alpha_t^*)^2]}{2} \le \mathcal{O}\left(\sqrt{\frac{\log(|I|)}{|I|}}\right) + \mathcal{O}\left(\sqrt{\frac{\sum_{t=r+1}^s \mathbb{E}[|\alpha_t^* - \alpha_{t-1}^*|]}{|I|}}\right)

    where p‾\underline{p} lower bounds the conditional density of βt\beta_t given past observations. Furthermore, if β↦P(Yt∈C^t(β)∣{βs}s<t)\beta \mapsto \mathbb{P}(Y_t \in \hat{C}_t(\beta) \mid \{\beta_s\}_{s<t}) is LL-Lipschitz, this yields an explicit bound on the local coverage gap ∣P(Yt∈C^t(αt)∣{βs}s<t)−(1−α)∣≤L∣αt−αt∗∣|\mathbb{P}(Y_t \in \hat{C}_t(\alpha_t) \mid \{\beta_s\}_{s<t}) - (1 - \alpha)| \le L |\alpha_t - \alpha_t^*|.

  4. Knowl 4 — Equivalence of Adaptive Conformal Inference and Online Gradient Descent on Pinball Loss

    model/method

    The standard Adaptive Conformal Inference (ACI) update equation:

    αt+1=αt+γ(α−errt),where errt=1{Yt∉C^t(αt)}\alpha_{t+1} = \alpha_t + \gamma (\alpha - \text{err}_t), \quad \text{where } \text{err}_t = \mathbf{1}\{Y_t \notin \hat{C}_t(\alpha_t)\}

    can be rewritten as an online gradient descent step with respect to the pinball loss. Let βt:=sup⁡{β:Yt∈C^t(β)}\beta_t := \sup\{\beta : Y_t \in \hat{C}_t(\beta)\} be the parameter producing the smallest prediction set that contains YtY_t, and define the pinball loss:

    ℓ(βt,θ):=α(βt−θ)−min⁡{0,βt−θ}.\ell(\beta_t, \theta) := \alpha(\beta_t - \theta) - \min\{0, \beta_t - \theta\}.

    The subgradient with respect to θ\theta is ∇θℓ(βt,αt)=errt−α\nabla_\theta \ell(\beta_t, \alpha_t) = \text{err}_t - \alpha (taking the subgradient to be 00 when βt=αt\beta_t = \alpha_t). Thus, the ACI update is equivalent to:

    αt+1=αt−γ∇θℓ(βt,αt).\alpha_{t+1} = \alpha_t - \gamma \nabla_\theta \ell(\beta_t, \alpha_t).

    This equivalence allows the step size γ\gamma to be learned dynamically using online convex optimization and expert tracking methods.

  5. Knowl 5 — Asymptotic Long-Term Coverage of DtACI with Decaying Hyperparameters

    theoretical result

    Consider DtACI running with time-varying hyperparameters ηt\eta_t and σt\sigma_t at step tt. Let γmin⁡:=min⁡1≤i≤kγi\gamma_{\min} := \min_{1 \le i \le k} \gamma_i and γmax⁡:=max⁡1≤i≤kγi\gamma_{\max} := \max_{1 \le i \le k} \gamma_i. For any horizon TT, the deviation of the cumulative average miscoverage from the nominal level α\alpha is bounded by:

    ∣1T∑t=1TE[errt]−α∣≤1+2γmax⁡Tγmin⁡+(1+2γmax⁡)2γmin⁡1T∑t=1Tηteηt(1+2γmax⁡)+21+γmax⁡γmin⁡1T∑t=1Tσt\left| \frac{1}{T}\sum_{t=1}^T \mathbb{E}[\text{err}_t] - \alpha \right| \le \frac{1 + 2\gamma_{\max}}{T \gamma_{\min}} + \frac{(1 + 2\gamma_{\max})^2}{\gamma_{\min}} \frac{1}{T}\sum_{t=1}^T \eta_t e^{\eta_t(1 + 2\gamma_{\max})} + 2\frac{1 + \gamma_{\max}}{\gamma_{\min}} \frac{1}{T}\sum_{t=1}^T \sigma_t

    where errt=1{Yt∉C^t(αt)}\text{err}_t = \mathbf{1}\{Y_t \notin \hat{C}_t(\alpha_t)\}.

    If the learning rate and mixing parameter decay to zero as t→∞t \to \infty, i.e., lim⁡t→∞ηt=0\lim_{t \to \infty} \eta_t = 0 and lim⁡t→∞σt=0\lim_{t \to \infty} \sigma_t = 0, then exact long-term empirical coverage holds almost surely:

    lim⁡T→∞1T∑t=1Terrt=a.s.α.\lim_{T \to \infty} \frac{1}{T}\sum_{t=1}^T \text{err}_t \stackrel{\text{a.s.}}{=} \alpha.

  6. Knowl 6 — Practical Hyperparameter Selection for DtACI

    model/method

    To implement DtACI without requiring prior knowledge of the underlying distribution shift:

    1. The target local interval is set to ∣I∣=500|I| = 500, yielding the mixing parameter σ=12∣I∣=0.001\sigma = \frac{1}{2|I|} = 0.001.
    2. Under an idealized stationary benchmark where βt∼Unif(0,1)\beta_t \sim \text{Unif}(0, 1) and αt≈α\alpha_t \approx \alpha, the expected squared pinball loss evaluates to:

    1∣I∣∑t=rsE[ℓ(βt,αt)2]≈Eβ∼Unif(0,1)[ℓ(β,α)2]=(1−α)2α23.\frac{1}{|I|}\sum_{t=r}^s \mathbb{E}[\ell(\beta_t, \alpha_t)^2] \approx \mathbb{E}_{\beta \sim \text{Unif}(0, 1)}[\ell(\beta, \alpha)^2] = \frac{(1 - \alpha)^2 \alpha^2}{3}.

    1. Substituting this into the theoretical bound yields the constant heuristic choice:

    η=3500log⁡(2k⋅500)+1(1−α)α.\eta = \sqrt{\frac{3}{500}} \frac{\sqrt{\log(2k \cdot 500) + 1}}{(1 - \alpha)\alpha}.

    Alternatively, η\eta can be tracked online via:

    ηt=log⁡(2k⋅500)+1∑s=t−501tE[ℓ(βs,αs)2].\eta_t = \sqrt{\frac{\log(2k \cdot 500) + 1}{\sum_{s=t-501}^t \mathbb{E}[\ell(\beta_s, \alpha_s)^2]}}.

    Both constant and online adaptive choices achieve near-identical empirical results across real-world datasets.

  7. Knowl 7 — Dynamic Regret of DtACI with Online Time-Varying Learning Rates

    theoretical result

    Let L∈NL \in \mathbb{N} denote a fixed target interval length and I=[r,s]I = [r, s] be an interval of length LL with r>Lr > L. Suppose DtACI is configured with mixing parameter σ=1/(2L)\sigma = 1/(2L) and time-varying expert learning rates ηt:=log⁡(2Lk)+1∑s=t−L+1tE[ℓ(βs,αs)2]\eta_t := \sqrt{\frac{\log(2Lk) + 1}{\sum_{s=t-L+1}^t \mathbb{E}[\ell(\beta_s, \alpha_s)^2]}}. If the normalized variability of ηt\eta_t over II satisfies:

    1Lηs∑t=rs∣ηt−ηs∣≤O(1L)and1Lηs∑t=rs∣ηt2−ηs2∣≤O(1L),\frac{1}{L \eta_s}\sum_{t=r}^s |\eta_t - \eta_s| \le \mathcal{O}\left(\frac{1}{\sqrt{L}}\right) \quad \text{and} \quad \frac{1}{L \eta_s}\sum_{t=r}^s |\eta_t^2 - \eta_s^2| \le \mathcal{O}\left(\frac{1}{\sqrt{L}}\right),

    then under the condition γi+1/γi≤2\gamma_{i+1}/\gamma_i \le 2 and γk≥1+1/L\gamma_k \ge \sqrt{1 + 1/L}, the average dynamic regret over II satisfies:

    1L∑t=rsE[ℓ(βt,αt)]−1L∑t=rsℓ(βt,αt∗)≤O(log⁡(L)L)+O(max⁡{γ1,1L∑t=r+1s∣αt∗−αt−1∗∣}).\frac{1}{L}\sum_{t=r}^s \mathbb{E}[\ell(\beta_t, \alpha_t)] - \frac{1}{L}\sum_{t=r}^s \ell(\beta_t, \alpha_t^*) \le \mathcal{O}\left(\sqrt{\frac{\log(L)}{L}}\right) + \mathcal{O}\left(\max\left\{ \gamma_1, \sqrt{\frac{1}{L}\sum_{t=r+1}^s |\alpha_t^* - \alpha_{t-1}^*|} \right\}\right).

    This guarantees that time-varying learning rates adapt to changing loss variance without sacrificing the optimal dynamic regret rate.

  8. Knowl 8 — Adaptivity of DtACI vs. Competitors Under Synthetic Regime Shifts

    empirical result

    In synthetic simulations with target coverage 1−α=0.901 - \alpha = 0.90 and candidate step sizes γ∈{0.001,0.002,0.004,0.008,0.016,0.032,0.064,0.128}\gamma \in \{0.001, 0.002, 0.004, 0.008, 0.016, 0.032, 0.064, 0.128\}, DtACI was evaluated against Aggregated ACI (AgACI) and Multivalid Conformal Prediction (MVP) under three data-generating regimes for Yt∼N(μt,1)Y_t \sim \mathcal{N}(\mu_t, 1):

    1. Stationary regime (μt=0\mu_t = 0): AgACI and MVP slightly outperformed DtACI by converging to a single stationary threshold, while DtACI maintained small bounded fluctuations due to its non-zero minimum step size.
    2. Smooth shift regime (autoregressive continuous drift μt+1=μt+0.5(μt−μt−1)+0.5ϵt\mu_{t+1} = \mu_t + 0.5(\mu_t - \mu_{t-1}) + 0.5\epsilon_t): DtACI and AgACI performed comparably well, closely tracking the time-varying target αt∗\alpha_t^*, while MVP failed to provide local adaptivity.
    3. Jump shifts regime (alternating between small oscillations in [−0.075,0.075][-0.075, 0.075] and large swings in [−1.5,1.5][-1.5, 1.5]): AgACI adjusted slowly to large jumps due to heavy historical weighting and failed to reduce its step size when small shifts resumed. In contrast, DtACI dynamically increased its effective step size γˉt=∑iptiγi\bar{\gamma}_t = \sum_i p_t^i \gamma_i during large shifts and rapidly decreased it during calm phases, achieving significantly smaller instantaneous coverage errors ∣P(Yt∈C^t(αt)∣αt)−0.90∣|\mathbb{P}(Y_t \in \hat{C}_t(\alpha_t) \mid \alpha_t) - 0.90| throughout.
  9. Knowl 9 — Stock Market Volatility Prediction under Distribution Shift

    empirical result

    DtACI was tested on predicting next-day squared price volatility Vt=((Pt−Pt−1)/Pt−1)2V_t = ((P_t - P_{t-1})/P_{t-1})^2 across four stocks (Nvidia, AMD, BlackBerry, Fannie Mae) using a GARCH(1,1) model fit on rolling 1250-day windows to construct 90% prediction sets (target 1−α=0.901 - \alpha = 0.90). Two conformity score types were evaluated:

    1. Normalized score: St(v)=∣v−(σ^tt)2∣/(σ^tt)2S_t(v) = |v - (\hat{\sigma}_t^t)^2| / (\hat{\sigma}_t^t)^2, producing a relatively stable optimal parameter αt∗\alpha_t^*.
    2. Unnormalized score: St(v)=∣v−(σ^tt)2∣S_t(v) = |v - (\hat{\sigma}_t^t)^2|, generating large shifts in αt∗\alpha_t^* during volatility spikes.

    Key empirical findings include:

    • DtACI maintained local 500-day rolling average coverage consistently near the 0.90 target across all stocks and both score types without requiring manual step-size selection.
    • The fixed baseline (holding αt=0.10\alpha_t = 0.10) suffered severe coverage drops down to 0.40–0.60 during high-volatility events like the 2008 financial crisis on unnormalized scores.
    • MVP showed minimal adaptivity to local shifts, performing similarly to the uncalibrated fixed baseline.
    • AgACI performed similarly to DtACI in stationary shift phases but lagged behind during sudden volatility surges (e.g., Fannie Mae post-2008).
    • Q-Q plots comparing the empirical quantiles of DtACI local coverage errors to an ideal independent Bernoulli(0.10)\text{Bernoulli}(0.10) process showed near-perfect alignment.
  10. Knowl 10 — Conditional Coverage Calibration and Non-Reactivity of DtACI

    empirical result

    To verify that DtACI's coverage is achieved by learning the underlying conditional optimal parameter αt∗\alpha_t^* rather than by oscillating reactively between extreme over- and under-coverage, empirical conditional coverage was computed across intervals B1,…,BmB_1, \dots, B_m partitioning [0,1][0, 1]:

    CondCoveragei:=1∣{t:αˉt∈Bi}∣∑t:αˉt∈Bierrt.\text{CondCoverage}_i := \frac{1}{|\{t : \bar{\alpha}_t \in B_i\}|} \sum_{t : \bar{\alpha}_t \in B_i} \text{err}_t.

    Evaluated on stock volatility forecasting at nominal coverage 1−α=0.901 - \alpha = 0.90, empirical conditional coverage remained consistently close to 0.90 across all observed bins of αˉt\bar{\alpha}_t. Block bootstrap 95% confidence intervals (block size 100, 100 resamples) covered the nominal 0.90 level across all bins for both normalized and unnormalized conformity scores, demonstrating that DtACI does not suffer from pathological oscillation artifacts.

Coverage note — Omitted the secondary COVID-19 real-data results (which mirrored the conclusions of the stock volatility experiments) as well as the implementation details of MVP (Appendix A) and the block bootstrap algorithm (Appendix B) to maintain high significance and granularity across the top knowls.

References

  1. 1.Rina Foygel Barber, Emmanuel J. Cand`es, Aaditya Ramdas, and Ryan J. Tibshirani. Predictive inference with the jackknife+. The Annals of Statistics, 49(1):486 – 507, 2021. doi: 10.1214/20-AOS1965. URL https://doi.org/10.1214/20-AOS1965.
  2. 2.Rina Foygel Barber, Emmanuel J. Cand`es, Aaditya Ramdas, and Ryan J. Tibshirani. Conformal prediction beyond exchangeability. The Annals of Statistics, 51(2):816 – 845, 2023. doi: 10.1214/23-AOS2276. URL https://doi.org/10.1214/23-AOS2276.
  3. 3.Osbert Bastani, Varun Gupta, Christopher Jung, Georgy Noarov, Ramya Ramalingam, and Aaron Roth. Practical adversarial multivalid conformal prediction. In Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=QNjyrDBx6tz.
  4. 4.Victor Chernozhukov, Kaspar W¨uthrich, and Zhu Yinchu. Exact and robust conformal inference methods for predictive machine learning with dependent data. In Proceedings of the 31st Conference On Learning Theory, volume 75, pages 732–749. PMLR, 06–09 Jul 2018. URL http://proceedings.mlr.press/v75/chernozhukov18a.html.
  5. 5.Rina Foygel Barber, Emmanuel J Cand`es, Aaditya Ramdas, and Ryan J Tibshirani. The limits of distribution-free conditional predictive inference. Information and Inference: A Journal of the IMA, 08 2020. ISSN 2049-8772. doi: 10.1093/imaiai/iaaa017. URL https://doi.org/10.1093/imaiai/iaaa017. iaaa017.
  6. 6.Alexander Gammerman and Vladimir Vovk. Hedging predictions in machine learning. The Computer Journal, 50(2):151–163, 2007. doi: 10.1093/comjnl/bxl065.
  7. 7.Isaac Gibbs and Emmanuel Cand`es. Adaptive conformal inference under distribution shift. In Advances in Neural Information Processing Systems, volume 34, pages 1660–1672. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper/2021/file/0d441de75945e5acbc865406fc9a2559-Paper.pdf.
  8. 8.Paula Gradu, Elad Hazan, and Edgar Minasyan. Adaptive regret for control of time-varying dynamics. In Proceedings of The 5th Annual Learning for Dynamics and Control Conference, volume 211, pages 560–572. PMLR, 15–16 Jun 2023. URL https://proceedings.mlr.press/v211/gradu23a.html.
  9. 9.Elad Hazan. Introduction to online convex optimization. arXiv preprint, 2019. arXiv:1909.05207.
  10. 10.Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton Earnshaw, Imran Haque, Sara M Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang. Wilds: A benchmark of in-the-wild distribution shifts. In Proceedings of the 38th International Conference on Machine Learning, volume 139, pages 5637–5664. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/koh21a.html.
  11. 11.Jing Lei and Larry Wasserman. Distribution-free prediction bands for non-parametric regression. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 76, 01 2014. doi: 10.1111/rssb.12021.
  12. 12.Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. Conformal prediction with temporal quantile adjustments. In Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=PM5gVmG2Jj.
  13. 13.Harris Papadopoulos. Inductive conformal prediction: Theory and application to neural networks. In Tools in Artificial Intelligence,, pages 315–330, 2008.
  14. 14.Harris Papadopoulos, Kostas Proedrou, Volodya Vovk, and Alex Gammerman. Inductive confidence machines for regression. In Machine Learning: ECML 2002, pages 345–356, Berlin, Heidelberg, 2002. Springer Berlin Heidelberg. ISBN 978-3-540-36755-0.
  15. 15.Aleksandr Podkopaev and Aaditya Ramdas. Distribution-free uncertainty quantification for classification under label shift. arXiv preprint, 2021. arXiv:2103.03323.
  16. 16.Alex Reinhart, Logan Brooks, Maria Jahja, Aaron Rumack, Jingjing Tang, Sumit Agrawal, Wael Al Saeed, Taylor Arnold, Amartya Basu, Jacob Bien, Angel A. Cabrera, Andrew Chin, Eu Jing Chua, Brian Clark, Sarah Colquhoun, Nat DeFries, David C. Farrow, Jodi Forlizzi, Jed Grabman, Samuel Gratzl, Alden Green, George Haff, Robin Han, Kate Harwood, Addison J. Hu, Raphael Hyde, Sangwon Hyun, Ananya Joshi, Jimi Kim, Andrew Kuznetsov, Wichada La Motte-Kerr, Yeon Jin Lee, Kenneth Lee, Zachary C. Lipton, Michael X. Liu, Lester Mackey, Kathryn Mazaitis, Daniel J. McDonald, Phillip McGuinness, Balasubramanian Narasimhan, Michael P. O’Brien, Natalia L. Oliveira, Pratik Patil, Adam Perer, Collin A. Politsch, Samyak Rajanala, Dawn Rucker, Chris Scott, Nigam H. Shah, Vishnu Shankar, James Sharpnack, Dmitry Shemetov, Noah Simon, Benjamin Y. Smith, Vishakha Srivastava, Shuyi Tan, Robert Tibshirani, Elena Tuzhilina, Ana Karina Van Nortwick, Val´erie Ventura, Larry Wasserman, Benjamin Weaver, Jeremy C. Weiss, Spencer Whitman, Kristin Williams, Roni Rosenfeld, and Ryan J. Tibshirani. An open repository of real-time covid-19 indicators. Proceedings of the National Academy of Sciences, 118(51):e2111452118, 2021. doi: 10.1073/pnas.2111452118. URL https://www.pnas.org/doi/abs/10.1073/pnas.2111452118.
  17. 17.Yaniv Romano, Evan Patterson, and Emmanuel Cand`es. Conformalized quantile regression. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/5103c3584b063c431bd1268e9b5e76fb-Paper.pdf.
  18. 18.Mauricio Sadinle, Jing Lei, and Larry Wasserman. Least ambiguous set-valued classifiers with bounded error levels. Journal of the American Statistical Association, 114(525): 223–234, 2019. doi: 10.1080/01621459.2017.1395341. URL https://doi.org/10.1080/01621459.2017.1395341.
  19. 19.C. Saunders, A. Gammerman, and V. Vovk. Transduction with confidence and credibility. In Sixteenth International Joint Conference on Artificial Intelligence (IJCAI ’99) (01/01/99), pages 722–726, 1999. URL https://eprints.soton.ac.uk/258961/.
  20. 20.Glenn Shafer and Vladimir Vovk. A tutorial on conformal prediction. Journal of Machine Learning Research, 9(12):371–421, 2008. URL http://jmlr.org/papers/v9/shafer08a.html.
  21. 21.R. J. Tibshirani. Can symptoms surveys improve covid-19 forecasts? https://delphi.cmu.edu/blog/2020/09/21/can-symptoms-surveys-improve-covid-19-forecasts/, 2020. Accessed: 2022-06-17.
  22. 22.Ryan J Tibshirani, Rina Foygel Barber, Emmanuel Cand`es, and Aaditya Ramdas. Conformal prediction under covariate shift. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/8fb21ee7a2207526da55a679f0332de2-Paper.pdf.
  23. 23.V. Vovk, A. Gammerman, and C. Saunders. Machine-learning applications of algorithmic randomness. In Sixteenth International Conference on Machine Learning (ICML-1999) (01/01/99), pages 444–453, 1999. URL https://eprints.soton.ac.uk/258960/.
  24. 24.Vladimir Vovk, Alex Gammerman, and Glenn Shafer. Algorithmic Learning in a Random World. Springer-Verlag, Berlin, Heidelberg, 2005. ISBN 0387001522.
  25. 25.Volodimir G. Vovk. Aggregating strategies. In Proceedings of the Third Annual Workshop on Computational Learning Theory, COLT ’90, page 371–386, San Francisco, CA, USA, 1990. Morgan Kaufmann Publishers Inc. ISBN 1558601465.
  26. 26.Olivier Wintenberger. Optimal learning with bernstein online aggregation. Mach. Learn., 106(1):119–141, jan 2017. ISSN 0885-6125. doi: 10.1007/s10994-016-5592-6. URL https://doi.org/10.1007/s10994-016-5592-6.
  27. 27.Yachong Yang, Arun Kumar Kuchibhotla, and Eric Tchetgen Tchetgen. Doubly robust calibration of prediction sets under covariate shift. Journal of the Royal Statistical Society Series B: Statistical Methodology, page qkae009, 03 2024. ISSN 1369-7412. doi: 10.1093/jrsssb/qkae009. URL https://doi.org/10.1093/jrsssb/qkae009.
  28. 28.Margaux Zaffran, Olivier Feron, Yannig Goude, Julie Josse, and Aymeric Dieuleveut. Adaptive conformal predictions for time series. In Proceedings of the 39th International Conference on Machine Learning, volume 162, pages 25834–25866. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/zaffran22a.html.

Citation

MLA
Gibbs, I., and E. J. Candès. “Conformal Inference for Online Prediction with Arbitrary Distribution Shifts”. Journal of Machine Learning Research, vol. 25, no. 162, 2024, pp. 1–6, https://www.jmlr.org/papers/v25/22-1218.html.
APA
Gibbs, I., & Candès, E. J. (2024). Conformal Inference for Online Prediction with Arbitrary Distribution Shifts. Journal of Machine Learning Research, 25(162), 1–36. https://www.jmlr.org/papers/v25/22-1218.html
Chicago
Gibbs, I., and E. J. Candès. 2024. “Conformal Inference for Online Prediction with Arbitrary Distribution Shifts”. Journal of Machine Learning Research 25 (162): 1–36. https://www.jmlr.org/papers/v25/22-1218.html.
Harvard
Gibbs, I. and Candès, E.J. (2024) “Conformal Inference for Online Prediction with Arbitrary Distribution Shifts”, Journal of Machine Learning Research, 25(162), pp. 1–36. Available at: https://www.jmlr.org/papers/v25/22-1218.html.
Vancouver
1. Gibbs I, Candès EJ (2024) Conformal Inference for Online Prediction with Arbitrary Distribution Shifts. Journal of Machine Learning Research 25:1–36

BibTeX

@article{JMLR:v25:22-1218,
  author  = {Isaac Gibbs and Emmanuel J. Cand{{\`e}}s},
  title   = {Conformal Inference for Online Prediction with Arbitrary Distribution Shifts},
  journal = {Journal of Machine Learning Research},
  year    = {2024},
  volume  = {25},
  number  = {162},
  pages   = {1--36},
  url     = {http://jmlr.org/papers/v25/22-1218.html}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/