Adaptive Time Series Reasoning via Segment Selection

Shvat MessicaJiawen ZhangKevin LiTheodoros TsiligkaridisMarinka Zitnik

article2026arXiv2 citations

Introduces ARTIST, a reinforcement learning framework that adaptively selects task-relevant time-series segments during inference to significantly improve accuracy on complex temporal reasoning benchmarks.

Listen

Modern time-series analysis increasingly requires artificial intelligence systems to answer complex natural-language questions, such as explaining why a patient's vital signs deteriorated or identifying triggers behind financial market shifts. Existing machine learning models typically process an entire, static sequence at once. This standard approach often degrades analytical accuracy because long sequences dilute critical local patterns with irrelevant data, failing to support dynamic multi-step reasoning where intermediate deductions should guide what data to inspect next.

The article demonstrates ARTIST, a framework that models time-series reasoning as a sequential decision process by interleaving analytical reasoning with adaptive segment selection during inference. To achieve this, the system uses a single policy divided into two complementary roles: a high-level controller that iteratively identifies and retrieves informative data intervals, and a low-level reasoner that generates deductions and answers conditioned on the selected segments. The model is optimized using supervised fine-tuning on structured reasoning traces, followed by collaborative self-play reinforcement learning that applies trajectory-level reliability rewards to the controller and correctness rewards to the reasoner.

The empirical evaluation shows substantial improvements over existing techniques across multiple domains. ARTIST improved average accuracy by 6.46 absolute percentage points over the strongest competing baseline across six diverse benchmarks spanning clinical, financial, and environmental domains, achieving gains up to 12.5 percentage points on tasks requiring localized reasoning. The reinforcement learning stage boosted accuracy from 63.61% under supervised fine-tuning alone to 69.26%. Crucially, the model achieved peak accuracy while consuming only 30% to 70% of the total time-series data, and testing on extended sequences showed that performance remained stable within 1.5 percentage points even when sequence length was tripled with uninformative signal.

These findings indicate that time-series analysis models perform significantly better when they selectively acquire task-relevant data instead of processing entire sequences simultaneously. Selectively focusing on critical intervals creates an interpretable, verifiable evidence trail connecting answers directly to specific temporal events, which mitigates the risk of hallucinated or diluted insights in high-stakes domains such as healthcare and operations. Furthermore, the decoupling of segment selection from step-by-step reasoning enables efficient processing that scales gracefully to long sequences without proportional increases in computational load.

Organizations developing or deploying automated time-series reasoning tools should adopt adaptive segment-retrieval mechanisms rather than monolithic sequence-encoding architectures. Before deploying this approach in high-throughput production environments, teams should run targeted pilot tests to balance accuracy gains against the additional inference latency resulting from multi-turn interactions. Subsequent development should focus on extending the framework to multivariate datasets, irregular sampling rates, and native vision-language backbones for specialized signals such as electroencephalograms.

Confidence in these findings is supported by rigorous evaluations across multiple domain benchmarks, ablation studies, and independent test runs. However, decision-makers should note that the current implementation is restricted to univariate time series and introduces higher inference latency than single-pass models due to iterative controller-reasoner rollouts.

No sufficiently relevant recommendations were found.

Cover for Adaptive Time Series Reasoning via Segment Selection

Abstract

Time series reasoning tasks often start with a natural language question and require targeted analysis of a time series. Evidence may span the full series or appear in a few short intervals, so the model must decide what to inspect. Most existing approaches encode the entire time series into a fixed representation before inference, regardless of whether or not the entire sequence is relevant. We introduce ARTIST, which formulates time-series reasoning as a sequential decision problem. ARTIST interleaves reasoning with adaptive temporal segment selection. It adopts a controller-reasoner architecture and uses reinforcement learning to train the controller role to select informative segments and the reasoner role to generate segment-conditioned reasoning traces and final answers. During inference, the model actively acquires task-relevant information instead of relying on a static summary of the full sequence. We use a novel hierarchical policy optimization approach for post-training that allows the model to excel in both segment selection and question-answering behavior. We evaluate ARTIST on six time-series reasoning benchmarks and compare it with large language models, vision-language models, and prior time-series reasoning systems. ARTIST improves average accuracy by 6.46 absolute percentage points over the strongest baseline. The largest gains appear on rare event localization and multi-segment reasoning tasks. Supervised fine-tuning improves performance, and reinforcement learning provides additional gains by optimizing question-adaptive segment selection. These results show that selective data use drives effective time-series reasoning.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 ARTIST
  • 3.1 Problem Setup and Notation
  • 3.2 Controller-Reasoner Roles in ARTIST
  • 3.3 Interaction Trajectory Rollout
  • 3.4 Reward Functions
  • 3.5 Hierarchical Policy Optimization
  • 4 Experiments
  • 4.1 Benchmarking Results
  • 4.2 Analysis of Time Series Data Utilization
  • 4.3 Ablation Study
  • 4.4 Illustration of Time Series Reasoning in ARTIST
  • 5 Conclusion
  • References
  • A Controller and Reasoner Objectives and Rewards
  • A.1 Controller Advantages
  • A.2 Controller Objective
  • A.3 Reasoner Advantages
  • A.4 Reasoner Objective
  • A.5 Combined Objective
  • A.6 Format Reward
  • B Datasets
  • C SFT CoT Data Creation
  • D Supervised Fine-Tuning Configuration
  • E RL Training Configuration
  • E.1 Training Cost: RL vs. SFT
  • F Prompts
  • G Modality Analysis on Sleep-QA
  • H Comparison with Dynamic Visual Search
  • I Inference Cost Analysis
  • J Scalability to Long Sequences
  • K Comparison with Time-Series Explainability Methods

Knowls

  1. Knowl 1 — ARTIST Controller-Reasoner Framework for Adaptive Time-Series Reasoning

    model/method

    ARTIST (Adaptive Reasoning for TIme-Series via Temporal Selection) is a framework that frames time-series question answering as a sequential decision process where a single policy model πθ\pi_\theta dynamically selects task-relevant temporal segments rather than encoding an entire time series statically.

    Given a natural-language question qq and a univariate time series T∈RH×1T \in \mathbb{R}^{H \times 1} of length HH, the model alternates between two roles invoked via role-specific prompts:

    1. Controller Role ( πθctl\,\pi_\theta^{\text{ctl}}): Acts as a high-level policy. At interaction round ii, given state xictl=(q,T,Si−1,ai−1,y^i−1)x_i^{\text{ctl}} = (q, T, S_{i-1}, a_{i-1}, \hat{y}_{i-1}) (where S0=∅S_0 = \emptyset, a0=∅a_0 = \emptyset, and y^0=∅\hat{y}_0 = \emptyset), the controller generates a reasoning trace uiu_i, a termination decision di∈{CONTINUE,ACCEPT}d_i \in \{\text{CONTINUE}, \text{ACCEPT}\}, and, if di=CONTINUEd_i = \text{CONTINUE}, a proposed segment si=Ttstart:tends_i = T_{t_{\text{start}}:t_{\text{end}}} with 0≤tstart<tend≤H0 \le t_{\text{start}} < t_{\text{end}} \le H: (ui,di,si)∼πθctl(⋅∣xictl)(u_i, d_i, s_i) \sim \pi_\theta^{\text{ctl}}(\cdot \mid x_i^{\text{ctl}}) The segment is appended to the cumulative segment set Si=Si−1∪{si}S_i = S_{i-1} \cup \{s_i\}. If di=ACCEPTd_i = \text{ACCEPT}, the process terminates and the previous reasoner output y^i−1\hat{y}_{i-1} is returned as the final answer.

    2. Reasoner Role ( πθrsn\,\pi_\theta^{\text{rsn}}): Acts as a low-level policy. When di=CONTINUEd_i = \text{CONTINUE}, it consumes only the natural-language question qq and accumulated segments SiS_i to produce an intermediate reasoning trace aia_i and answer hypothesis y^i\hat{y}_i: (ai,y^i)∼πθrsn(⋅∣q,Si)(a_i, \hat{y}_i) \sim \pi_\theta^{\text{rsn}}(\cdot \mid q, S_i)

    A complete interaction trajectory terminated at step LL is denoted τ={(ui,si,di,ai,y^i)}i=1L\tau = \{(u_i, s_i, d_i, a_i, \hat{y}_i)\}_{i=1}^L with sub-trajectories τctl={(ui,si,di)}i=1L\tau_{\text{ctl}} = \{(u_i, s_i, d_i)\}_{i=1}^L and final reasoner trajectory τrsn,L=(aL,y^L)\tau_{\text{rsn}, L} = (a_L, \hat{y}_L).

  2. Knowl 2 — Reliability-Based and Role-Specific Reward Formulations in ARTIST

    equation

    In the ARTIST framework, training signals are tailored to the distinct roles of the controller and reasoner:

    1. Answer Correctness: Given question qq, ground-truth answer y∗y^*, and prediction y^\hat{y}, binary correctness is defined as: C(q,y∗,y^)=1[y^=y∗]C(q, y^*, \hat{y}) = \mathbf{1}[\hat{y} = y^*]

    2. Reasoner Reliability Score: For a segment set SS, reliability D(q,S,y∗)D(q, S, y^*) measures the expected accuracy of the reasoner across NN independent stochastic rollouts y^(n)∼πθrsn(⋅∣q,S)\hat{y}^{(n)} \sim \pi_\theta^{\text{rsn}}(\cdot \mid q, S): D(q,S,y∗)=1N∑n=1NC(q,y∗,y^(n))D(q, S, y^*) = \frac{1}{N} \sum_{n=1}^N C(q, y^*, \hat{y}^{(n)})

    3. Controller Format and Total Reward: Let fi∈[−1,1]f_i \in [-1, 1] be the per-step format compliance score at round ii, and let Iviol(τctl)∈{0,1}\mathbb{I}_{\text{viol}}(\tau_{\text{ctl}}) \in \{0, 1\} indicate a critical formatting violation (such as invalid JSON, missing tool calls, or proposing out-of-bound segments). The trajectory-level format score is: Fctl(τctl)=(1−Iviol(τctl))⋅(1L∑i=1Lfi)−Iviol(τctl)F_{\text{ctl}}(\tau_{\text{ctl}}) = (1 - \mathbb{I}_{\text{viol}}(\tau_{\text{ctl}})) \cdot \left(\frac{1}{L} \sum_{i=1}^L f_i\right) - \mathbb{I}_{\text{viol}}(\tau_{\text{ctl}}) The controller reward assigns a hard failure of −1-1 on critical format failure and otherwise rewards reliability and formatting: Rctl(τctl,D)={−1if Fctl(τctl)<0,wDD+wfFctl(τctl)otherwise,R_{\text{ctl}}(\tau_{\text{ctl}}, D) = \begin{cases} -1 & \text{if } F_{\text{ctl}}(\tau_{\text{ctl}}) < 0, \\ w_D D + w_f F_{\text{ctl}}(\tau_{\text{ctl}}) & \text{otherwise}, \end{cases} where wDw_D and wfw_f are weighting hyperparameters.

    4. Reasoner Reward: For final-round reasoner trajectory τrsn=(a,y^)\tau_{\text{rsn}} = (a, \hat{y}) with correctness c=C(q,y∗,y^)c = C(q, y^*, \hat{y}) and format score Frsn(τrsn)∈[−1,1]F_{\text{rsn}}(\tau_{\text{rsn}}) \in [-1, 1]: Rrsn(τrsn,c)=wcc+weFrsn(τrsn)R_{\text{rsn}}(\tau_{\text{rsn}}, c) = w_c c + w_e F_{\text{rsn}}(\tau_{\text{rsn}}) where wcw_c and wew_e are weighting hyperparameters.

  3. Knowl 3 — Hierarchical Policy Optimization Algorithm

    algorithm

    The Hierarchical Policy Optimization method in ARTIST decouples the optimization of the high-level controller and low-level reasoner via nested rollouts and variance-guided sampling, executing a joint gradient update on shared parameters θ\theta.

    For each question-series pair, GG interaction trajectories are sampled. At the terminating round L(g)L^{(g)} of each trajectory, the reasoner is resampled NN times under the final segment list SL(g)(g)S_{L^{(g)}}^{(g)}. Controller advantages are calculated across all GG trajectories to credit multi-step segment acquisition, while reasoner advantages are computed within a single trajectory group g∗g^* selected proportionally to the empirical outcome variance rσ(g)r_\sigma^{(g)}.

    Input: Initial policy parameters θinit\theta_{\text{init}}, training batch of question-series pairs (q,T,y∗)(q, T, y^*), trajectory group size GG, reasoner sample count NN
    Output: Updated policy parameters θ\theta
    θ←θinit\theta \leftarrow \theta_{\text{init}}
    for each batch do
        for g=1,2,…,Gg = 1, 2, \dots, G do
            Sample full interaction trajectory τ(g)\tau^{(g)} of length L(g)L^{(g)} using πθ\pi_\theta
            for n=1,2,…,Nn = 1, 2, \dots, N do
                Sample final reasoner rollout τrsn,L(g)(g,n)=(aL(g)(g,n),y^L(g)(g,n))∼πθrsn(⋅∣q,SL(g)(g))\tau_{\text{rsn}, L^{(g)}}^{(g,n)} = (a_{L^{(g)}}^{(g,n)}, \hat{y}_{L^{(g)}}^{(g,n)}) \sim \pi_\theta^{\text{rsn}}(\cdot \mid q, S_{L^{(g)}}^{(g)})
                Compute reasoner reward rrsn(g,n)←Rrsn(τrsn,L(g)(g,n),C(q,y∗,y^L(g)(g,n)))r_{\text{rsn}}^{(g,n)} \leftarrow R_{\text{rsn}}(\tau_{\text{rsn}, L^{(g)}}^{(g,n)}, C(q, y^*, \hat{y}_{L^{(g)}}^{(g,n)}))
            end for
            Compute mean correctness rμ(g)←1N∑n=1NC(q,y∗,y^L(g)(g,n))r_\mu^{(g)} \leftarrow \frac{1}{N} \sum_{n=1}^N C(q, y^*, \hat{y}_{L^{(g)}}^{(g,n)})
            Compute variance rσ(g)←1N∑n=1N(C(q,y∗,y^L(g)(g,n))−rμ(g))2r_\sigma^{(g)} \leftarrow \frac{1}{N} \sum_{n=1}^N (C(q, y^*, \hat{y}_{L^{(g)}}^{(g,n)}) - r_\mu^{(g)})^2
            Compute controller reward rctl(g)←Rctl(τctl(g),rμ(g))r_{\text{ctl}}^{(g)} \leftarrow R_{\text{ctl}}(\tau_{\text{ctl}}^{(g)}, r_\mu^{(g)})
        end for
        Compute controller advantages {A^ctl(g)}g=1G\{\hat{A}_{\text{ctl}}^{(g)}\}_{g=1}^G over all GG rollouts
        Sample single trajectory index g∗∼{1,…,G}g^* \sim \{1, \dots, G\} with probability p(g)∝rσ(g)p(g) \propto r_\sigma^{(g)}
        Compute reasoner advantages {A^rsn(g∗,n)}n=1N\{\hat{A}_{\text{rsn}}^{(g^*, n)}\}_{n=1}^N over the NN resampled rollouts of group g∗g^*
        Compute controller loss Jctl(θ)J_{\text{ctl}}(\theta) over all rounds i∈{1,…,L(g)}i \in \{1, \dots, L^{(g)}\} for all g∈{1,…,G}g \in \{1, \dots, G\}
        Compute reasoner loss Jrsn(θ)J_{\text{rsn}}(\theta) over final round L(g∗)L^{(g^*)} for rollouts n∈{1,…,N}n \in \{1, \dots, N\}
        Update parameters θ\theta via gradient ascent on J(θ)=Jctl(θ)+Jrsn(θ)J(\theta) = J_{\text{ctl}}(\theta) + J_{\text{rsn}}(\theta)
    end for
    return θ\theta
  4. Knowl 4 — Joint Policy Optimization Objectives and Group Advantage Formulations

    equation

    The policy parameters θ\theta of ARTIST are optimized by maximizing a combined objective J(θ)=Jctl(θ)+Jrsn(θ)J(\theta) = J_{\text{ctl}}(\theta) + J_{\text{rsn}}(\theta) using group-relative advantages:

    1. Controller Advantages and Objective: For GG interaction trajectories, the mean μctl\mu_{\text{ctl}} and standard deviation σctl\sigma_{\text{ctl}} of controller rewards {rctl(g)}g=1G\{r_{\text{ctl}}^{(g)}\}_{g=1}^G are computed, yielding group-relative advantages: A^ctl(g)={0if σctl<ϵ,rctl(g)−μctlσctl+ϵotherwise,\hat{A}_{\text{ctl}}^{(g)} = \begin{cases} 0 & \text{if } \sigma_{\text{ctl}} < \epsilon, \\ \frac{r_{\text{ctl}}^{(g)} - \mu_{\text{ctl}}}{\sigma_{\text{ctl}} + \epsilon} & \text{otherwise}, \end{cases} where ϵ>0\epsilon > 0 is a stability constant. For a batch of BB questions, where trajectory gg of question bb has L(g)L^{(g)} rounds with Ti(g)T_i^{(g)} controller output tokens octl,i,t(g)o_{\text{ctl}, i, t}^{(g)} at round ii, the controller objective is: Jctl(θ)=1B∑b=1B1G∑g=1G1L(g)∑i=1L(g)1Ti(g)∑t=1Ti(g)A^ctl(g)log⁡πθ(octl,i,t(g) | qb,{octl,i,j(g)}j=1t−1)J_{\text{ctl}}(\theta) = \frac{1}{B} \sum_{b=1}^B \frac{1}{G} \sum_{g=1}^G \frac{1}{L^{(g)}} \sum_{i=1}^{L^{(g)}} \frac{1}{T_i^{(g)}} \sum_{t=1}^{T_i^{(g)}} \hat{A}_{\text{ctl}}^{(g)} \log \pi_\theta\left(o_{\text{ctl}, i, t}^{(g)} \,\middle|\, q_b, \{o_{\text{ctl}, i, j}^{(g)}\}_{j=1}^{t-1}\right)

    2. Reasoner Advantages and Objective: For the chosen trajectory group g∗g^*, mean μrsn\mu_{\text{rsn}} and standard deviation σrsn\sigma_{\text{rsn}} of reasoner rewards {rrsn(g∗,n)}n=1N\{r_{\text{rsn}}^{(g^*, n)}\}_{n=1}^N define: A^rsn(g∗,n)={0if σrsn<ϵ,rrsn(g∗,n)−μrsnσrsn+ϵotherwise.\hat{A}_{\text{rsn}}^{(g^*, n)} = \begin{cases} 0 & \text{if } \sigma_{\text{rsn}} < \epsilon, \\ \frac{r_{\text{rsn}}^{(g^*, n)} - \mu_{\text{rsn}}}{\sigma_{\text{rsn}} + \epsilon} & \text{otherwise}. \end{cases} Optimizing only over the T(n)T^{(n)} tokens orsn,t(n)o_{\text{rsn}, t}^{(n)} of the final round L(g∗)L^{(g^*)} with a reference SFT policy πref\pi_{\text{ref}} gives: Jrsn(θ)=1B∑b=1B1N∑n=1N1T(n)∑t=1T(n)A^rsn(g∗,n)log⁡πθ(orsn,t(n) | qb,{orsn,j(n)}j=1t−1)−DKL(πθ∥πref)J_{\text{rsn}}(\theta) = \frac{1}{B} \sum_{b=1}^B \frac{1}{N} \sum_{n=1}^N \frac{1}{T^{(n)}} \sum_{t=1}^{T^{(n)}} \hat{A}_{\text{rsn}}^{(g^*, n)} \log \pi_\theta\left(o_{\text{rsn}, t}^{(n)} \,\middle|\, q_b, \{o_{\text{rsn}, j}^{(n)}\}_{j=1}^{t-1}\right) - D_{\text{KL}}(\pi_\theta \parallel \pi_{\text{ref}})

  5. Knowl 5 — Time-Series Reasoning Benchmark Accuracy and F1 Evaluation

    data/table

    ARTIST was evaluated across six time-series reasoning benchmarks: Etiological Reasoning (ETI), Right Whale Call Detection (RCW), ECG-QA, Sleep-QA, TSQA, and TRQA. Performance was compared against general LLMs with serialized inputs (GPT-5, LLaMA-3-8B, Qwen3-8B, Qwen3-14B) and fine-tuned encoder-LLM models (ChatTS-14B, OpenTSLM-4B, ITFormer-4B).

    Model ETI RCW ECG QA SLEEP QA TSQA TRQA Avg.
    Acc F1 Acc F1 Acc F1 Acc F1 Acc F1 Acc F1 Acc F1
    Random Guess 25.00 25.00 50.00 50.00 50.00 50.00 16.67 16.67 29.67 29.67 37.13 37.13 34.74 34.74
    GPT-5 w/ stats 63.54 63.72 32.74 31.08 53.96 49.62 0.49 1.49 35.75 32.24 25.00 28.70 35.25 34.48
    Llama-3 8B w/ stats 35.50 34.39 62.83 45.82 50.00 48.90 22.06 13.34 44.93 44.12 55.50 53.02 45.14 39.93
    Qwen3-14B w/ stats 42.00 47.34 22.57 22.91 54.58 51.79 3.43 4.48 24.15 22.20 29.00 30.32 29.29 29.84
    ChatTS-14B + SFT 50.50 40.69 73.89 33.10 53.47 24.31 26.47 14.60 46.38 42.65 69.00 55.57 53.28 35.15
    OpenTSLM-4B + SFT 82.69 82.66 65.49 38.29 69.50 41.00 35.37 18.99 47.50 35.81 76.25 69.36 62.80 47.68
    ITFormer-4B + SFT 84.62 84.60 67.31 57.95 57.31 49.91 33.62 15.77 49.50 23.62 80.12 74.22 62.08 51.01
    ARTIST
    + SFT 85.12 85.11 69.75 61.46 56.31 55.68 28.13 17.94 60.06 57.13 82.26 62.32 63.61 56.61
    + SFT + RL 87.03 87.10 77.00 50.00 69.81 52.67 36.63 19.21 62.00 58.66 83.06 78.02 69.26 57.61
    Improvement +2.41 +2.50 +3.11 +3.51 +3.14 +3.89 +1.26 +0.22 +12.50 +11.91 +2.94 +3.80 +6.46 +6.60

    ARTIST with SFT and RL achieves the highest average accuracy (69.26%) and F1 (57.61%), outperforming the strongest baseline per dataset by an average of +6.46 percentage points in accuracy and +6.60 in F1. The largest single gain is on TSQA (+12.50 percentage points in accuracy).

  6. Knowl 6 — Ablation Study of ARTIST Algorithmic Components

    data/table

    An ablation study on ECG-QA and RCW benchmarks isolates the contribution of each core algorithmic component of ARTIST:

    Method / Configuration ECG-QA Acc (%) RCW Acc (%) Average Acc (%)
    ARTIST (Full Model) 69.81 77.00 73.41
    Reasoner Only (Static full sequence) 65.33 62.88 64.11
    Controller-only RL (Frozen Reasoner) 60.81 68.13 64.47
    w/o Reliability Reward (Single-rollout) 52.50 51.44 51.97
    w/o Trajectory-based Objective (Myopic) 55.19 67.06 61.13
    w/o Variance-guided Sampling (Stochastic) 68.13 72.75 70.44

    Key takeaways from the ablations:

    1. Removing Adaptive Selection (Reasoner Only) decreases average accuracy by 9.30%, demonstrating that attending to the entire series introduces confounding noise.
    2. Freezing Reasoner during RL (Controller-only RL) drops performance by 8.94% due to distribution shift between SFT and dynamic RL segments.
    3. Removing Reliability Reward causes the largest drop of 21.44%, showing that multi-sample consistency is critical to prevent stochastic single-rollout rewards from corrupting segment selection.
    4. Myopic Step Objective (w/o Trajectory Objective) drops accuracy by 12.28%, confirming that credit assignment across all interaction steps is necessary for multi-round localization.
    5. Standard Sampling (w/o Variance-Guided Sampling) causes a 2.97% drop.
  7. Knowl 7 — Non-Monotonic Relationship Between Sequence Utilization and Reasoning Accuracy

    empirical result

    In ARTIST, reasoning performance does not increase monotonically with the fraction of the time series inspected. Instead, peak accuracy occurs when the model adaptively selects a targeted sub-sequence:

    • On long sequences (Sleep-QA and TRQA), accuracy peaks when the model retrieves between 30% and 50% of the time series. Questions that prompt 90% to 100% sequence utilization exhibit significantly lower reasoning accuracy.
    • On shorter sequences (TSQA), the optimal utilization range shifts higher to approximately 70%, reflecting the need for broader proportional coverage over concise series.

    This behavior demonstrates that full-sequence consumption dilutes salient local evidence with irrelevant temporal variations.

  8. Knowl 8 — Computational and Performance Invariance to Extended Sequence Lengths

    data/table

    The scalability of ARTIST to inputs longer than those seen during training was evaluated on RCW by extending length from 4,000 to 8,000 and 12,000 timesteps via noisy repeated tiling. The task-relevant information remained solely within [0,4000)[0, 4000). ARTIST was not retrained on the extended lengths.

    Sequence Length (HH) Accuracy (%) F1 Score (%) Inference Time / example (8 runs, min)
    4,000 77.00 50.00 1.880
    8,000 75.66 51.21 1.895
    12,000 76.11 48.37 1.910

    Across a 3×\times increase in sequence length (from 4K to 12K):

    • Accuracy changes by less than 1.5 percentage points (77.00% to 76.11%).
    • Per-example inference time increases by only 1.6% (1.880 min to 1.910 min across 8 runs).

    Because the reasoner conditions strictly on selected segments and the number of question-relevant regions is determined by the query rather than sequence length HH, computational cost is decoupled from total sequence length.

  9. Knowl 9 — Input Modality Effect in Biomedical Time Series Reasoning

    data/table

    To analyze why vision-language models (VLMs) like TimeMaster performed well on Sleep-QA (EEG signals) compared to tokenized time-series models, an ablation replaced the default Qwen3-4B backbone of ARTIST with a vision-language backbone (Qwen2.5-VL-3B) that receives rendered plots, while preserving the controller-reasoner segment selection mechanism:

    Method Accuracy (%) F1 Score (%)
    ARTIST (VLM backbone, SFT) 65.00 39.22
    TimeMaster 47.51 32.98
    ARTIST (Tokenized backbone, SFT) 28.13 17.94

    Under identical SFT training, the VLM-instantiated ARTIST improves accuracy from 28.13% (tokenized) to 65.00%, surpassing TimeMaster (47.51%). This confirms that the baseline performance disparity on Sleep-QA is driven by visual versus tokenized input modality representations rather than a deficiency in adaptive segment selection.

  10. Knowl 10 — Alignment Between Adaptive Segment Selection and Time-Series Explainability Ground Truth

    data/table

    The temporal segments selected by ARTIST's controller during question answering were evaluated against ground-truth salient regions on synthetic XAI benchmarks (FreqShape and SeqCombSingle), and compared against continuous attribution methods (TimeX and TimeX++):

    FreqShape Benchmark SHR ↑\uparrow GTC ↑\uparrow Redundancy ↓\downarrow SER ↑\uparrow
    ARTIST 0.7340 ±\pm 0.0205 0.6804 ±\pm 0.0189 0.0138 ±\pm 0.0062 0.6250 ±\pm 0.0342
    TimeX 0.6252 ±\pm 0.0065 0.5600 ±\pm 0.0066 0.0826 ±\pm 0.0198 0.6100 ±\pm 0.0141
    TimeX++ 0.6430 ±\pm 0.0092 0.5933 ±\pm 0.0057 0.0601 ±\pm 0.0217 0.6880 ±\pm 0.0259
    SeqCombSingle Benchmark
    ARTIST 0.4900 ±\pm 0.0137 0.3406 ±\pm 0.0134 0.0067 ±\pm 0.0066 0.6467 ±\pm 0.0172
    TimeX 0.7095 ±\pm 0.0149 0.1793 ±\pm 0.0225 0.0274 ±\pm 0.0218 0.6690 ±\pm 0.0297
    TimeX++ 0.7340 ±\pm 0.0034 0.2006 ±\pm 0.0238 0.0205 ±\pm 0.0149 0.7180 ±\pm 0.0067

    Metrics include Segment Hit Rate (SHR), Ground-Truth Coverage (GTC), Redundancy (fraction of non-overlapping selected segments), and Sufficient Evidence Rate (SER). ARTIST achieves substantially lower Redundancy (0.0138 on FreqShape, 0.0067 on SeqCombSingle) and higher GTC (0.6804 and 0.3406), producing tight, highly focused segment selections compared to broader continuous XAI attributions.

  11. Knowl 11 — Inference Latency Overhead and Univariate Limitations of ARTIST

    limitation

    The design of ARTIST introduces specific operational limitations:

    1. Inference Latency: Because ARTIST requires multi-turn Controller-Reasoner interactions per query rather than a single forward pass, inference wall-clock time is higher than single-pass models. For example, on TRQA, ARTIST takes 1.68 minutes per example (across 8 evaluation runs) compared to 1.26 minutes for OpenTSLM-4B and 1.29 minutes for ITFormer-4B; on ECG-QA, ARTIST takes 2.00 minutes versus 0.78 minutes for OpenTSLM-4B.
    2. Univariate Setting: The current formulation and evaluation are restricted to univariate time series (V=1V=1) with regular sampling intervals. Extension to multivariate systems and irregular sampling dynamics remains unaddressed.

Coverage note — None was omitted.

References

  1. 1.Alnegheimish, S., Nguyen, L., Berti-Equille, L., and Veeramachaneni, K. Large language models can be zero-shot anomaly detectors for time series? arXiv preprint arXiv:2405.14755, 2024.
  2. 2.Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923, 2025.
  3. 3.Cai, J., Xie, Y., Lim, G., Yin, Y., Zimmermann, R., and Ng, S.-K. Self-perturbed anomaly-aware graph dynamics for multivariate time-series anomaly detection. Advances in Neural Information Processing Systems, 38: 84785–84805, 2026.
  4. 4.Chen, L., Prabhudesai, M., Fragkiadaki, K., Liu, H., and Pathak, D. Self-questioning language models. arXiv preprint arXiv:2508.03682, 2025.
  5. 5.Chow, W., Gardiner, L., Hallgr'ımsson, H. T., Xu, M. A., and Ren, S. Y. Towards time series reasoning with llms. arXiv preprint arXiv:2409.11376, 2024b.
  6. 6.Das, A., Kong, W., Sen, R., and Zhou, Y. A decoder-only foundation model for time-series forecasting. In Forty-first International Conference on Machine Learning, 2024.
  7. 7.Fang, W., Liu, S., Zhou, Y., Zhang, K., Zheng, T., Chen, K., Song, M., and Tao, D. Serl: Self-play reinforcement learning for large language models with limited data. arXiv preprint arXiv:2505.20347, 2025.
  8. 8.Grattafiori, A. et al. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783.
  9. 9.He, H., Yi, K., Ma, Y., Zhang, Q., Niu, Z., and Pang, G. Sempo: Lightweight foundation models for time series forecasting. arXiv preprint arXiv:2510.19710, 2025a.
  10. 10.He, Y., Huang, C., Li, Z., Huang, J., and Yang, Y. Visplay: Self-evolving vision-language models from images. arXiv preprint arXiv:2511.15661, 2025b.
  11. 11.Huang, C., Yu, W., Wang, X., Zhang, H., Li, Z., Li, R., Huang, J., Mi, H., and Yu, D. R-zero: Self-evolving reasoning llm from zero data. arXiv preprint arXiv:2508.05004, 2025.
  12. 12.Jing, B., Chen, S., Zheng, L., Liu, B., Li, Z., Zou, J., Wei, T., Liu, Z., Zeng, Z., Qiu, R., et al. Trqa: Time series reasoning question and answering benchmark.
  13. 13.Kim, H., Mok, J., Lee, D., Lew, J., Kim, S., and Yoon, S. Causality-aware contrastive learning for robust multivariate time-series anomaly detection. In International Conference on Machine Learning, 2025.
  14. 14.Klissarov, M., Bagaria, A., Luo, Z., Konidaris, G., Precup, D., and Machado, M. C. Discovering temporal structure: An overview of hierarchical reinforcement learning. arXiv preprint arXiv:2506.14045, 2025.
  15. 15.Kong, Y., Yang, Y., Hwang, Y., Du, W., Zohren, S., Wang, Z., Jin, M., and Wen, Q. Time-mqa: Time series multi-task question answering with context enhancement. arXiv preprint arXiv:2503.01875, 2025a.
  16. 16.Kong, Y., Yang, Y., Wang, S., Liu, C., Liang, Y., Jin, M., Zohren, S., Pei, D., Liu, Y., and Wen, Q. Position: Empowering time series reasoning with multimodal llms. arXiv preprint arXiv:2502.01477, 2025b.
  17. 17.Langer, P., Kaar, T., Rosenblattl, M., Xu, M. A., Chow, W., Maritsch, M., Verma, A., Han, B., Kim, D. S., Chubb, H., et al. Opentslm: Time-series language models for reasoning over multivariate medical text-and time-series data. arXiv preprint arXiv:2510.02410, 2025.
  18. 18.Lei, P., Song, J., Hao, Y., Chen, T., Zhang, Y., JIA, L., Li, Y., et al. Itformer: Bridging time series and natural language for multi-modal qa with large-scale multi-task dataset. In Forty-second International Conference on Machine Learning, 2025.
  19. 19.Li, C., Wu, W., Zhang, H., Xia, Y., Mao, S., Dong, L., Vulić, I., and Wei, F. Imagine while reasoning in space: Multimodal visualization-of-thought. arXiv preprint arXiv:2501.07542, 2025.
  20. 20.Liu, B., Jin, C., Kim, S., Yuan, W., Zhao, W., Kulikov, I., Li, X., Sukhbaatar, S., Lanchantin, J., and Weston, J. Spice: Self-play in corpus environments improves reasoning. arXiv preprint arXiv:2510.24684, 2025a.
  21. 21.Liu, H., Liu, C., and Prakash, B. A. A picture is worth a thousand numbers: Enabling llms reason about time series via visualization. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 7486–7518, 2025b.
  22. 22.Liu, S., Wei, C., Zhou, X., and Chen, H. Spectral-aware reservoir computing for fast and accurate time series classification. In International Conference on Machine Learning, 2025c.
  23. 23.Liu, Z., Dong, Y., Rao, Y., Zhou, J., and Lu, J. Chain-of-spot: Interactive reasoning improves large vision-language models, 2024a. URL https://arxiv.org/abs/2403.12966.
  24. 24.Liu, Z., Wang, T., Shi, J., Zheng, X., Chen, Z., Song, L., Dong, W., Obeysekera, J., Shirani, F., and Luo, D. Timex++: Learning time-series explanations with information bottleneck, 2024b. URL https://arxiv.org/abs/2405.09308.
  25. 25.Liu, Z., Luo, Y., Li, B., Eldele, E., Wu, M., and Ma, Q. Learning soft sparse shapes for efficient time-series classification. In Proceedings of the 42nd International Conference on Machine Learning, pp. 39032–39059, 2025d.
  26. 26.Liu, Z., Ni, J., Tang, X., Lau, M. S., Yin, W., and Jin, W. Can large language models adequately perform symbolic reasoning over time series? arXiv preprint arXiv:2508.03963, 2025e.
  27. 27.Luo, Y., Zhou, Y., Cheng, M., Wang, J., Wang, D., Pan, T., and Zhang, J. Time series forecasting as reasoning: A slow-thinking approach with reinforced llms. arXiv preprint arXiv:2506.10630, 2025.
  28. 28.Merrill, M. A., Tan, M., Gupta, V., Hartvigsen, T., and Althoff, T. Language models still struggle to zero-shot reason about time series. In Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 3512–3533, 2024.
  29. 29.Oh, J., Lee, G., Bae, S., Kwon, J.-m., and Choi, E. Ecg-qa: A comprehensive question answering dataset combined with electrocardiogram. Advances in Neural Information Processing Systems, 36:66277–66288, 2023.
  30. 30.Pouliou, A., Papageorgiou, V. E., Petmezas, G., Pessoa, D., Paiva, R. P., Maglaveras, N., and Tsaklidis, G. A new approach for sleep stage identification combining hidden markov models and eeg signal processing. Journal of Medical and Biological Engineering, pp. 1–12, 2025.
  31. 31.Queen, O., Hartvigsen, T., Koker, T., He, H., Tsiligkaridis, T., and Zitnik, M. Encoding time-series explanations through self-supervised model behavior consistency, 2023. URL https://arxiv.org/abs/2306.02109.
  32. 32.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  33. 33.Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.
  34. 34.Shen, H., Zhao, K., Zhao, T., Xu, R., Zhang, Z., Zhu, M., and Yin, J. Zoomeye: Enhancing multimodal llms with human-like zooming capabilities through tree-based image exploration, 2025. URL https://arxiv.org/abs/2411.16044.
  35. 35.Singh, A., Fry, A., Perelman, A., Tart, A., Ganesh, A., El-Kishky, A., McLaughlin, A., Low, A., Ostrow, A., Ananthram, A., et al. Openai gpt-5 system card. arXiv preprint arXiv:2601.03267, 2025.
  36. 36.Su, A., Wang, H., Ren, W., Lin, F., and Chen, W. Pixel reasoner: Incentivizing pixel space reasoning via curiosity-driven reinforcement learning. Advances in Neural Information Processing Systems, 38:8222–8251, 2026.
  37. 37.Sun, X., Zhou, H., and Li, C. Multivariate time series anomaly detection with idempotent reconstruction. Advances in Neural Information Processing Systems, 38: 160900–160958, 2026.
  38. 38.Tran, V.-H., Doan, N. P., Zhang, Z., Pham, T., Nguyen, P. H., Nguyen, X., Vandierendonck, H., Assent, I., and Mai, T. S. Mix: A multi-view time-frequency interactive explanation framework for time series classification. Advances in Neural Information Processing Systems, 38: 49181–49226, 2026.
  39. 39.Wang, H., Yang, Y., Hu, J., Zhu, M., and Chen, W. V-zero: Self-improving multimodal reasoning with zero annotation. arXiv preprint arXiv:2601.10094, 2026.
  40. 40.Wang, Q., Liu, B., Zhou, T., Shi, J., Lin, Y., Chen, Y., Li, H. H., Wan, K., and Zhao, W. Vision-zero: Scalable vlm self-improvement via strategic gamified self-play. arXiv preprint arXiv:2509.25541, 2025a.
  41. 41.Wang, Y., Qiu, Y., Chen, P., Zhao, K., Shu, Y., Rao, Z., Pan, L., Yang, B., and Guo, C. Towards a general time series forecasting model with unified representation and adaptive transfer. In Proceedings of the 42nd International Conference on Machine Learning, volume 267, pp. 64127–64151, 2025b.
  42. 42.Wu, Y., Wang, Y., Tang, S., Wu, W., He, T., Ouyang, W., Torr, P., and Wu, J. Dettoolchain: A new prompting paradigm to unleash detection ability of mllm, 2024. URL https://arxiv.org/abs/2403.12488.
  43. 43.Xie, Z., Li, Z., He, X., Xu, L., Wen, X., Zhang, T., Chen, J., Shi, R., and Pei, D. Chatts: Aligning time series with llms via synthetic data for enhanced understanding and reasoning. arXiv preprint arXiv:2412.03104, 2024.
  44. 44.Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025a.
  45. 45.Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023. URL https://arxiv.org/abs/2310.11441.
  46. 46.Yang, Y., Zhang, D., Liang, Y., Lu, H., Chen, G., and Li, H. Not all data are good labels: On the self-supervised labeling for time series forecasting. arXiv preprint arXiv:2502.14704, 2025b.
  47. 47.Yang, Z., Shen, W., Li, C., Chen, R., Wan, F., Yan, M., Quan, X., and Huang, F. Spell: Self-play reinforcement learning for evolving long-context language models. arXiv preprint arXiv:2509.23863, 2025c.
  48. 48.Yu, Q., Zhang, Z., Zhu, R., Yuan, Y., Zuo, X., Yue, Y., Dai, W., Fan, T., Liu, G., Liu, L., et al. Dapo: An open-source llm reinforcement learning system at scale. arXiv preprint arXiv:2503.14476, 2025a.
  49. 49.Yu, W., Liang, Z., Huang, C., Panaganti, K., Fang, T., Mi, H., and Yu, D. Guided self-evolving llms with minimal human supervision. arXiv preprint arXiv:2512.02472, 2025b.
  50. 50.Yue, Z., Upasani, K., Yang, X., Ge, S., Nie, S., Mao, Y., Liu, Z., and Wang, D. Dr. zero: Self-evolving search agents without training data. arXiv preprint arXiv:2601.07055, 2026.
  51. 51.Zhang, J., Feng, L., Guo, X., Wu, Y., Dong, Y., and Xu, D. Timemaster: Training time-series multimodal llms to reason via reinforcement learning. arXiv preprint arXiv:2506.13705, 2025a.
  52. 52.Zhang, R., Xu, Z., Ma, C., Yu, C., Tu, W.-W., Tang, W., Huang, S., Ye, D., Ding, W., Yang, Y., et al. A survey on self-play methods in reinforcement learning. arXiv preprint arXiv:2408.01072, 2025b.
  53. 53.Zhang, X., Gao, Z., Zhang, B., Li, P., Zhang, X., Liu, Y., Yuan, T., Wu, Y., Jia, Y., Zhu, S.-C., et al. Chain-of-focus: Adaptive visual search and zooming for multimodal reasoning via rl. arXiv preprint arXiv:2505.15436, 2025c.
  54. 54.Zhao, A., Wu, Y., Yue, Y., Wu, T., Xu, Q., Lin, M., Wang, S., Wu, Q., Zheng, Z., and Huang, G. Absolute zero: Reinforced self-play reasoning with zero data, 2025. URL https://arxiv. org/abs/2505.03335, 2025.
  55. 55.Zheng, Z., Yang, M., Hong, J., Zhao, C., Xu, G., Yang, L., Shen, C., and Yu, X. Deepeyes: Incentivizing ”thinking with images” via reinforcement learning, 2026. URL https://arxiv.org/abs/2505.14362.

Citation

MLA
Messica, S., et al. “Adaptive Time Series Reasoning via Segment Selection”. arXiv, 2026, http://arxiv.org/abs/2602.18645v3.
APA
Messica, S., Zhang, J., Li, K., Tsiligkaridis, T., & Zitnik, M. (2026). Adaptive Time Series Reasoning via Segment Selection. arXiv. http://arxiv.org/abs/2602.18645v3
Chicago
Messica, S., J. Zhang, K. Li, T. Tsiligkaridis, and M. Zitnik. 2026. “Adaptive Time Series Reasoning via Segment Selection”. arXiv. http://arxiv.org/abs/2602.18645v3.
Harvard
Messica, S. et al. (2026) “Adaptive Time Series Reasoning via Segment Selection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2602.18645v3.
Vancouver
1. Messica S, Zhang J, Li K, Tsiligkaridis T, Zitnik M (2026) Adaptive Time Series Reasoning via Segment Selection. arXiv

BibTeX

@article{messica2026adaptive,
  title = {Adaptive Time Series Reasoning via Segment Selection},
  author = {Messica, Shvat and Zhang, Jiawen and Li, Kevin and Tsiligkaridis, Theodoros and Zitnik, Marinka},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2602.18645v3},
  eprint = {2602.18645}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/