Adaptive Time Series Reasoning via Segment Selection
Shvat MessicaJiawen ZhangKevin LiTheodoros TsiligkaridisMarinka Zitnik
Introduces ARTIST, a reinforcement learning framework that adaptively selects task-relevant time-series segments during inference to significantly improve accuracy on complex temporal reasoning benchmarks.
Modern time-series analysis increasingly requires artificial intelligence systems to answer complex natural-language questions, such as explaining why a patient's vital signs deteriorated or identifying triggers behind financial market shifts. Existing machine learning models typically process an entire, static sequence at once. This standard approach often degrades analytical accuracy because long sequences dilute critical local patterns with irrelevant data, failing to support dynamic multi-step reasoning where intermediate deductions should guide what data to inspect next.
The article demonstrates ARTIST, a framework that models time-series reasoning as a sequential decision process by interleaving analytical reasoning with adaptive segment selection during inference. To achieve this, the system uses a single policy divided into two complementary roles: a high-level controller that iteratively identifies and retrieves informative data intervals, and a low-level reasoner that generates deductions and answers conditioned on the selected segments. The model is optimized using supervised fine-tuning on structured reasoning traces, followed by collaborative self-play reinforcement learning that applies trajectory-level reliability rewards to the controller and correctness rewards to the reasoner.
The empirical evaluation shows substantial improvements over existing techniques across multiple domains. ARTIST improved average accuracy by 6.46 absolute percentage points over the strongest competing baseline across six diverse benchmarks spanning clinical, financial, and environmental domains, achieving gains up to 12.5 percentage points on tasks requiring localized reasoning. The reinforcement learning stage boosted accuracy from 63.61% under supervised fine-tuning alone to 69.26%. Crucially, the model achieved peak accuracy while consuming only 30% to 70% of the total time-series data, and testing on extended sequences showed that performance remained stable within 1.5 percentage points even when sequence length was tripled with uninformative signal.
These findings indicate that time-series analysis models perform significantly better when they selectively acquire task-relevant data instead of processing entire sequences simultaneously. Selectively focusing on critical intervals creates an interpretable, verifiable evidence trail connecting answers directly to specific temporal events, which mitigates the risk of hallucinated or diluted insights in high-stakes domains such as healthcare and operations. Furthermore, the decoupling of segment selection from step-by-step reasoning enables efficient processing that scales gracefully to long sequences without proportional increases in computational load.
Organizations developing or deploying automated time-series reasoning tools should adopt adaptive segment-retrieval mechanisms rather than monolithic sequence-encoding architectures. Before deploying this approach in high-throughput production environments, teams should run targeted pilot tests to balance accuracy gains against the additional inference latency resulting from multi-turn interactions. Subsequent development should focus on extending the framework to multivariate datasets, irregular sampling rates, and native vision-language backbones for specialized signals such as electroencephalograms.
Confidence in these findings is supported by rigorous evaluations across multiple domain benchmarks, ablation studies, and independent test runs. However, decision-makers should note that the current implementation is restricted to univariate time series and introduces higher inference latency than single-pass models due to iterative controller-reasoner rollouts.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). Introduces the reinforcement learning framework for interleaving internal reasoning with dynamic information acquisition actions, establishing the direct algorithmic blueprint adapted by ARTIST for temporal segment retrieval.
- Paper: Position: What Can Large Language Models Tell Us about Time Series Analysis, Ming Jin et al. (2024). Frames the foundational shift toward treating large language models as interactive, reasoning-driven agents for time series analysis rather than static numeric predictors.
- Paper: SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training, Tianzhe Chu et al. (2025). Demonstrates why combining supervised fine-tuning with reinforcement learning is essential for out-of-distribution reasoning generalization, providing the theoretical motivation behind ARTIST's two-stage post-training pipeline.
- Paper: TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning, Xiangyu Zeng 0004 et al. (2025). Demonstrates the necessity of temporal grounding and active segment localization to reduce hallucinations and improve reasoning across long sequential contexts.
- Paper: Decision Transformer: Reinforcement Learning via Sequence Modeling, Lili Chen et al. (2021). Provides the foundational paradigm of formulating sequential decision-making and policy optimization within autoregressive sequence models.
No sufficiently relevant recommendations were found.
