TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
Jingang QuDavid HolzmllerGal VaroquauxMarine Le Morvan
Introduces TabICL, a tabular foundation model featuring a two-stage attention architecture that scales in-context learning to 500,000 samples, running up to ten times faster than TabPFNv2 while outperforming CatBoost on large classification datasets.
Tabular data underpins core business and operational systems across industries such as finance and healthcare, where tree-based algorithms like gradient-boosted decision trees have long been the standard modeling approach. Recent advances have introduced tabular foundation models that perform in-context learning, generating predictions in a single pass without model retraining. However, existing tabular foundation models such as TabPFNv2 suffer from severe computational bottlenecks due to alternating row and column attention mechanisms, rendering them impractical for datasets with tens or hundreds of thousands of samples. The article evaluates whether in-context learning can be effectively scaled to larger tabular datasets and introduces TabICL, a scalable tabular foundation model designed specifically for classification tasks on large data.
The developers created a novel two-stage architecture that decouples feature embedding from sample-level reasoning. TabICL first uses a shareable Set Transformer to build distribution-aware representations of each column, followed by a row-wise transformer equipped with rotary positional embeddings to aggregate features into compact row vectors while preventing representation collapse. A final transformer performs in-context learning solely over these condensed row vectors and their associated labels. The model was pretrained purely on synthetic datasets generated via causal and tree-based structures, using a three-stage curriculum learning schedule that progressively scaled training set sizes up to 60,000 samples. The architecture was evaluated across 200 real-world classification datasets from the TALENT benchmark against more than 30 deep learning and tree-based baselines.
The evaluation demonstrates that TabICL matches the predictive accuracy of TabPFNv2 on small and medium datasets while running systematically faster by 1.5 to 10 times. On large datasets containing more than 10,000 samples, TabICL outperforms both TabPFNv2 and traditional gradient boosting models such as CatBoost, achieving a top average rank in log-loss and area under the curve metrics. Across the entire benchmark, TabICL achieved the highest median relative accuracy improvement over baseline neural networks while requiring an average fit-and-predict time of only 1.1 seconds per 1,000 samples on a single graphics processing unit. In comparison, tuning standard models like CatBoost or deep neural networks required 3 to 7 minutes per 1,000 samples. Furthermore, by pairing TabICL with a hierarchical classification strategy, the model successfully scaled to problems with more than 10 classes without retraining the underlying feature representations.
These findings indicate that tabular foundation models can eliminate the need for costly hyperparameter tuning pipelines in enterprise workflows, drastically cutting down development timelines and computational expenses while improving probability calibration. Organizations should consider pilot deployments of TabICL for rapid tabular classification tasks, particularly where fast turnarounds or real-time inference are required. Decision-makers should note that the model is currently restricted to classification problems, requires ensembling over column permutations to restore ordering invariance, and experiences slow inference per sample compared to lightweight shallow models. Nevertheless, confidence in the reported performance is high across standard tabular benchmarks, and integrating hybrid approaches like decision-tree partitioning can further scale TabICL to massive datasets.
- Paper: Accurate predictions on small data with a tabular foundation model, Noah Hollmann et al. (2025). TabICL scales the same tabular in-context-learning approach introduced by TabPFN, so this paper clarifies the foundation and small-data setting that motivate TabICL’s larger-data design.
No sufficiently relevant recommendations were found.
