TabICL: A Tabular Foundation Model for In-Context Learning on Large Data

Jingang QuDavid HolzmllerGal VaroquauxMarine Le Morvan

article2025ICML114 citationsBest Paper Award of ECML-PKDD 2025

Introduces TabICL, a tabular foundation model featuring a two-stage attention architecture that scales in-context learning to 500,000 samples, running up to ten times faster than TabPFNv2 while outperforming CatBoost on large classification datasets.

Listen

Tabular data underpins core business and operational systems across industries such as finance and healthcare, where tree-based algorithms like gradient-boosted decision trees have long been the standard modeling approach. Recent advances have introduced tabular foundation models that perform in-context learning, generating predictions in a single pass without model retraining. However, existing tabular foundation models such as TabPFNv2 suffer from severe computational bottlenecks due to alternating row and column attention mechanisms, rendering them impractical for datasets with tens or hundreds of thousands of samples. The article evaluates whether in-context learning can be effectively scaled to larger tabular datasets and introduces TabICL, a scalable tabular foundation model designed specifically for classification tasks on large data.

The developers created a novel two-stage architecture that decouples feature embedding from sample-level reasoning. TabICL first uses a shareable Set Transformer to build distribution-aware representations of each column, followed by a row-wise transformer equipped with rotary positional embeddings to aggregate features into compact row vectors while preventing representation collapse. A final transformer performs in-context learning solely over these condensed row vectors and their associated labels. The model was pretrained purely on synthetic datasets generated via causal and tree-based structures, using a three-stage curriculum learning schedule that progressively scaled training set sizes up to 60,000 samples. The architecture was evaluated across 200 real-world classification datasets from the TALENT benchmark against more than 30 deep learning and tree-based baselines.

The evaluation demonstrates that TabICL matches the predictive accuracy of TabPFNv2 on small and medium datasets while running systematically faster by 1.5 to 10 times. On large datasets containing more than 10,000 samples, TabICL outperforms both TabPFNv2 and traditional gradient boosting models such as CatBoost, achieving a top average rank in log-loss and area under the curve metrics. Across the entire benchmark, TabICL achieved the highest median relative accuracy improvement over baseline neural networks while requiring an average fit-and-predict time of only 1.1 seconds per 1,000 samples on a single graphics processing unit. In comparison, tuning standard models like CatBoost or deep neural networks required 3 to 7 minutes per 1,000 samples. Furthermore, by pairing TabICL with a hierarchical classification strategy, the model successfully scaled to problems with more than 10 classes without retraining the underlying feature representations.

These findings indicate that tabular foundation models can eliminate the need for costly hyperparameter tuning pipelines in enterprise workflows, drastically cutting down development timelines and computational expenses while improving probability calibration. Organizations should consider pilot deployments of TabICL for rapid tabular classification tasks, particularly where fast turnarounds or real-time inference are required. Decision-makers should note that the model is currently restricted to classification problems, requires ensembling over column permutations to restore ordering invariance, and experiences slow inference per sample compared to lightweight shallow models. Nevertheless, confidence in the reported performance is high across standard tabular benchmarks, and integrating hybrid approaches like decision-tree partitioning can further scale TabICL to massive datasets.

No sufficiently relevant recommendations were found.

Cover for TabICL: A Tabular Foundation Model for In-Context Learning on Large Data

Abstract

The long-standing dominance of gradient-boosted decision trees on tabular data is currently challenged by tabular foundation models using In-Context Learning (ICL): setting the training data as context for the test data and predicting in a single forward pass without parameter updates. While TabPFNv2 foundation model excels on tables with up to 10K samples, its alternating column- and row-wise attentions make handling large training sets computationally prohibitive. So, can ICL be effectively scaled and deliver a benefit for larger tables? We introduce TabICL, a tabular foundation model for classification, pretrained on synthetic datasets with up to 60K samples and capable of handling 500K samples on affordable resources. This is enabled by a novel two-stage architecture: a column-then-row attention mechanism to build fixed-dimensional embeddings of rows, followed by a transformer for efficient ICL. Across 200 classification datasets from the TALENT benchmark, TabICL is on par with TabPFNv2 while being systematically faster (up to 10 times), and significantly outperforms all other approaches. On 53 datasets with over 10K samples, TabICL surpasses both TabPFNv2 and CatBoost, demonstrating the potential of ICL for large data. Pretraining code, inference code, and pre-trained models are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Foundation Models and In-Context Learning
  • 2.2 Tabular Deep Learning Models
  • 2.3 Tabular Foundation Models
  • 3 The TabICL Architecture
  • 3.1 High-level Structure: Embedding then ICL
  • 3.2 Distribution-aware Column-wise Embedding
  • 3.3 Context-aware Row-wise Interaction
  • 3.4 Dataset-wise In-Context Learning
  • 3.5 Computational Complexity Analysis
  • 4 Pretraining and Inference
  • 4.1 Improved Pretraining Synthetic Datasets
  • 4.2 Curriculum Learning for Large-scale Pretraining
  • 4.3 Hierarchical Class-extension Strategy
  • 4.4 Memory-efficient Inference
  • 5 Experiments
  • 5.1 Benchmark
  • 5.2 Results
  • 5.3 Ablation Studies
  • 6 Conclusion
  • References
  • A Systematic Comparison between TabICL and TabPFNv2
  • A.1 Architecture
  • A.2 Pretraining
  • A.3 Scalability and Efficiency
  • B Further Experiments
  • C Synthetic Datasets for Pretraining
  • C.1 SCM prior with more activation functions
  • C.2 Tree-based SCM prior
  • D Rotary Positional Embedding
  • E Setup of TabICL
  • E.1 Pretraining details
  • E.2 Memory-efficient inference
  • E.3 Hierarchical Class-extension Strategy
  • F Excluded Development Datasets of TabPFNv2
  • G Random Forest Extension of TabPFNv2 for Large Datasets
  • H Effect of Ensemble Size on TabPFNv2 and TabICL
  • I Average Performance and Rankings

Knowls

  1. Knowl 1 — TabICL separates row representation learning from in-context prediction

    model/method

    TabICL is a pretrained tabular classification model that first converts each sample's features, without using labels, into a fixed-width row embedding. It then combines training-row embeddings with their labels and uses a separate transformer to predict labels for test rows. This late-fusion design shares the same feature embeddings across classification subproblems and lets the ICL stage operate on one token per row rather than on the original feature-by-sample table.

  2. Knowl 2 — Distribution-aware column embedding uses a shared set-conditioned hypernetwork

    model/method

    For a table with nn samples and mm features, TabICL applies the same column encoder to every feature. For a column cj∈Rnc_j\in\mathbb{R}^n, the encoder uses the set of training-cell values to produce a weight and bias for each cell, then forms its dd-dimensional embedding:

    (Wj,Bj)=TFcol(cj),ej=Wj⊙cj+Bj.(W_j,B_j)=TF_{\mathrm{col}}(c_j),\qquad e_j=W_j\odot c_j+B_j.

    Here Wj,Bj,ej∈Rn×dW_j,B_j,e_j\in\mathbb{R}^{n\times d}, scalar multiplication is broadcast across embedding dimensions, and ⊙\odot is elementwise multiplication. The encoder is a Set Transformer composed of three induced self-attention blocks (ISABs): it linearly projects the cell values to U∈Rn×dU\in\mathbb{R}^{n\times d}; kk learned inducing vectors query the training rows to form an induced representation; then each row, including test rows, queries that representation to produce per-cell outputs. Only training rows are keys and values in the first attention step, so test values do not determine the training-derived distributional context. The model uses d=128d=128, k=128k=128 inducing vectors, and four attention heads. This shared, set-conditioned mapping is intended to capture column-specific distributional regularities without a separate learned embedding module for each feature.

  3. Knowl 3 — Row-wise interaction compresses features and uses rotary positions to prevent collapse

    model/method

    After column encoding, each sample has mm feature embeddings in R128\mathbb{R}^{128}. TabICL prepends four learned [CLS] tokens to each sample and processes the resulting sequence with a three-layer, eight-head transformer, TFrowTF_{\mathrm{row}}. Concatenating the four final [CLS] outputs gives a 512-dimensional row embedding. Rotary positional embeddings (RoPE) are applied to the feature positions in attention, with rotation angle θi=p/1000002i/d\theta_i=p/100000^{2i/d} for position pp, dimension-pair index ii, and embedding dimension d=128d=128. RoPE breaks the symmetry that can otherwise make rows indistinguishable when features share the same distribution; the paper illustrates this failure and its mitigation on the balance-scale dataset. Because RoPE makes predictions depend on column order, inference averages predictions over shuffled column orders to approximately recover permutation invariance.

  4. Knowl 4 — Dataset-wise ICL predicts test labels from compressed row embeddings

    model/method

    Let H∈Rn×512H\in\mathbb{R}^{n\times512} denote the row embeddings produced by TabICL for nn samples. Training labels are one-hot encoded, mapped into the same representation space, and added to the corresponding training embeddings. A 12-layer, four-head transformer, TFiclTF_{\mathrm{icl}}, processes the training and test representations: training rows can attend to one another, while test rows can attend only to training rows. A two-layer MLP maps the resulting test representations to class probabilities. Thus, the full set of test predictions is produced in one ICL forward pass, with labels introduced only at this final stage.

  5. Knowl 5 — Collapsing features reduces TabICL's attention cost for large tables

    theoretical result

    For a table with nn rows, mm features, and a fixed number kk of column-encoder inducing vectors, TabICL's column encoder costs O(nkm)O(nkm), its row-wise feature attention costs O(m2n)O(m^2n), and its dataset-wise ICL attention costs O(n2)O(n^2). With fixed kk, the total time complexity is O(m2n+n2)O(m^2n+n^2). The key reduction comes from performing ICL after compressing each row to a fixed-width embedding. The paper contrasts this with TabPFNv2's O(m2n+n2m)O(m^2n+n^2m) complexity, which retains the feature dimension during row-wise attention over samples.

  6. Knowl 6 — Synthetic pretraining combines structural causal and tree-based priors

    model/method

    TabICL is pretrained exclusively on synthetic classification data generated from structural causal models (SCMs). A sampled directed acyclic graph specifies dependencies among variables, and each variable is generated as a function of its parent variables plus independent noise. To incorporate tree-model inductive biases, 30% of generated datasets use tree-based SCMs and 70% use the standard SCM prior. In a tree-based SCM, an XGBoost multi-output regressor is fitted at each graph layer using the layer inputs and independent standard-normal fake targets; its predictions become the child variables. The number of estimators is sampled as min⁡{4,1+Z}\min\{4,1+Z\} and maximum depth as min⁡{4,2+Z′}\min\{4,2+Z'\}, where Z,Z′Z,Z' are independent exponential random variables with rate 0.50.5. The standard SCM prior is also diversified with 15 additional activation functions, including non-monotone and discontinuous functions, and random functions constructed from random Fourier-style features; standardization and random rescaling are applied before activations.

  7. Knowl 7 — Curriculum pretraining scales synthetic tasks from 1,024 to 60,000 rows

    experimental setup

    TabICL uses three pretraining stages, with 512 synthetic datasets per optimization step and at most 100 features and 10 classes per dataset. Stage 1 trains for 160,000 steps on datasets of 1,024 rows with micro-batch size NB=4N_B=4. Stage 2 trains for 2,000 steps with NB=1N_B=1 and dataset sizes sampled log-uniformly from 1,000 to 40,000; activation checkpointing is used above 10,000 rows, and feature counts are reduced when needed to avoid memory exhaustion. Stage 3 trains for 50 steps with NB=1N_B=1 and sizes sampled uniformly from 40,000 to 60,000, freezing all components except TFiclTF_{\mathrm{icl}}. The complete pretraining run took 20 days on three 40-GB A100 GPUs and produced approximately 82 million synthetic datasets.

  8. Knowl 8 — Hierarchical prediction extends TabICL to more than ten classes

    model/method

    Because TabICL's pretraining tasks contain at most 10 classes, the model handles a larger classification problem by recursively dividing its classes into groups of at most 10. Internal nodes predict probabilities over their child groups; leaves predict probabilities over the classes they contain. For a class cc, its final probability is the product of the conditional probabilities along the unique path from the root to the leaf containing cc. A tree for kk classes has depth ⌈log⁡10k⌉\lceil\log_{10}k\rceil. All node-level tasks share the same label-independent row embeddings and the same ICL transformer; only their local targets differ. On the benchmark's 12 datasets with more than 10 classes, this extension gave TabICL the second-best mean normalized accuracy among the compared methods.

  9. Knowl 9 — Dynamic batching and offloading enable large-table inference

    model/method

    TabICL uses FlashAttention and dynamically selects batch sizes for its column, row, and ICL transformers based on sequence length and available GPU memory. Peak activation memory is estimated by fitting a polynomial of the form MEM=α1b+α2s+α3bs+α4\mathrm{MEM}=\alpha_1b+\alpha_2s+\alpha_3bs+\alpha_4, where bb is the batch size, ss is sequence length, and the fitted coefficients are specific to each transformer. Here bb means number of columns for column encoding, number of samples for row interaction, and number of datasets for ICL; ss means number of samples for column encoding and ICL, and number of features for row interaction. Intermediate activations can also be offloaded to CPU memory or disk. In reported inference measurements, a dataset with 100,000 samples and 500 features used about 5 GB of GPU memory and 25 GB of CPU memory; a 500,000-sample, 500-feature dataset required less than 14 GB of GPU memory and about 120 GB of CPU memory, with disk offloading available to reduce CPU-memory demand.

  10. Knowl 10 — Benchmark comparisons use training-only ICL ensembles and untuned baselines

    experimental setup

    The evaluation uses the TALENT tabular classification benchmark. The authors exclude 15 datasets used for TabPFNv2 development and focus their principal comparison on 171 datasets with at most 10 classes. Dataset splits contain 64% training, 16% validation, and 20% test data. TabICL and TabPFNv2 use training data only, without validation-based tuning, and average 32 predictions over shuffled columns and classes and different preprocessors; TabICL uses z-normalization with or without a power transform. TabPFNv2 is limited to 30,000 training samples by subsampling to avoid memory problems, whereas TabICL uses all available training samples. In the benchmark timing comparison, automatic mixed precision is disabled for all methods.

  11. Knowl 11 — TabICL ranks with TabPFNv2 while requiring less prediction time

    empirical result

    Across the 171 benchmark datasets with at most 10 classes, TabICL and TabPFNv2 were the leading methods and were not significantly different from one another in accuracy-based rankings. Their average ranks were 6.95 for TabICL and 7.17 for TabPFNv2, where lower is better; for AUC the ranks were 5.59 and 5.99, and for log loss they were 4.21 and 4.47. The paper reports that both ICL models significantly outperformed accuracy-tuned competitors on log loss, with no significant difference between the two ICL models. TabICL's geometric-mean training-plus-inference time was 1.1 seconds per 1,000 samples; the corresponding reported comparison was about three minutes for CPU CatBoost tuning and about seven minutes for GPU RealMLP or ModernNCA. In direct GPU comparisons with TabPFNv2, TabICL was about 1.5 times faster on small datasets and 3–10 times faster on large ones; for 10,000 samples and 100 features, the reported times were about 20 seconds and 100 seconds, respectively.

  12. Knowl 12 — TabICL performs strongly on datasets exceeding 10,000 samples

    empirical result

    On the 53 benchmark datasets with more than 10,000 samples, TabICL's average accuracy rank was 8.64, compared with 9.68 for CatBoost and 11.94 for TabPFNv2; lower rank indicates better performance. Its average AUC rank was 7.67, compared with 8.13 for CatBoost and 11.18 for TabPFNv2, and its log-loss rank was 6.00, compared with 8.13 and 8.75, respectively. The sample-size ranking analysis also shows that TabICL remains competitive as dataset size grows, unlike TabPFNv2, whose performance and memory use become limiting at larger sample counts. These results support the paper's claim that pretrained ICL can compete with tree-based methods in large-sample regimes.

  13. Knowl 13 — Ablations support tree-based priors and large-sample curriculum learning

    empirical result

    Adding tree-based SCMs to the synthetic pretraining prior improved TabICL's performance across the reported accuracy, AUC, and log-loss comparisons. Curriculum learning also improved average benchmark rank across its stages: TabICL moved from rank 11.4 after stage 1 to 7.46 after stage 2 and 6.95 after stage 3. The authors note a slight performance decrease on some small datasets during curriculum learning, even as performance improved overall and on large datasets.

  14. Knowl 14 — TabICL remains order-sensitive and is limited to classification

    limitation

    TabICL's RoPE feature-position encoding breaks exact invariance to column permutations, although prediction ensembling over shuffled columns is used to approximately restore it. Inference can still be slow despite being faster than TabPFNv2 in the reported comparisons, and the released method addresses classification rather than regression. The benchmark also limits how broadly the comparisons can be interpreted: it uses mean imputation for missing values, does not provide explicit categorical-feature metadata to TabPFNv2, and compares ICL ensembles with single models for other deep-learning methods.

Coverage note — The paper's detailed activation-function catalogue, fitted memory-regression coefficients, and per-dataset benchmark tables are omitted because their implementation details or individual entries are secondary to the main model, scalability, and comparative results captured here.

References

  1. 1.Ansari, A. F., Stella, L., Turkmen, C., Zhang, X., Mercado, P., Shen, H., Shchur, O., Rangapuram, S. S., Arango, S. P., Kapoor, S., et al. Chronos: Learning the language of time series. arXiv preprint arXiv:2403.07815, 2024.
  2. 2.Barbero, F., Vitvitskyi, A., Perivolaropoulos, C., Pascanu, R., and Velicković, P. Round and Round We Go! What makes Rotary Positional Encodings useful?, October 2024.
  3. 3.Bordt, S., Nori, H., and Caruana, R. Elephants never forget: Testing language models for memorization of tabular data. arXiv preprint arXiv:2403.06644, 2024.
  4. 4.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language Models are Few-Shot Learners, July 2020.
  5. 5.Cartella, F., Anunciacao, O., Funabiki, Y., Yamaguchi, D., Akishita, T., and Elshocht, O. Adversarial Attacks for Tabular Data: Application to Fraud Detection and Imbalanced Data, January 2021.
  6. 6.Chen, T. and Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 785–794, August 2016. doi: 10.1145/2939672.2939785.
  7. 7.Da Costa, N., Pfortner, M., Da Costa, L., and Hennig, P. Sample path regularity of gaussian processes from the covariance kernel. arXiv preprint arXiv:2312.14886, 2023.
  8. 8.den Breejen, F., Bae, S., Cha, S., and Yun, S.-Y. Why In-Context Learning Transformers are Tabular Data Classifiers, May 2024.
  9. 9.Dorogush, A. V., Ershov, V., and Gulin, A. CatBoost: Gradient boosting with categorical features support, October 2018.
  10. 10.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  11. 11.Fang, X., Xu, W., Tan, F. A., Zhang, J., Hu, Z., Qi, Y., Nickleach, S., Socolinsky, D., Sengamedu, S., and Faloutsos, C. Large language models(llms) on tabular data: Prediction, generation, and understanding – a survey, 2024.
  12. 12.Feuer, B., Schirrmeister, R. T., Cherepanova, V., Hegde, C., Hutter, F., Goldblum, M., Cohen, N., and White, C. TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks. arXiv preprint arXiv:2402.11137, 2024.
  13. 13.Gardner, J., Perdomo, J. C., and Schmidt, L. Large Scale Transfer Learning for Tabular Data via Language Modeling, June 2024.
  14. 14.Garg, S., Tsipras, D., Liang, P., and Valiant, G. What Can Transformers Learn In-Context? A Case Study of Simple Function Classes, August 2023.
  15. 15.Gneiting, T. and Raftery, A. E. Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477):359–378, 2007.
  16. 16.Gorishniy, Y., Rubachev, I., and Babenko, A. On embeddings for numerical features in tabular deep learning. Advances in Neural Information Processing Systems, 35:24991–25004, 2022.
  17. 17.Gorishniy, Y., Kotelnikov, A., and Babenko, A. TabM: Advancing Tabular Deep Learning with Parameter-Efficient Ensembling, November 2024.
  18. 18.Grinsztajn, L. Réconcilier l’apprentissage profond avec les données tabulaires. PhD thesis, Université Paris-Saclay, 2024.
  19. 19.Grinsztajn, L., Oyallon, E., and Varoquaux, G. Why do tree-based models still outperform deep learning on tabular data?, July 2022.
  20. 20.Hegselmann, S., Buendia, A., Lang, H., Agrawal, M., Jiang, X., and Sontag, D. Tabllm: Few-shot classification of tabular data with large language models. In International Conference on Artificial Intelligence and Statistics, pp. 5549–5581. PMLR, 2023.
  21. 21.Hendrycks, D. and Gimpel, K. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  22. 22.Hollmann, N., Müller, S., Eggensperger, K., and Hutter, F. Tabpfn: A transformer that solves small tabular classification problems in a second. arXiv preprint arXiv:2207.01848, 2022.
  23. 23.Hollmann, N., Müller, S., Purucker, L., Krishnakumar, A., Körfer, M., Hoo, S. B., Schirrmeister, R. T., and Hutter, F. Accurate predictions on small data with a tabular foundation model. Nature, 637(8045):319–326, January 2025. ISSN 1476-4687. doi: 10.1038/s41586-024-08328-6.
  24. 24.Johnson, A. E. W., Pollard, T. J., Shen, L., Lehman, L.-w. H., Feng, M., Ghassemi, M., Moody, B., Szolovits, P., Anthony Celi, L., and Mark, R. G. MIMIC-III, a freely accessible critical care database. Scientific Data, 3(1):160035, May 2016. ISSN 2052-4463. doi: 10.1038/sdata.2016.35.
  25. 25.Kim, M. J., Grinsztajn, L., and Varoquaux, G. CARTE: Pretraining and Transfer for Tabular Learning, May 2024.
  26. 26.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  27. 27.Klambauer, G., Unterthiner, T., Mayr, A., and Hochreiter, S. Self-normalizing neural networks. Neural Information Processing Systems, 30, 2017.
  28. 28.Koshil, M., Nagler, T., Feurer, M., and Eggensperger, K. Towards Localization via Data Embedding for TabPFN. 2024.
  29. 29.Kossen, J., Band, N., Lyle, C., Gomez, A. N., Rainforth, T., and Gal, Y. Self-attention between datapoints: Going beyond individual input-output pairs in deep learning. Advances in Neural Information Processing Systems, 34:28742–28756, 2021.
  30. 30.Krizhevsky, A. Convolutional deep belief networks on cifar-10.
  31. 31.Lee, J., Lee, Y., Kim, J., Kosiorek, A. R., Choi, S., and Teh, Y. W. Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks, May 2019.
  32. 32.Ma, J., Thomas, V., Hosseinzadeh, R., Kamkari, H., Labach, A., Cresswell, J. C., Golestan, K., Yu, G., Volkovs, M., and Caterini, A. L. TabDPT: Scaling Tabular Foundation Models, October 2024a.
  33. 33.Ma, J., Thomas, V., Yu, G., and Caterini, A. In-Context Data Distillation with TabPFN. arXiv preprint arXiv:2402.06971, 2024b.
  34. 34.Mikolov, T., Chen, K., Corrado, G., and Dean, J. Efficient Estimation of Word Representations in Vector Space, September 2013.
  35. 35.Müller, A., Curino, C., and Ramakrishnan, R. MotherNet: A Foundational Hypernetwork for Tabular Classification, December 2023.
  36. 36.Müller, S., Hollmann, N., Arango, S. P., Grabocka, J., and Hutter, F. Transformers Can Do Bayesian Inference, August 2024.
  37. 37.Rahimi, A. and Recht, B. Random features for large-scale kernel machines. Neural Information Processing Systems, 20, 2007.
  38. 38.Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S., and He, Y. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning, April 2021.
  39. 39.Rubachev, I., Kartashev, N., Gorishniy, Y., and Babenko, A. Tabred: Analyzing pitfalls and filling the gaps in tabular deep learning benchmarks. arXiv preprint arXiv:2406.19380, 2024.
  40. 40.Siegler, R. Balance Scale. UCI Machine Learning Repository, 1976. DOI: https://doi.org/10.24432/C5488X.
  41. 41.Silla, C. N. and Freitas, A. A. A survey of hierarchical classification across different application domains. Data Mining and Knowledge Discovery, 22(1):31–72, January 2011. ISSN 1573-756X. doi: 10.1007/s10618-010-0175-9.
  42. 42.Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y. RoFormer: Enhanced Transformer with Rotary Position Embedding, November 2023.
  43. 43.Thawani, A., Pujara, J., Szekely, P. A., and Ilievski, F. Representing Numbers in NLP: A Survey and a Vision, March 2021.
  44. 44.Thomas, V., Ma, J., Hosseinzadeh, R., Golestan, K., Yu, G., Volkovs, M., and Caterini, A. Retrieval & Fine-Tuning for In-Context Tabular Models, June 2024.
  45. 45.van Breugel, B. and van der Schaar, M. Why Tabular Foundation Models Should Be a Research Priority, June 2024.
  46. 46.von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M. Transformers learn in-context by gradient descent, May 2023.
  47. 47.Xie, S. M., Raghunathan, A., Liang, P., and Ma, T. An Explanation of In-context Learning as Implicit Bayesian Inference, July 2022.
  48. 48.Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H. Effective Long-Context Scaling of Foundation Models, November 2023.
  49. 49.Xu, D., Cirit, O., Asadi, R., Sun, Y., and Wang, W. Mixture of In-Context Prompters for Tabular PFNs, May 2024.
  50. 50.Ye, H.-J., Yin, H.-H., and Zhan, D.-C. Modern Neighborhood Components Analysis: A Deep Tabular Baseline Two Decades Later, July 2024.
  51. 51.Ye, H.-J., Liu, S.-Y., Cai, H.-R., Zhou, Q.-L., and Zhan, D.-C. A Closer Look at Deep Learning Methods on Tabular Datasets, January 2025.
  52. 52.Zeng, Y., Kang, W., and Mueller, A. C. Tabflex: Scaling tabular learning to millions with linear attention. In NeurIPS 2024 Third Table Representation Learning Workshop, 2024.
  53. 53.Zhou, C., Li, Q., Li, C., Yu, J., Liu, Y., Wang, G., Zhang, K., Ji, C., Yan, Q., He, L., Peng, H., Li, J., Wu, J., Liu, Z., Xie, P., Xiong, C., Pei, J., Yu, P. S., and Sun, L. A comprehensive survey on pretrained foundation models: A history from BERT to ChatGPT. International Journal of Machine Learning and Cybernetics, November 2024. ISSN 1868-808X. doi: 10.1007/s13042-024-02443-6.
  54. 54.Zhu, B., Shi, X., Erickson, N., Li, M., Karypis, G., and Shoaran, M. XTab: Cross-table Pretraining for Tabular Transformers, May 2023.

Citation

MLA
Qu, J., et al. “TabICL: A Tabular Foundation Model for In-Context Learning on Large Data”. arXiv, 2025, http://arxiv.org/abs/2502.05564v2.
APA
Qu, J., Holzmüller, D., Varoquaux, G., & Morvan, M. L. (2025). TabICL: A Tabular Foundation Model for In-Context Learning on Large Data. arXiv. http://arxiv.org/abs/2502.05564v2
Chicago
Qu, J., D. Holzmüller, G. Varoquaux, and M. L. Morvan. 2025. “TabICL: A Tabular Foundation Model for In-Context Learning on Large Data”. arXiv. http://arxiv.org/abs/2502.05564v2.
Harvard
Qu, J. et al. (2025) “TabICL: A Tabular Foundation Model for In-Context Learning on Large Data”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.05564v2.
Vancouver
1. Qu J, Holzmüller D, Varoquaux G, Morvan ML (2025) TabICL: A Tabular Foundation Model for In-Context Learning on Large Data. arXiv

BibTeX

@article{qu2025tabicl,
  title = {TabICL: A Tabular Foundation Model for In-Context Learning on Large Data},
  author = {Qu, Jingang and Holzmüller, David and Varoquaux, Gaël and Morvan, Marine Le},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.05564v2},
  eprint = {2502.05564}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/