Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In Efrain Garay’s 2026 benchmark, TabICL recorded a higher AUC than tuned XGBoost on all 14 selected classification datasets, including after XGBoost was retuned to optimize AUC. That result applies to this experiment—not to every tabular problem, and not to TabPFN: TabPFN’s results varied by dataset.

What the 14-for-14 result means

Garay tested TabICL 2.x, TabPFN 2.2.1, default XGBoost, and XGBoost tuned with randomized search. The headline finding is specifically about TabICL’s AUC direction against tuned XGBoost: Garay reports TabICL ahead on all 14 datasets after the XGBoost search was scored for AUC. The reported mean AUC gap in that rerun was 0.0106.

AUC measures how well a classifier ranks positive cases above negative ones across possible decision thresholds. It is not the same as accuracy, which counts correct classifications at a chosen threshold. The 14/14 claim therefore does not mean TabICL made more correct predictions on every dataset, nor that both TabICL and TabPFN beat XGBoost on every comparison.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the benchmark was set up

Garay selected 14 classification datasets from the Grinsztajn tabular benchmark, capped each at 3,000 rows, and reported medians over five seeds. The tuned XGBoost configuration used 25 randomized-search iterations and three-fold cross-validation. The benchmark measured fitting and prediction time separately. Garay describes these smaller datasets as a favorable setting for in-context tabular models.

The reported setup used XGBoost 3.4.1, PyTorch 2.9.1, a 16 GB NVIDIA GeForce RTX 4070 Ti SUPER, and 14 CPU cores allocated to XGBoost. Those details matter for reproducibility; they are not requirements or hardware recommendations.

Why the XGBoost rerun matters

The first tuned-XGBoost search optimized accuracy, even though the comparison emphasized AUC. That metric mismatch could disadvantage XGBoost on an AUC comparison. Garay therefore reran the search using ROC AUC as the scoring metric. The reported TabICL lead remained in all 14 datasets, while the mean gap shifted from 0.0114 in the first comparison to 0.0106 after the AUC-scored rerun.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Accuracy results were less uniform: Garay reports TabICL had the higher median accuracy on 12 of 14 datasets, but only about seven of those comparisons remained outside the seed-to-seed spread. The author also reports the same direction in 68 of 70 per-seed comparisons. These are the benchmark author’s summaries, not an independent uncertainty analysis.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the dataset examples show

Individual results varied, including cases where TabPFN led TabICL. The following are Garay’s displayed seed-0 AUC examples, not five-seed medians:

Dataset TabICL TabPFN Tuned XGBoost
Credit 0.7667 0.7578 0.7533
HELOC 0.7222 0.7300 0.7078
Default of credit 0.6956 0.6967 0.6944
Bank marketing 0.7944 0.7967 0.7833

Bank marketing illustrates why the aggregate claim should not be stretched: TabPFN’s displayed score was slightly higher than TabICL’s, while both exceeded tuned XGBoost in that example.

“Does not train” still involves prior training and inference

TabPFN and TabICL use in-context learning. Their models are pretrained before a new dataset arrives; at prediction time, the new table’s training rows are supplied as context. They do not ordinarily update model weights through gradient-based fitting for each new dataset in the way conventional training does. Garay summarizes the concept this way: “A tabular foundation model is pretrained on millions of synthetic tables generated on purpose.” That is the author’s conceptual description, not a precise claim about every version.

A software API may still expose a method named fit. The name alone does not establish that the method performs gradient descent on task-specific weights. Nor does “does not train” mean no prior training or no computation when making predictions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fitting time is only part of the compute trade-off

Because these models condition on training rows at prediction time, they shift work away from conventional dataset-specific fitting and can spend more time on inference. In Garay’s displayed 419-column Bioresponse example, TabICL scored 0.8667 AUC and took 6.0 seconds to predict. Several other displayed examples had prediction times around 0.6–0.8 seconds. These are measurements from that setup, not general speed guarantees or evidence of a maximum supported feature width.

For a real workload, measure both fitting and prediction on representative data, including the expected number of prediction batches. A method that reduces setup time may still be a poor operational fit if repeated inference is slow or costly.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Where the result may not transfer

The experiment does not settle which model is best for larger datasets, different feature types, other preprocessing choices, or a different XGBoost search budget. The row cap and selected benchmark suite define the scope; they are not a random sample of every tabular task. Garay notes that test sets had 900 rows and estimates AUC standard error near 0.01, making the per-seed direction more informative than any single small margin.

Use the comparison as evidence for a promising option on similarly sized classification tasks, not as a replacement rule. For deployment or consequential decisions such as credit, readmission, or biological-response prediction, benchmark on data that represents the intended population and assess calibration, threshold-specific errors, fairness, and operational requirements separately.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reproducing or applying the comparison

  1. Review Garay’s benchmark script and results alongside the benchmark article to confirm the exact preprocessing and dataset handling before interpreting the reported numbers.
  2. For a comparable test, keep the metric aligned across model selection and evaluation: if AUC is the target, tune XGBoost using ROC AUC rather than accuracy.
  3. Use the same train/test partitions and repeat across seeds; report both aggregate results and per-seed variability.
  4. Measure fitting and prediction separately on your own row counts, feature widths, hardware, and inference volume.
  5. Check current implementation versions, supported limits, installation instructions, and license terms in the official TabICL project and official TabPFN project. The benchmark’s package versions are historical details of Garay’s experiment, not a statement of current defaults.

For background distinct from Garay’s comparison, the TabPFN Nature paper reports favorable results on its own small-tabular benchmarks against tuned baselines; it does not independently validate the 14-dataset TabICL result.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.