Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In Efrain Garay’s 2026 benchmark, TabICL recorded a higher AUC than tuned XGBoost on all 14 selected classification datasets, including after XGBoost was retuned to optimize AUC. That result applies to this experiment—not to every tabular problem, and not to TabPFN: TabPFN’s results varied by dataset.
What the 14-for-14 result means
Garay tested TabICL 2.x, TabPFN 2.2.1, default XGBoost, and XGBoost tuned with randomized search. The headline finding is specifically about TabICL’s AUC direction against tuned XGBoost: Garay reports TabICL ahead on all 14 datasets after the XGBoost search was scored for AUC. The reported mean AUC gap in that rerun was 0.0106.
AUC measures how well a classifier ranks positive cases above negative ones across possible decision thresholds. It is not the same as accuracy, which counts correct classifications at a chosen threshold. The 14/14 claim therefore does not mean TabICL made more correct predictions on every dataset, nor that both TabICL and TabPFN beat XGBoost on every comparison.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsHow the benchmark was set up
Garay selected 14 classification datasets from the Grinsztajn tabular benchmark, capped each at 3,000 rows, and reported medians over five seeds. The tuned XGBoost configuration used 25 randomized-search iterations and three-fold cross-validation. The benchmark measured fitting and prediction time separately. Garay describes these smaller datasets as a favorable setting for in-context tabular models.
#1 Best Overall
The reported setup used XGBoost 3.4.1, PyTorch 2.9.1, a 16 GB NVIDIA GeForce RTX 4070 Ti SUPER, and 14 CPU cores allocated to XGBoost. Those details matter for reproducibility; they are not requirements or hardware recommendations.
Why the XGBoost rerun matters
The first tuned-XGBoost search optimized accuracy, even though the comparison emphasized AUC. That metric mismatch could disadvantage XGBoost on an AUC comparison. Garay therefore reran the search using ROC AUC as the scoring metric. The reported TabICL lead remained in all 14 datasets, while the mean gap shifted from 0.0114 in the first comparison to 0.0106 after the AUC-scored rerun.
Rank #2
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
Accuracy results were less uniform: Garay reports TabICL had the higher median accuracy on 12 of 14 datasets, but only about seven of those comparisons remained outside the seed-to-seed spread. The author also reports the same direction in 68 of 70 per-seed comparisons. These are the benchmark author’s summaries, not an independent uncertainty analysis.
Free tools Windows power users keep installed
One-click scans. No signup required.
What the dataset examples show
Individual results varied, including cases where TabPFN led TabICL. The following are Garay’s displayed seed-0 AUC examples, not five-seed medians:
Rank #3
| Dataset | TabICL | TabPFN | Tuned XGBoost |
|---|---|---|---|
| Credit | 0.7667 | 0.7578 | 0.7533 |
| HELOC | 0.7222 | 0.7300 | 0.7078 |
| Default of credit | 0.6956 | 0.6967 | 0.6944 |
| Bank marketing | 0.7944 | 0.7967 | 0.7833 |
Bank marketing illustrates why the aggregate claim should not be stretched: TabPFN’s displayed score was slightly higher than TabICL’s, while both exceeded tuned XGBoost in that example.
“Does not train” still involves prior training and inference
TabPFN and TabICL use in-context learning. Their models are pretrained before a new dataset arrives; at prediction time, the new table’s training rows are supplied as context. They do not ordinarily update model weights through gradient-based fitting for each new dataset in the way conventional training does. Garay summarizes the concept this way: “A tabular foundation model is pretrained on millions of synthetic tables generated on purpose.” That is the author’s conceptual description, not a precise claim about every version.
Rank #4
A software API may still expose a method named fit. The name alone does not establish that the method performs gradient descent on task-specific weights. Nor does “does not train” mean no prior training or no computation when making predictions.
Fitting time is only part of the compute trade-off
Because these models condition on training rows at prediction time, they shift work away from conventional dataset-specific fitting and can spend more time on inference. In Garay’s displayed 419-column Bioresponse example, TabICL scored 0.8667 AUC and took 6.0 seconds to predict. Several other displayed examples had prediction times around 0.6–0.8 seconds. These are measurements from that setup, not general speed guarantees or evidence of a maximum supported feature width.
Best Value
For a real workload, measure both fitting and prediction on representative data, including the expected number of prediction batches. A method that reduces setup time may still be a poor operational fit if repeated inference is slow or costly.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Where the result may not transfer
The experiment does not settle which model is best for larger datasets, different feature types, other preprocessing choices, or a different XGBoost search budget. The row cap and selected benchmark suite define the scope; they are not a random sample of every tabular task. Garay notes that test sets had 900 rows and estimates AUC standard error near 0.01, making the per-seed direction more informative than any single small margin.
Use the comparison as evidence for a promising option on similarly sized classification tasks, not as a replacement rule. For deployment or consequential decisions such as credit, readmission, or biological-response prediction, benchmark on data that represents the intended population and assess calibration, threshold-specific errors, fairness, and operational requirements separately.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Reproducing or applying the comparison
- Review Garay’s benchmark script and results alongside the benchmark article to confirm the exact preprocessing and dataset handling before interpreting the reported numbers.
- For a comparable test, keep the metric aligned across model selection and evaluation: if AUC is the target, tune XGBoost using ROC AUC rather than accuracy.
- Use the same train/test partitions and repeat across seeds; report both aggregate results and per-seed variability.
- Measure fitting and prediction separately on your own row counts, feature widths, hardware, and inference volume.
- Check current implementation versions, supported limits, installation instructions, and license terms in the official TabICL project and official TabPFN project. The benchmark’s package versions are historical details of Garay’s experiment, not a statement of current defaults.
For background distinct from Garay’s comparison, the TabPFN Nature paper reports favorable results on its own small-tabular benchmarks against tuned baselines; it does not independently validate the 14-dataset TabICL result.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

