← All stories
● Covered by 1 source · 1 reportMedium impact1 neutral

TabPFN and TabICL Outperform Tuned XGBoost on 14 Tabular Datasets Without Training

🔄 Updated 4d ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • TabPFN/TabICL beat tuned XGBoost on 14 datasets.
  • Models predict without training on new tabular data.
  • Performance holds up to 32,000 rows.
  • Hyperparameter tuning may become less critical.

Tabular Foundation Models Tested

A comparison was conducted between TabPFN and TabICL, which are tabular foundation models, and a tuned XGBoost model. The evaluation used fourteen datasets from the Grinsztajn benchmark, maintaining consistent data splits and processing times for all models.

Performance Results

The tabular foundation models, TabPFN and TabICL, won on all fourteen datasets. This performance advantage persisted even with up to 32,000 rows of data. The key distinction is that these models predict on new data without requiring specific training on that data.

How Tabular Foundation Models Work

Tabular foundation models are pretrained on millions of synthetically generated tables. When presented with a new table, they do not adjust their internal weights. Instead, they process the training rows as context and generate predictions in a single forward pass, similar to in-context learning in language models. This process, while internally referred to as 'fit', does not involve gradient descent but rather a data copy operation.

Implications for Machine Learning Practice

The results indicate a potential change in how machine learning is applied to tabular data. If these models consistently outperform tuned boosting methods without requiring dataset-specific training, the mandatory step of searching for hyperparameters could become less critical. This could simplify and accelerate the deployment of models for tabular prediction tasks.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~34 min · 27 stories · Oct 02

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

Tabular Foundation Models (TabPFN and TabICL) demonstrated superior performance over tuned XGBoost across 14 tabular datasets, despite not undergoing training on the specific data. This suggests a potential shift in machine learning practices for tabular data, reducing the need for extensive hyperparameter tuning.