Conference Presentation, Product Demonstration
The Foundation Model Revolution for Structured Data | Frank Hutter, Prior Labs | RAISE Summit 2026
Market Context & Problem Definition
- Tabular foundation models (TFMs) are an emerging AI category that has seen exponential growth in the last 18 months, addressing a gap left by Large Language Models (LLMs).
- LLMs excel at unstructured data (text, code, summarization) but fail at structured enterprise data (databases, CSVs, time series, forecasts, risk scoring) due to poor statistical understanding and numerical accuracy.
- Traditional machine learning workflows require 3–6 months to solve new tabular problems, with 60% of time spent on governance, data cleaning, and feature engineering.
- Legacy models suffer from zero knowledge transfer; when underlying data distributions change, models must be retrained from scratch, often resulting in "intern-built" legacy systems that are difficult to govern.
Prior Labs & Team Composition
- Founded 18 months ago by CEO Frank Hutter, CTO Noah, and CEO Souraj Gambier to build world-class TFMs combining deep learning, statistics, and data science.
- Advisors include Yann LeCun and Bernhard Schölkopf.
- The company has a research-heavy team of 40 selected from 10,000 applicants, boasting over 1 million collective citations, two Forbes 30 Under 30s, and a paper published in Nature (2025) that became the most cited paper in that journal for the year.
- Open Source Strategy: The model has been downloaded over 3 million times and holds 7,000 GitHub stars.
Technical Architecture & Capabilities
- TFN (TabPFN): An "in-context learner" that takes training data (X_train, Y_train) and test data (X_test) in a single forward pass to predict missing values (Y_test) without retraining.
- Performance Metrics: Trained on hundreds of millions of synthetic datasets to prevent leakage; uses cross-entropy loss to ensure smooth decision surfaces and avoid overfitting.
- Efficiency: Reduces model deployment time from six months to seconds; operates with approximately 10 million parameters (compared to the billions in LLMs).
- Key Advantages over LLMs:
- Understands table invariances (e.g., row/column swapping) for better data efficiency.
- Deterministic output with no hallucinations, crucial for regulated industries.
- Significantly faster (millions of times) for data science tasks due to optimized token handling.
- Capabilities: Supports causality, interpretability, time series, relational databases, and text integration; fine-tunable for specific enterprise data.
Benchmark Performance & Case Studies
- TabArena Leaderboard: TFN3, the company's latest model, leads on the leading tabular benchmark using "thinking mode" (test-time compute) for higher accuracy and speed.
- Hitachi: Used for preventive maintenance on Spain's high-speed rail; achieved a 40% reduction in errors compared to custom models and allowed a single model to serve multiple track sections.
- Credit Agricole (Credit Plus): Improved car loan approval rates without raising risk thresholds; the model's logic was distilled into interpretable tree-based models for regulatory compliance.
- Oxford Cancer Analytics: Built a lung disease prediction platform; achieved state-of-the-art performance using only 50% of the data required by competing models, addressing data scarcity.
Synergy with LLMs & Product Integration
- MCP Server: Enables integration with AI agents; LLMs handle natural language interfaces (e.g., "predict demand") while TFNs handle the actual data processing and prediction.
- Demo Results: LLM-only agents achieved ~50% accuracy on churn prediction; integrating TFNs significantly improved prediction quality while maintaining a seamless user experience within dashboards (e.g., Databricks integration).
- Scalability: Enables the shift from hundreds of static models to individual models generated on-the-fly for every customer using LLMs to generate datasets and TFNs to predict.
Interpretability, Trust & Regulatory Compliance
- Design Philosophy: TFMs prioritize trustworthiness, robustness, fairness, and causal reasoning.
- Decision Surfaces: Training with cross-entropy loss creates smooth decision surfaces, making trends more interpretable than the jagged surfaces of tree-based models.
- Distillation: High-performance models can be distilled into tree-based models or MLPs for regulators with minimal performance loss.
- Feature Importance: Supports forward-backward passes to determine feature and data point importance, overcoming limitations of XGBoost.
Strategic Acquisition & Future Outlook
- Acquisition Agreement: SAP and Prior Labs signed a definitive agreement for SAP to acquire Prior Labs 18 months after its founding.
- Investment Terms: SAP will invest over €1 billion over the next four years to scale Prior Labs into a leading European AI lab.
- Operational Independence: Prior Labs will retain its brand, team, customers, offices, open-weight models, and publications.
- Strategic Goal: Position the company as a leader in "sovereign AI" in Europe, addressing the need for massive compute investment and open-source support.
- Forward-Looking Statement: The leadership asserts that Europe can move rapidly to become a world leader in specialized foundation models that revolutionize tabular predictions.