
NVIDIA Kumo Tabular: Open Tabular Foundation Model Tops 4 Benchmarks
NVIDIA launches Kumo Tabular, an open foundation model for tabular data, on Hugging Face in three sizes (28M to 215M parameters) with no training, fine-tuning, or feature engineering required. It ranks first on four standard benchmarks (TabArena, BeyondArena, TALENT, ScoringBench) with an ELO of 1950 and is 17x faster than LimiX-2, under the OpenMDW-1.1 license for commercial use.
Executive Overview
NVIDIA has officially announced the launch of NVIDIA Kumo Tabular, an open foundation model for tabular data, available immediately on Hugging Face as part of the NVIDIA Kumo Structured suite. The model predicts new row labels in a single forward pass without training, fine-tuning, or feature engineering, and it ranks first on four global standard benchmarks. With three sizes (28M to 215M parameters) and an OpenMDW-1.1 license for commercial use, Kumo Tabular sets a new performance bar for tabular machine learning.
📊 Official Data & Technical Specifications Sheet
| Technical Aspect | Confirmed Official Data |
|---|---|
| 💰 Pricing & Usage Cost | Completely free — open-source model under OpenMDW-1.1 license for commercial use with no licensing fees. The only cost is GPU consumption for inference (compatible with RTX 6000 Pro). |
| 🌐 Platforms & Immediate Availability | Hugging Face (weights: huggingface.co/nvidia/Kumo-Tabular), GitHub (code: github.com/NVIDIA/structured-data-models), NVIDIA GPU-native library for structured data, local execution via GPU. |
| ⚡ Performance & Speed Metrics | ELO 1950 on TabArena (1st place), 17x faster than LimiX-2 on RTX 6000 Pro, ELO 1418 on BeyondArena with 7.78% Improvability, 1st place on TALENT and ScoringBench. |
| 🛡️ Security & Breach Resistance | Follows NVIDIA Trustworthy AI policies. Does not process free text (only numerical and categorical columns), reducing prompt injection attack surface. Recommended to validate accuracy and calibration on held-out data before deployment. |
| 🧠 Context Window | Handles context tables up to 60,000 rows and 100 columns in the third training stage. A single pass covers up to 10 classes, expandable to any number of classes via Error-Correcting Output Codes. |
| 🌍 Arabic Language & Regional Support | Does not directly support raw text (numerical and categorical only). Arabic text can be encoded as categorical columns or numerical features via built-in preprocessing recipes. Available globally, including the Arab region, via Hugging Face. |
Deep-Dive Features & Architecture
Kumo Tabular features a Transformer architecture built around the table structure itself, using column, row, and context attention inspired by TabICL and TabPFN. The model consists of three stages: Cell Embedding, where a set of cells is transformed into a token via Fourier features with separate weights for numerical and categorical values; Row Embedding, through alternating column attention (linear cost with number of rows) and row attention with 4 learnable [CLS] tokens; and finally In-context Learning, where query rows attend only to context rows, allowing reuse of keys and values for subsequent predictions via Test-GQA.
A key innovation is the Length-aware Attention Temperature feature, which addresses attention fading as tables grow: each query is scaled by a temperature that grows logarithmically with the number of keys, with a coefficient learned separately for each attention head. The model was trained exclusively on synthetic tables generated from a Structural Causal Model (SCM) through six steps, and Kumo Tabular-Small/Medium/Large saw approximately 35/71/137 million synthetic tables respectively. For regression, the model outputs 999 quantiles to derive point predictions and uncertainty estimates.
Benchmark & Competitive Performance
On the TabArena benchmark, Kumo Tabular achieved first place with an ELO score of 1950, outperforming tuned boosted trees, AutoGluon, and the latest tabular foundation models, with a speed 17 times faster than LimiX-2 in a unified evaluation setup on RTX 6000 Pro. On BeyondArena, it reached an ELO of 1418 with an Improvability score of 7.78%, taking the top spot. On TALENT, it achieved the best overall ranking with average ranks of 6.67 for classification accuracy, 3.98 for classification log-loss, and 4.22 for regression RMSE. On ScoringBench, dedicated to predictive distributions, Kumo Tabular-Large and Medium took first and second place by average rank.
Industry Impact & Enterprise Adoption
The launch of Kumo Tabular represents a significant shift for enterprises dealing with structured data. By eliminating the need for training, fine-tuning, and feature engineering, it drastically reduces the time and expertise required to deploy predictive models on tabular datasets. The OpenMDW-1.1 license allows unrestricted commercial use, making it attractive for industries such as finance, healthcare, retail, and manufacturing where tabular data is ubiquitous. The model's ability to handle up to 60,000 rows and 100 columns in a single forward pass, combined with its 17x speed advantage over LimiX-2, positions it as a cost-effective solution for large-scale inference. Furthermore, its compatibility with NVIDIA RTX 6000 Pro GPUs ensures that organizations can run it locally, addressing data privacy and latency concerns. While it does not natively support raw Arabic text, the ability to encode Arabic text as categorical or numerical features opens doors for regional adoption, especially in markets where structured data is predominant.
Conclusion
NVIDIA Kumo Tabular sets a new standard for open tabular foundation models, combining state-of-the-art performance across four benchmarks with a permissive commercial license and immediate availability on Hugging Face. Its innovative architecture, including length-aware attention temperature and in-context learning, enables efficient and accurate predictions without the overhead of traditional machine learning pipelines. For developers and enterprises worldwide, including the Arab region, Kumo Tabular offers a powerful, free, and scalable tool for tabular data tasks, with the only cost being GPU inference. As the model gains traction, it is poised to accelerate AI adoption in data-rich industries.
Media Source: Hugging Face | البيان الرسمي للشركة: المصدر الأصلي | Fact Verification & Analysis: AI Tools Oasis
Frequently Asked Questions
NVIDIA Kumo Tabular is an open foundation model for tabular data that predicts new row labels in a single forward pass without training, fine-tuning, or feature engineering. It supports both classification and regression and is available in three sizes: 28M and 215M parameters. It is part of the NVIDIA Kumo Structured suite.
The model is immediately available on Hugging Face at huggingface.co/nvidia/Kumo-Tabular, with source code on GitHub at github.com/NVIDIA/structured-data-models. It runs via NVIDIA's new GPU-native library for structured data, which loads weights from the Hub on first use.
The model is licensed under the OpenMDW License Agreement version 1.1, which explicitly permits commercial use. This makes it available for companies and developers to integrate into production applications without commercial licensing restrictions.
Kumo Tabular tops four benchmarks: TabArena with an ELO of 1950 and 17x faster than LimiX-2 on RTX 6000 Pro; BeyondArena with an ELO of 1418 and 7.78% Improvability; TALENT with the best overall ranking (average rank 6.67 for classification, 3.98 for log-loss, 4.22 for RMSE); and ScoringBench where Large and Medium took first and second place.
Kumo Tabular operates on numerical and categorical columns only, while text, images, and timestamps can be converted to features via built-in preprocessing recipes. There is no direct support for Arabic as raw text, but Arabic text can be encoded as categorical columns or numerical features before feeding to the model.

AI Tools Oasis Team
Bringing you the latest news and analysis in the world of Artificial Intelligence with accuracy and credibility. Follow us for all updates.
