NVIDIA Releases Kumo Tabular Models
NVIDIA introduces Kumo Tabular, open tabular foundation models ranging from 28M to 215M parameters with commercial licensing support.

Stock photo for illustration only, not from the actual event
- NVIDIA releases Kumo Tabular as open tabular foundation models
- Available in Small, Medium, and Large sizes, spanning 28M to 215M parameters
- Powered by NVIDIA's open-source structured-data-models (SDM) library
- Model weights ship under the commercial-friendly OpenMDW-1.1 license
NVIDIA has officially announced the release of Kumo Tabular, a set of tabular foundation models designed to predict new rows in a single forward pass. The models are available in three distinct sizes—Small, Medium, and Large—ranging from approximately 28 million to 215 million parameters. They operate through NVIDIA's structured-data-models (SDM) library, an open-source framework built specifically for structured data models.
Regarding deployment and accessibility, Kumo Tabular stands out due to its permissive licensing. The model weights ship under the OpenMDW-1.1 license, which explicitly permits commercial use. Meanwhile, the underlying SDM codebase uses the Apache-2.0 license, requiring Python 3.11 and PyTorch 2.7, with execution examples targeting CUDA GPUs.

Stock photo for illustration only, not from the actual event
The SDM library serves as a GPU-native framework for structured data foundation models and preprocessing pipelines. Alongside Kumo Tabular, the library includes models such as TabICLv2, Google's TabFM, and KumoRelational designed for multi-table datasets. All these models share a unified in-context learning interface built on top of a TableTensor container, while also managing preprocessing, ensembling, and many-class prediction tasks.
Kumo Tabular's architecture is structured around tabular layouts, utilizing column, row, and in-context attention mechanisms. Its processing pipeline consists of three main stages, allowing keys and values to be computed once and reused efficiently. The model head outputs class probabilities or 999 quantiles for regression tasks, delivering both point predictions and uncertainty estimates simultaneously.
The emergence of tabular foundation models represents a major shift in enterprise machine learning. While large language models dominate unstructured text, the vast majority of business data resides in relational databases and structured tables. By releasing models like Kumo Tabular under permissive commercial terms such as OpenMDW-1.1, NVIDIA is lowering the barrier for organizations to leverage advanced deep learning on tabular datasets without building pipelines from scratch.
During its pretraining phase, Kumo Tabular was trained entirely on synthetic tables sampled from Structural Causal Models (SCMs). The generator incorporated real-world complexities such as missing values, high-cardinality categories, and conflicting duplicate rows. Training proceeded across three stages, scaling context lengths from 1,024 up to 60,000 rows with up to 100 columns.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment