Skip to content

OpenTabular/PreTab

Repository files navigation

PreTab: Tabular Preprocessing Made Simple

PreTab is a modular, scikit-learn compatible preprocessing library for tabular data. A single Preprocessor detects numerical and categorical columns and turns them into model-ready features, and every strategy is also available as a standalone transformer: splines, neural basis expansions, piecewise-linear encoding, binning, language embeddings, and temporal features. Because it speaks the sklearn API, PreTab drops straight into Pipeline and ColumnTransformer workflows and accepts any sklearn transformer alongside its own.

Why PreTab?

  • Familiar interface. A scikit-learn fit/transform/fit_transform API that drops into existing pipelines and works with both pandas.DataFrame and numpy.ndarray inputs.
  • Automatic feature handling. Feature-type detection and per-feature strategies let you describe intent once instead of wiring transformers by hand.
  • Beyond scaling. Spline bases, neural basis maps, and piecewise-linear encoding turn raw numerical columns into expressive representations.
  • Categoricals done right. Ordinal and one-hot encoding, pretrained language embeddings, and custom binning cover both low- and high-cardinality columns.
  • Composable and extensible. Every strategy is a standalone transformer you can import, compose, or subclass, and any sklearn transformer works out of the box.

πŸƒ Quickstart

import numpy as np
import pandas as pd

from pretab import Preprocessor

df = pd.DataFrame({
    "age": np.random.randint(18, 65, size=100),
    "income": np.random.normal(60_000, 15_000, size=100).astype(int),
    "city": np.random.choice(["Berlin", "Munich", "Hamburg"], size=100),
})
y = np.random.randn(100)

# Global strategies: PLE for numerics, integer codes for categoricals
preprocessor = Preprocessor(numerical_method="ple", categorical_method="int")

X = preprocessor.fit_transform(df, y)   # dict of transformed feature blocks

print({k: v.shape for k, v in X.items()})
# {'num_age': (100, 7), 'num_income': (100, 7), 'cat_city': (100, 1)}

That's it. PreTab detects feature types, fits a strategy per column, and returns ready-to-use arrays.

Works with pandas and numpy. Pass a DataFrame or an array, and PreTab infers numerical vs. categorical columns for you.

Mix strategies per column. Swap the global methods for a feature_preprocessing map, for example {"age": "ple", "income": "rbf", "city": "one-hot"}, and PreTab fits each column with its own strategy in a single pass. See Usage for a full example.

Available Transformers

PreTab groups its transformers into four families. Each one follows the standard fit / transform API and is importable from pretab.transformers.

Splines

Transformer Basis Best for
CubicSplineTransformer B-spline basis Smooth non-linear numerical effects
NaturalCubicSplineTransformer Natural cubic spline Smooth effects with linear tails
PSplineTransformer Penalized B-spline Smoothness with a penalty matrix
TensorProductSplineTransformer Tensor-product spline Interactions between two features
ThinPlateSplineTransformer Thin-plate regression spline Smooth multivariate surfaces

Feature maps

Transformer Basis Best for
RBFExpansionTransformer Radial basis functions Localized, kernel-like features
ReLUExpansionTransformer ReLU basis Piecewise-linear neural features
SigmoidExpansionTransformer Sigmoid basis Smooth saturating features
TanhExpansionTransformer Tanh basis Zero-centered saturating features

Encoding and binning

Transformer Method Best for
PLETransformer Piecewise linear encoding (supervised) Strong numerical encoding for models
CustomBinTransformer Rule- or tree-based binning Discretizing numerical or code values
OneHotFromOrdinalTransformer One-hot from ordinal codes One-hot on pre-encoded categoricals
LanguageEmbeddingTransformer Pretrained language embeddings High-cardinality, semantic columns

Temporal

Transformer Method Best for
CyclicalTimeTransformer Sine/cosine encoding Hour, day, month and cyclic fields
LagFeatureTransformer Lagged values Time-series lag features
RollingStatsTransformer Rolling window statistics Moving averages and rolling summary

Strategy strings. Inside the Preprocessor you select these by short name (for example "ple", "rbf", "one-hot", "pretrained"). See the User Guide for the full list.

πŸ“š Documentation

Full documentation: pretab.readthedocs.io

Quick Links

πŸ› οΈ Installation

Basic installation:

pip install pretab

With language-embedding support:

pip install "pretab[embeddings]"   # adds sentence-transformers

Lightweight by default. The embeddings extra pulls in sentence-transformers and PyTorch, so install it only if you use the pretrained categorical strategy.

Requirements: Python 3.10 to 3.13.

From source:

git clone https://github.com/OpenTabular/PreTab
cd PreTab
pip install -e ".[dev]"

Usage

The Preprocessor

The Preprocessor is the high-level entry point. Set a global strategy per feature type, or override individual columns with feature_preprocessing.

from pretab import Preprocessor

# Per-feature configuration overrides the global defaults
preprocessor = Preprocessor(
    feature_preprocessing={
        "age": "ple",
        "income": "rbf",
        "experience": "quantile",
        "city": "one-hot",
    },
    task="regression",
)

X_dict = preprocessor.fit_transform(df, y)               # {"num_age": ..., "cat_city": ...}
X_array = preprocessor.transform(df, return_array=True)  # single stacked ndarray

preprocessor.get_feature_info(verbose=True)              # inspect resolved strategies

get_feature_info(verbose=True) prints the resolved layout so you can confirm every column at a glance:

feature     kind         pipeline                        dim   cats
-------------------------------------------------------------------
age         numerical    imputer -> minmax -> ple          7      -
income      numerical    imputer -> minmax -> rbf          7      -
experience  numerical    imputer -> minmax -> quantile     1      -
city        categorical  imputer -> onehot -> to_float     4      4

Two output formats. transform returns a dict of feature blocks by default (keys prefixed num_ and cat_), or a single stacked array when you pass return_array=True.

Standalone transformers

Each transformer works on its own and composes with any sklearn estimator.

import numpy as np

from pretab.transformers import PLETransformer

x = np.random.randn(100, 1)
y = np.random.randn(100, 1)

x_ple = PLETransformer(output_dim=15, task="regression").fit_transform(x, y)
assert x_ple.shape[1] == 15

Some transformers are supervised. PLETransformer uses the target y during fit to place its bin edges, so pass y whenever you fit it.

Inside an sklearn Pipeline

Because every transformer follows the sklearn API, you can drop them into a Pipeline or ColumnTransformer.

from sklearn.compose import ColumnTransformer
from sklearn.linear_model import Ridge
from sklearn.pipeline import Pipeline

from pretab.transformers import NaturalCubicSplineTransformer, RBFExpansionTransformer

features = ColumnTransformer([
    ("age", NaturalCubicSplineTransformer(output_dim=10), ["age"]),
    ("income", RBFExpansionTransformer(), ["income"]),
])

model = Pipeline([("features", features), ("ridge", Ridge())])
model.fit(df[["age", "income"]], y)

Spline penalty matrices

Spline transformers expose their penalty matrix for penalized (smoothing) models.

import numpy as np

from pretab.transformers import ThinPlateSplineTransformer

x = np.random.randn(100, 1)

tp = ThinPlateSplineTransformer(output_dim=15)
x_tp = tp.fit_transform(x)
penalty = tp.get_penalty_matrix()   # (output_dim, output_dim) smoothing penalty

Advanced Features

Automatic feature-type detection

By default PreTab inspects each column and classifies it as numerical or categorical. String and object columns are treated as categorical, low-cardinality integer columns are categorical, and integer columns with enough distinct values stay numerical. Tune the behavior with cat_cutoff and treat_all_integers_as_numerical.

preprocessor = Preprocessor(
    treat_all_integers_as_numerical=False,
    cat_cutoff=0.03,
)

Language embeddings for categoricals

The pretrained strategy encodes categorical values with a sentence-transformer, which helps with high-cardinality or semantically rich columns.

preprocessor = Preprocessor(
    feature_preprocessing={"job_title": "pretrained"},
)

Optional dependency. Install with pip install "pretab[embeddings]" before using the pretrained strategy.

Custom binning

CustomBinTransformer supports both rule-based edges and tree-based bins learned from the target.

preprocessor = Preprocessor(
    numerical_method="custombin",
    output_dim=32,
)

πŸ“„ License

PreTab is licensed under the MIT License. See LICENSE for details.

🀝 Contributing

Contributions are welcome, whether you are fixing bugs, adding transformers, or improving the docs. See the Contributing Guide to get started, and please follow our Code of Conduct.

git clone https://github.com/OpenTabular/PreTab
cd PreTab
pip install -e ".[dev]"

πŸ“ž Support

About

PreTab is a flexible, scikit-learn-based library for advanced tabular preprocessing, feature expansion, and embedding integration.

Topics

Resources

Code of conduct

Stars

Watchers

Forks

Releases

Packages

Used by

Contributors

Languages