Skip to content

Preprocessing

AutoEncoder

AutoEncoder(*, cardinality_threshold: int = 20, continuous_scaler: Literal['standard', 'minmax'] = 'standard', handle_unknown: Literal['ignore', 'raise'] = 'ignore', mean_center_ohe: bool = False, column_types: Sequence[Literal['categorical', 'continuous', 'ordinal']] | None = None, binary: Literal['single', 'both'] = 'single', drop: Literal['first'] | None = None, block_weighting: Literal['none', 'equal_variance'] = 'none', column_levels: Sequence[Sequence[Any] | None] | None = None)

Bases: BaseEstimator, TransformerMixin

Column-wise router that applies missing-aware OHE or scaling.

After fitting, feature_groups_ gives the input column of every output column and feature_kinds_ the kind of every input column ("categorical", "continuous" or "ordinal"). Ordinal columns are never inferred: list them in column_types, and they are coded as scaled numeric scores (see :class:MissingAwareOrdinalEncoder) with the continuous_scaler. binary, drop and block_weighting are passed to the dense one-hot encoder (see :class:MissingAwareOneHotEncoder). Dense fits expose encoding_schema_ for encoding-aware categorical CV; sparse fits leave it None (component CV accepts dense input only). column_levels declares external nominal vocabularies or ordinal level orders; leave continuous columns and training-inferred vocabularies None.

MissingAwareOneHotEncoder

MissingAwareOneHotEncoder(*, handle_unknown: Literal['ignore', 'raise'] = 'ignore', mean_center: bool = False, dtype: type = float, binary: Literal['single', 'both'] = 'single', drop: Literal['first'] | None = None, block_weighting: Literal['none', 'equal_variance'] = 'none', categories: Sequence[Sequence[Any] | None] | None = None)

Bases: BaseEstimator, TransformerMixin

One-hot encode categorical columns while respecting missing values.

Args: handle_unknown: "ignore" encodes unseen categories as all zeros; "raise" raises. mean_center: Subtract each output column's observed mean. dtype: Output dtype. binary: "single" codes a two-level variable as one indicator for its second level (a reference-level drop); "both" codes it with one indicator per level, like every other variable. The v1/EHS analyses one-hot encoded both levels. drop: "first" drops the indicator of each variable's first category in categories_ (the first level seen), for every variable with at least two levels, so a complete block no longer sums to one; None (default) keeps every level. inverse_transform restores the dropped level as one minus the sum of the others. block_weighting: "equal_variance" scales each variable's block so its total observed variance at fit time is one, so every variable contributes comparable variance whatever its number of levels; "none" (default) leaves the indicators unscaled. The weights are in block_weights_ and are undone by inverse_transform. categories: Optional externally declared category order per column; None entries infer levels from training observations only.

After fitting, feature_groups_ gives the input column of every output column and feature_kinds_ the kind of every input column, so model selection can hold out whole variables (#250). encoding_schema_ also records category order, reference drop, fitted means and weights for categorical decoding and scoring (#266). Centering is fitted once; transforming new data retains those means.

MissingAwareSparseOneHotEncoder

MissingAwareSparseOneHotEncoder(*, handle_unknown: Literal['ignore', 'raise'] = 'ignore', mean_center: bool = False, dtype: type = float)

Bases: BaseEstimator, TransformerMixin

Sparse one-hot encoder for categorical columns.

Assumptions: - Input is sparse (CSR/CSC) with one column. - Observed entries are stored; missing entries are absent. - Categories must be numeric to round-trip through sparse matrices. - mean_center adjusts stored values per category without densifying.

MissingAwareStandardScaler

MissingAwareStandardScaler()

Bases: _BaseScaler

Standardize continuous columns ignoring missing entries.

MissingAwareMinMaxScaler

MissingAwareMinMaxScaler()

Bases: _BaseScaler

Scale features to [0, 1] range while ignoring missing entries.

MissingAwareLogTransformer

MissingAwareLogTransformer(*, offset: float = 1.0)

Apply log1p (or log(x + offset)) while preserving NaN entries.

This is a pre-scaling transform: compress right-skewed features before standardisation so that the standard deviation better reflects the bulk of the data rather than extreme tails.

Args: offset: Additive constant before taking the log. The default 1.0 gives the standard log1p transform.

fit

fit(x: ndarray, mask: Mask | None = None) -> MissingAwareLogTransformer

Record input width (stateless transform).

transform

transform(x: ndarray, mask: Mask | None = None) -> np.ndarray

Apply log(x + offset) to observed entries.

inverse_transform

inverse_transform(z: ndarray, mask: Mask | None = None) -> np.ndarray

Reverse via exp(z) - offset.

MissingAwarePowerTransformer

MissingAwarePowerTransformer(*, standardize: bool = True)

Yeo-Johnson variance-stabilising transform preserving NaN entries.

Supports mixed-sign data. The transform is applied element-wise with a per-column lmbda parameter estimated by maximum likelihood on observed values.

Args: standardize: If True (default), z-score the transformed output so each feature has zero mean and unit variance.

fit

fit(x: ndarray, mask: Mask | None = None) -> MissingAwarePowerTransformer

Estimate per-column Yeo-Johnson lambda by profile MLE.

transform

transform(x: ndarray, mask: Mask | None = None) -> np.ndarray

Apply Yeo-Johnson transform and optional standardisation.

inverse_transform

inverse_transform(z: ndarray, mask: Mask | None = None) -> np.ndarray

Reverse transform: un-standardise then invert Yeo-Johnson.

MissingAwareWinsorizer

MissingAwareWinsorizer(*, lower_quantile: float = 0.01, upper_quantile: float = 0.99)

Clip features at fitted percentiles while preserving NaN entries.

Winsorization prevents outlier-driven variance inflation before scaling, so that z-scoring better equalises feature scales.

Args: lower_quantile: Lower clipping quantile (default 0.01). upper_quantile: Upper clipping quantile (default 0.99).

Note: This transform is lossy — inverse_transform is a no-op (returns data unchanged) because the original tail values cannot be recovered.

fit

fit(x: ndarray, mask: Mask | None = None) -> MissingAwareWinsorizer

Compute per-column clipping bounds from observed values.

transform

transform(x: ndarray, mask: Mask | None = None) -> np.ndarray

Clip observed entries to fitted bounds.

inverse_transform

inverse_transform(z: ndarray, mask: Mask | None = None) -> np.ndarray

No-op: winsorization is lossy.

Returns a copy of the input unchanged.

check_data

check_data(x: ndarray, *, mask: Mask | None = None, column_names: Sequence[str] | None = None, skewness_threshold: float = 2.0, outlier_mad_threshold: float = 5.0, near_zero_var_eps: float = 1e-10, missing_fraction_warn: float = 0.5, warn: bool = False, mcar_test: bool = False) -> DataReport

Run preflight diagnostics on a data matrix before VBPCA fitting.

Checks focus on scale comparability — conditions that cause individual features to dominate the decomposition — rather than distributional shape. Categorical encodings are also checked: one-hot blocks are detected (see suggested_feature_groups), and rare levels, single-indicator binary variables and ordinal-looking integer codes are flagged.

Args: x: Data matrix of shape (n_samples, n_features). mask: Optional boolean observation mask with the same shape as x. Entries are included only when the mask is true and the data value is finite. column_names: Optional feature names for readable messages. skewness_threshold: Absolute skewness above which a feature is flagged (default 2.0). outlier_mad_threshold: Number of MADs from the median beyond which an entry is considered an outlier (default 5.0). near_zero_var_eps: Variance threshold below which a feature is flagged as near-zero-variance (default 1e-10). missing_fraction_warn: Per-feature and per-row missing fraction above which a warning is emitted (default 0.5). warn: If True, also emit :func:warnings.warn for each issue. mcar_test: If True, screen for missingness that depends on observed values: each continuous column is compared between rows where another column is missing and rows where it is observed (Welch's t-test, Bonferroni-corrected). A flagged pair is evidence against MCAR; no flag does not establish MCAR.

Returns: A :class:DataReport with warnings, per-feature summary, and suggested pre-transforms.

DataReport dataclass

DataReport(warnings: list[str] = list(), summary: dict[str, dict[str, float]] = dict(), suggested_pretransforms: dict[str | int, str] = dict(), passed: bool = True, suggested_feature_groups: list[int] | None = None, missingness: dict[str, Any] = dict())

Result of :func:check_data preflight validation.

Attributes: warnings: Human-readable diagnostic messages. summary: Per-feature statistics dictionary. suggested_pretransforms: Mapping of column index (or name) to a suggested transform string (e.g. "log1p"). passed: True when no warnings were raised. suggested_feature_groups: Variable label per column when one-hot blocks were detected (pass as CVConfig(feature_groups=...)), otherwise None. missingness: Overall and per-row missing fractions, complete rows, the number of distinct missingness patterns, rows above the missing-fraction threshold and, with mcar_test=True, the pairs tested and flagged by the MCAR screen.

Fitted encoding metadata

EncodingSchema dataclass

EncodingSchema(variables: tuple[EncodedVariable, ...])

Snapshot of dense encoding metadata, in original variable order.

Obtain this from a fitted encoder's encoding_schema_. Category order, centering and weighting belong to that fit and must match the encoded data supplied to component CV. This schema does not fit preprocessing per fold.

Attributes: variables: One description per original variable, including numeric variables that are excluded from nominal categorical scoring.

feature_groups property

feature_groups: tuple[int, ...]

Original variable index of each encoded feature.

validate

validate(n_features: int) -> None

Check that metadata describes every encoded feature exactly once.

Args: n_features: Number of rows of the model's features-by-samples input.

Raises: ValueError: If columns, kind, category dimensions or affine metadata are inconsistent with the supplied matrix.

EncodedVariable dataclass

EncodedVariable(kind: Literal['categorical', 'continuous', 'ordinal'], columns: tuple[int, ...], categories: tuple[object, ...] = (), reference: int | None = None, offset: tuple[float, ...] = (), weight: float = 1.0)

Describe one original variable in an encoded feature matrix.

Attributes: kind: Original variable kind; only categorical variables receive nominal categorical scores. columns: Encoded feature indices, in category order with the reference omitted when applicable. categories: Fitted nominal category or ordinal level order. reference: Index of an omitted category, or None for full coding. offset: Unweighted means (or numeric minima) subtracted before scaling. weight: Positive multiplier applied to the block or numeric column.

decode

decode(block: ndarray) -> np.ndarray

Undo a categorical block's affine encoding and reference drop.

Args: block: Encoded features-by-samples values for this variable.

Returns: Full category-by-samples values before clipping or normalization.

Raises: ValueError: If the variable is not categorical or dimensions differ.

decode_numeric

decode_numeric(block: ndarray) -> np.ndarray

Restore continuous values or unrounded ordinal level-index scores.

Args: block: Encoded features-by-samples values for this numeric variable.

Returns: Values in continuous input units or ordinal level-index units.

Raises: ValueError: If numeric affine metadata or dimensions are unavailable.