Preprocessing¶
AutoEncoder ¶
AutoEncoder(*, cardinality_threshold: int = 20, continuous_scaler: Literal['standard', 'minmax'] = 'standard', handle_unknown: Literal['ignore', 'raise'] = 'ignore', mean_center_ohe: bool = False, column_types: Sequence[Literal['categorical', 'continuous', 'ordinal']] | None = None, binary: Literal['single', 'both'] = 'single', drop: Literal['first'] | None = None, block_weighting: Literal['none', 'equal_variance'] = 'none', column_levels: Sequence[Sequence[Any] | None] | None = None)
Bases: BaseEstimator, TransformerMixin
Column-wise router that applies missing-aware OHE or scaling.
After fitting, feature_groups_ gives the input column of every
output column and feature_kinds_ the kind of every input column
("categorical", "continuous" or "ordinal"). Ordinal columns
are never inferred: list them in column_types, and they are coded as
scaled numeric scores (see :class:MissingAwareOrdinalEncoder) with the
continuous_scaler. binary, drop and
block_weighting are passed to the dense one-hot encoder (see
:class:MissingAwareOneHotEncoder).
Dense fits expose encoding_schema_ for encoding-aware categorical CV;
sparse fits leave it None (component CV accepts dense input only).
column_levels declares external nominal vocabularies or ordinal level
orders; leave continuous columns and training-inferred vocabularies None.
MissingAwareOneHotEncoder ¶
MissingAwareOneHotEncoder(*, handle_unknown: Literal['ignore', 'raise'] = 'ignore', mean_center: bool = False, dtype: type = float, binary: Literal['single', 'both'] = 'single', drop: Literal['first'] | None = None, block_weighting: Literal['none', 'equal_variance'] = 'none', categories: Sequence[Sequence[Any] | None] | None = None)
Bases: BaseEstimator, TransformerMixin
One-hot encode categorical columns while respecting missing values.
Args:
handle_unknown: "ignore" encodes unseen categories as all zeros;
"raise" raises.
mean_center: Subtract each output column's observed mean.
dtype: Output dtype.
binary: "single" codes a two-level variable as one indicator for
its second level (a reference-level drop); "both" codes it
with one indicator per level, like every other variable. The
v1/EHS analyses one-hot encoded both levels.
drop: "first" drops the indicator of each variable's first
category in categories_ (the first level seen), for every
variable with at least two levels, so a complete block no longer
sums to one; None (default) keeps every level.
inverse_transform restores the dropped level as one minus the
sum of the others.
block_weighting: "equal_variance" scales each variable's block so
its total observed variance at fit time is one, so every variable
contributes comparable variance whatever its number of levels;
"none" (default) leaves the indicators unscaled. The weights
are in block_weights_ and are undone by inverse_transform.
categories: Optional externally declared category order per column;
None entries infer levels from training observations only.
After fitting, feature_groups_ gives the input column of every
output column and feature_kinds_ the kind of every input column, so
model selection can hold out whole variables (#250).
encoding_schema_ also records category order, reference drop, fitted
means and weights for categorical decoding and scoring (#266). Centering
is fitted once; transforming new data retains those means.
MissingAwareSparseOneHotEncoder ¶
MissingAwareSparseOneHotEncoder(*, handle_unknown: Literal['ignore', 'raise'] = 'ignore', mean_center: bool = False, dtype: type = float)
Bases: BaseEstimator, TransformerMixin
Sparse one-hot encoder for categorical columns.
Assumptions:
- Input is sparse (CSR/CSC) with one column.
- Observed entries are stored; missing entries are absent.
- Categories must be numeric to round-trip through sparse matrices.
- mean_center adjusts stored values per category without
densifying.
MissingAwareStandardScaler ¶
Bases: _BaseScaler
Standardize continuous columns ignoring missing entries.
MissingAwareMinMaxScaler ¶
Bases: _BaseScaler
Scale features to [0, 1] range while ignoring missing entries.
MissingAwareLogTransformer ¶
Apply log1p (or log(x + offset)) while preserving NaN entries.
This is a pre-scaling transform: compress right-skewed features before standardisation so that the standard deviation better reflects the bulk of the data rather than extreme tails.
Args:
offset: Additive constant before taking the log. The default
1.0 gives the standard log1p transform.
MissingAwarePowerTransformer ¶
Yeo-Johnson variance-stabilising transform preserving NaN entries.
Supports mixed-sign data. The transform is applied element-wise
with a per-column lmbda parameter estimated by maximum
likelihood on observed values.
Args:
standardize: If True (default), z-score the transformed
output so each feature has zero mean and unit variance.
MissingAwareWinsorizer ¶
Clip features at fitted percentiles while preserving NaN entries.
Winsorization prevents outlier-driven variance inflation before scaling, so that z-scoring better equalises feature scales.
Args:
lower_quantile: Lower clipping quantile (default 0.01).
upper_quantile: Upper clipping quantile (default 0.99).
Note:
This transform is lossy — inverse_transform is a no-op
(returns data unchanged) because the original tail values
cannot be recovered.
check_data ¶
check_data(x: ndarray, *, mask: Mask | None = None, column_names: Sequence[str] | None = None, skewness_threshold: float = 2.0, outlier_mad_threshold: float = 5.0, near_zero_var_eps: float = 1e-10, missing_fraction_warn: float = 0.5, warn: bool = False, mcar_test: bool = False) -> DataReport
Run preflight diagnostics on a data matrix before VBPCA fitting.
Checks focus on scale comparability — conditions that cause
individual features to dominate the decomposition — rather than
distributional shape. Categorical encodings are also checked: one-hot
blocks are detected (see suggested_feature_groups), and rare levels,
single-indicator binary variables and ordinal-looking integer codes are
flagged.
Args:
x: Data matrix of shape (n_samples, n_features).
mask: Optional boolean observation mask with the same shape as x.
Entries are included only when the mask is true and the data value
is finite.
column_names: Optional feature names for readable messages.
skewness_threshold: Absolute skewness above which a feature is
flagged (default 2.0).
outlier_mad_threshold: Number of MADs from the median beyond
which an entry is considered an outlier (default 5.0).
near_zero_var_eps: Variance threshold below which a feature is
flagged as near-zero-variance (default 1e-10).
missing_fraction_warn: Per-feature and per-row missing fraction
above which a warning is emitted (default 0.5).
warn: If True, also emit :func:warnings.warn for each
issue.
mcar_test: If True, screen for missingness that depends on
observed values: each continuous column is compared between rows
where another column is missing and rows where it is observed
(Welch's t-test, Bonferroni-corrected). A flagged pair is
evidence against MCAR; no flag does not establish MCAR.
Returns:
A :class:DataReport with warnings, per-feature summary, and
suggested pre-transforms.
DataReport
dataclass
¶
DataReport(warnings: list[str] = list(), summary: dict[str, dict[str, float]] = dict(), suggested_pretransforms: dict[str | int, str] = dict(), passed: bool = True, suggested_feature_groups: list[int] | None = None, missingness: dict[str, Any] = dict())
Result of :func:check_data preflight validation.
Attributes:
warnings: Human-readable diagnostic messages.
summary: Per-feature statistics dictionary.
suggested_pretransforms: Mapping of column index (or name) to a
suggested transform string (e.g. "log1p").
passed: True when no warnings were raised.
suggested_feature_groups: Variable label per column when one-hot
blocks were detected (pass as CVConfig(feature_groups=...)),
otherwise None.
missingness: Overall and per-row missing fractions, complete rows,
the number of distinct missingness patterns, rows above the
missing-fraction threshold and, with mcar_test=True, the
pairs tested and flagged by the MCAR screen.
Fitted encoding metadata¶
EncodingSchema
dataclass
¶
Snapshot of dense encoding metadata, in original variable order.
Obtain this from a fitted encoder's encoding_schema_. Category order,
centering and weighting belong to that fit and must match the encoded data
supplied to component CV. This schema does not fit preprocessing per fold.
Attributes: variables: One description per original variable, including numeric variables that are excluded from nominal categorical scoring.
feature_groups
property
¶
Original variable index of each encoded feature.
validate ¶
Check that metadata describes every encoded feature exactly once.
Args: n_features: Number of rows of the model's features-by-samples input.
Raises: ValueError: If columns, kind, category dimensions or affine metadata are inconsistent with the supplied matrix.
EncodedVariable
dataclass
¶
EncodedVariable(kind: Literal['categorical', 'continuous', 'ordinal'], columns: tuple[int, ...], categories: tuple[object, ...] = (), reference: int | None = None, offset: tuple[float, ...] = (), weight: float = 1.0)
Describe one original variable in an encoded feature matrix.
Attributes:
kind: Original variable kind; only categorical variables receive
nominal categorical scores.
columns: Encoded feature indices, in category order with the reference
omitted when applicable.
categories: Fitted nominal category or ordinal level order.
reference: Index of an omitted category, or None for full coding.
offset: Unweighted means (or numeric minima) subtracted before scaling.
weight: Positive multiplier applied to the block or numeric column.
decode ¶
Undo a categorical block's affine encoding and reference drop.
Args: block: Encoded features-by-samples values for this variable.
Returns: Full category-by-samples values before clipping or normalization.
Raises: ValueError: If the variable is not categorical or dimensions differ.
decode_numeric ¶
Restore continuous values or unrounded ordinal level-index scores.
Args: block: Encoded features-by-samples values for this numeric variable.
Returns: Values in continuous input units or ordinal level-index units.
Raises: ValueError: If numeric affine metadata or dimensions are unavailable.