agentsclimarketplace

Feature engineering

Skill param087/agent-ml-skills/skills/feature-engineering

Production-grade Machine Learning, Data Science & MLOps skills for AI coding agents (Codex, Claude Code, Cursor, OpenCode). One npx command to install.

Install
npx -y skills add param087/agent-ml-skills --skill feature-engineering

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when creating, encoding, scaling, or selecting features for ML models. Covers categorical encoding, numeric transforms, datetime/text/aggregation features, and leakage-safe target encoding.

SKILL.md

3.1 KB, as published. Nobody here has run it

Feature Engineering

Overview

Feature engineering is where most model performance is won or lost. The aim is to express the signal in a form the model can use, while never letting information from the target or the test set leak into a feature.

When to use

  • After cleaning, before/iterating with modeling.
  • A model plateaus and you suspect under-expressed signal.
  • You have raw datetime, text, or relational data to turn into columns.

Encoding categoricals

CardinalityEncoderNotes
Low (<15), tree modelOne-hot or native categoricalLightGBM/CatBoost handle natively
Low, linear modelOne-hotDrop-first to avoid collinearity
High (>15)Target/leave-one-out encodingMust be cross-fitted to avoid leakage
Ordinal meaningOrdinal mapPreserve order (low<med<high)

Numeric transforms

  • Skewed positive valueslog1p or Box-Cox/Yeo-Johnson.
  • ScalingStandardScaler for linear/NN, none needed for trees.
  • Binning → only when the relationship is genuinely non-monotonic.
  • Interactions → products/ratios of features with domain meaning (e.g., price / sqft).

Datetime features

ts = df["event_time"]
df["hour"] = ts.dt.hour
df["dayofweek"] = ts.dt.dayofweek
df["is_weekend"] = ts.dt.dayofweek.ge(5).astype(int)
df["month"] = ts.dt.month
# Cyclical encoding so 23:00 and 00:00 are close
import numpy as np
df["hour_sin"] = np.sin(2 * np.pi * df["hour"] / 24)
df["hour_cos"] = np.cos(2 * np.pi * df["hour"] / 24)

Leakage-safe target encoding

Target encoding must be fit out-of-fold, never on the rows it encodes:

from sklearn.model_selection import KFold
import numpy as np

def target_encode_oof(train, col, target, n_splits=5, smoothing=10):
    oof = np.zeros(len(train))
    prior = train[target].mean()
    kf = KFold(n_splits=n_splits, shuffle=True, random_state=42)
    for tr_idx, val_idx in kf.split(train):
        agg = train.iloc[tr_idx].groupby(col)[target].agg(["mean", "count"])
        smooth = (agg["mean"] * agg["count"] + prior * smoothing) / (agg["count"] + smoothing)
        oof[val_idx] = train.iloc[val_idx][col].map(smooth).fillna(prior).values
    return oof

Pitfalls

  • Target encoding fit on all rows → severe leakage, inflated CV, collapse in production.
  • Scaling fit on full data → use Pipeline so scaler fits on train folds only.
  • Aggregations over the whole timeline in time-series → only use past data (rolling windows with proper shift).
  • Creating thousands of features then trusting noisy importance — prefer a few well-motivated features + regularization.

Hand-off

Deliver a documented feature set (name, source, rationale) and ensure all transforms are wrapped in a fitted Pipeline for the model-evaluation and model-serving skills to reuse.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.