Skip to content

Feature Engineering

Quick overview Feature engineering is the process of transforming raw data into numerical features a model can digest — missing value handling, categorical encoding, numerical scaling, feature transforms, combination and selection. This article provides a complete checklist and the iterative principle of "start with a simple baseline, then add features based on evidence."

Feature Engineering ​

Concept Definition: The "Nutrition Recipe" for Models ​

A feature is the input raw material for a model. Feature engineering is the entire process of transforming raw data (raw tables, logs, text, images) into numerical features that a model can effectively use. With the same data and model, good vs. poor feature engineering can produce differences of several orders of magnitude — in the classic ML era, there was a famous saying: "feature engineering sets the ceiling, and the model just approaches that ceiling."

After deep learning emerged, neural networks can learn features automatically (representation learning), but feature engineering is far from obsolete:

  • Tabular data (finance, e-commerce, healthcare) remains the dominant business scenario, and tree models / linear models all rely on hand-crafted features;
  • Data preprocessing (normalization, encoding, augmentation) in deep models is still fundamentally feature engineering;
  • Domain knowledge (e.g., "number of purchases in the user's last 7 days") is prior knowledge the model can never learn on its own — it must come from humans.

A Full Feature Engineering Checklist ​

In processing order, feature engineering is divided into five layers:

Raw data → 1. Cleaning & missing values → 2. Scaling & transforms → 3. Categorical encoding → 4. Feature construction → 5. Feature selection

1. Missing Value Handling ​

StrategyApproachSuitable For
DeletionDrop rows/columnsExtremely high missing ratio (>80%) or irrelevant to target
Statistical imputationMean/median/modeRandomly missing numerical data (simple baseline)
Model-based imputationKNN/regression to predict missing valuesMissing pattern correlates with other features
Flag missingAdd an "is missing" indicator columnMissingness itself carries information (e.g., "no income filled in")
Domain-based imputationFill using business rules (e.g., 0, default value)Clear business semantics

Three Pitfalls of Missing Value Handling

  1. Mean imputation underestimates variance — the distribution is flattened, and the model learns incorrect uncertainty;
  2. Imputation statistics must only be computed from the training set (otherwise data leakage);
  3. "Missing" itself is often a signal — e.g., users with "no occupation info" often have higher risk; adding an is_missing column often works wonders.

2. Numerical Scaling and Transforms ​

MethodFormulaCharacteristicsSuitable For
Standardization (Z-score)(x-μ)/σMean 0, variance 1, preserves distribution shapeLinear models, SVM, neural networks (default)
Normalization (Min-Max)(x-min)/(max-min)Maps to [0,1], sensitive to outliersWhen a fixed value range is needed
Robust scaling(x-median)/IQRResistant to outliersData with extreme values
Log/power transformslog(x), x^λ (Box-Cox)Compress long tails, correct right-skewIncome, price, click counts, and other skewed data

Tree models don't need scaling (decision trees split by thresholds, independent of scale) — one of the reasons tree models are "convenient" for tabular data. Linear models, KNN, SVM, and neural networks must scale, otherwise features with large scales will dominate distance/gradient computations.

3. Categorical Encoding ​

MethodApproachCharacteristicsSuitable For
Label EncodingMap categories to integers 0,1,2…Introduces artificial ordinal relationshipOrdered categories (low/medium/high)
One-Hot EncodingOne 0/1 column per categoryNo ordinal assumption, but column count explodesUnordered low-cardinality categories (default)
Target EncodingEncode by the target mean for that categoryHigh information density, but prone to overfitting (needs encoding within cross-validation)High-cardinality categories (cities, user IDs)
Frequency encodingEncode by category occurrence countSimple and robust, captures popularityHigh-cardinality categories
EmbeddingLearn dense vector representationsStrongest but heaviestDeep models

The conflict between one-hot and high cardinality is the core pain point of tabular modeling: 1000 cities → 1000 one-hot columns, sparse and bloated. In this case, target encoding / frequency encoding / embeddings are better solutions.

4. Feature Construction (The Source of Domain Features) ​

  • Statistical features: aggregation (mean/max/min/std/count/percentage) — "mean spending amount in the user's last 30 days";
  • Time features: year/month/day/weekday/holiday/time-since-last-event/time-of-day — seasonal and periodic signals;
  • Combined features: feature cross (e.g., "city × product category") — interaction priors the model can't learn, partially auto-learned by tree models;
  • Ratio features: conversion rate, average order value, growth rate — turning absolute numbers into relative quantities with stronger business meaning;
  • Text/image features: TF-IDF, embeddings, statistics (length, word count).

The core principle of feature construction: start from the business problem. "In this business scenario, what metrics do experts look at when making decisions?" — feature-izing expert decision criteria is the most effective source of features.

5. Feature Selection: Removing the Useless Ones ​

Too many features (especially after one-hot encoding, tens of thousands of columns) introduce noise, slow training, and harm interpretability. Three categories of methods:

CategoryMethodCharacteristics
FilterVariance threshold, correlation coefficient, chi-square, mutual informationFast, independent of the model, ignores interactions
WrapperForward/backward selection, recursive feature elimination (RFE)Considers model performance, computationally expensive
EmbeddedL1 regularization (Lasso), tree feature importance, SHAPAuto-selects during training, the practice favorite

Practice favorite: tree model feature importance + correlation deduplication — fast, stable, and consistent with the model. L1 regularization has built-in sparsity; see Overfitting and Regularization.

Feature Engineering for Different Data Types ​

Data TypeCore Feature Engineering
TabularAll of the above: missing values, scaling, encoding, aggregation, combination
Time seriesLag features, rolling statistics (rolling mean/std), differencing, time decomposition (trend/seasonal/residual)
TextCleaning (tokenization, stopword removal), TF-IDF/BOW, word/sentence embeddings, length and structural features
ImagesNormalization, uniform sizing, data augmentation (flip/crop/color jitter), pre-trained feature extraction
Categories/IDsFrequency encoding, target encoding, embedding, graph-structured features (user-item interactions)

Principles and Workflow of Feature Engineering ​

Iterative principle: start with a simple baseline, then add features based on evidence

The correct workflow is not "get all features right the first time," but:

  1. Start with a minimal viable feature set (raw numerical columns + simple encoding) and establish a baseline;
  2. Use error analysis and feature importance to find improvement directions (which error types are most common? which features are useful?);
  3. Add one feature group at a time and confirm it actually boosts metrics on the validation set before adding the next;
  4. Features only growing without removal leads to overfitting and maintenance disaster — do feature selection regularly.

"I built 200 features" isn't an achievement; "every feature has evidence for its existence" is.

Train/Inference Consistency (The Root of Feature Leakage) ​

The ultimate trap of feature engineering is inconsistency between training features and online inference features:

  • Used "future information" during training (e.g., using the monthly mean to predict the beginning of the month) → the feature is unavailable at inference;
  • Imputation/scaling statistics were computed on the full dataset during training, but online only gets new samples;
  • Feature delay: online features arrive later than the time point assumed in training.

Countermeasure: every statistic for feature computation should depend only on "data visible at the prediction moment"; use a feature store to unify the feature computation logic for training and inference. See Data and Data Engineering and MLOps and Model Deployment.

Feature Engineering vs. Representation Learning ​

A relationship worth clarifying: deep learning uses representation learning to automatically learn features from raw data (CNNs learn edges → parts → objects, Transformers learn word meaning → syntax → semantics), automating part of feature engineering. But:

  • Deep models learn "general representations," but domain priors (e.g., "weekends have higher conversion") still need humans to provide;
  • On structured tabular data, representation learning has not yet universally beaten the combination of hand-crafted features + tree models;
  • In practice, the two are often combined: pre-trained models extract features (embeddings) → feed into tree models / linear models.

So the conclusion is not "feature engineering is obsolete," but rather the focus of feature engineering shifts from "hand-designing all features" to "designing data pipelines + selecting/composing auto-learned features + injecting domain priors."

Tradeoffs ​

  • Number of features vs. model complexity: more features require the model to guard against overfitting (L1, depth, regularization). The goal of feature engineering is not "more" but "useful."
  • Hand-crafted vs. auto representations: small data, need interpretability → hand-crafted; large data, unstructured → representation learning; the middle ground → hybrid.
  • Feature engineering cost vs. return: an excellent domain feature (e.g., "days since last purchase") can boost metrics more than ten rounds of hyperparameter tuning — do feature engineering first, then tune, in that order.
  • Interpretability: the more domain-aligned features (naming, explainability) the more trustworthy the model; embedding features are hard to explain but stronger. Trade off per scenario.

Further Reading ​

References ​