Skip to content

Data and Data Engineering

Quick overview 80% of an ML project's time is spent on data. This article systematically covers data collection, cleaning, labeling, versioning, data leakage, plus feature stores and data pipelines — the full engineering scope from "getting data" to "ready to train."

Data and Data Engineering ​

Concept Definition: Data Sets the Model's Ceiling ​

A saying circulates in the machine learning community: "garbage in, garbage out." No matter how strong the model is, it can't fill the gaps left by bad data. The real-world time allocation of ML projects also confirms this: data preparation and cleaning typically take 60–80% of the time, while modeling itself is only a small part.

In the ML context, data engineering refers to all the work of transforming raw data into high-quality, reproducible, versioned training data and inference data — collection, cleaning, labeling, splitting, pipelines, storage.

The Data Lifecycle Panorama ​

text
Business systems / logs / external data
       │ 1. Collect (logs, DBs, APIs, crawlers, purchased)
       ▼
   Raw data storage (data lake / warehouse)
       │ 2. Clean (missing/outliers/duplicates/formats)
       ▼
   Clean data (versionable)
       │ 3. Label (needed for supervised learning)
       ▼
   Train/validation/test splits (leakage-proof)
       │ 4. Feature engineering (see Feature Engineering chapter)
       ▼
   Feature store (unified for training/inference)
       │ 5. Train → Evaluate → Deploy (see MLOps chapter)
       ▼
   Monitoring → Data drift → Re-collect/re-label (closed loop)

Collection and Storage ​

Three Sources of Data ​

SourceExamplesCaveats
Business systemsTransactions, clicks, behavior logsData definitions change as products evolve (schema drift)
External dataPublic datasets, APIs, crawlersCopyright, compliance (GDPR), stability
Labeled dataHuman annotation, crowdsourcing, remote supervisionLabeling cost and quality are the core bottlenecks

Storage Selection ​

  • Data lake: store in raw format (object storage), flexible but messy;
  • Data warehouse: structured, queried fast after ETL (BigQuery, Snowflake, ClickHouse);
  • Feature store: unified training/inference features (Feast, Tecton, ByteDance FeatureStore) — solving the #1 engineering incident: "training features differ from online features."

Why feature stores matter

During training, you used the feature "user's average spending in the last 30 days." How do you guarantee the same definition and same time point during online inference? A feature store unified the definition, computation, and freshness of features; both training and inference pull from it — this is the watershed of MLOps maturity. See MLOps and Model Deployment.

Data Cleaning: Eight Forms of Dirty Data ​

Dirty DataExampleHandling
Missing valuesUser didn't fill in ageImpute/delete/flag; see Feature Engineering
OutliersA house price of 100 millionDomain judgment + percentile truncation
DuplicatesTwo records for the same orderDeduplicate (watch for legitimate "natural duplicates")
Format inconsistencyDates: 2023/1/1 vs 01-01-2023Unified parsing and standardization
Unit inconsistencyMeters vs. feet, yuan vs. dollarsUnit normalization (easy to get wrong!)
Encoding chaosChinese garbled text, emojisUnified encoding (UTF-8)
Logical contradictionsAge 5 but has a driver's licenseRule validation
Data driftPurchasing distribution changes dramatically before/after Double 11Monitoring + per-group modeling

The principle of cleaning is "auditability": every cleaning decision (what was deleted, what was imputed) must be traceable and re-runnable, otherwise you'll never know why the model "suddenly got better/worse" — it might be that the cleaning logic changed, not that the model improved.

Data Labeling: The Bottleneck of Supervised Learning ​

Labeling quality directly determines the model's ceiling. Three common approaches:

ApproachCostQualitySuitable For
Human labeling (in-house / crowdsourcing)HighHighCore business data
Remote supervision (rule-based / knowledge base auto-labeling)LowMedium (noisy)Large-scale initial labeling
Weak supervision / pre-labeling + human reviewMediumMedium-highThe mainstream approach of labeling tools

Three disciplines of labeling:

  1. Labeling guidelines first: ambiguous cases ("does this image count as a cat?") must have rules defined beforehand, otherwise different annotators do their own thing;
  2. Consistency checks: have multiple annotators label the same batch of samples and compute labeling consistency (Cohen's Kappa); low consistency means the guidelines are unclear;
  3. Annotator bias: differences in standards among annotators become "noisy labels" that the model learns — consider majority voting and annotator review.

Data Splitting and Data Leakage ​

The Iron Rules of Splitting ​

  • Never touch the test set during iteration (see Model Evaluation and Validation);
  • Split criteria must match the real prediction scenario:
    • i.i.d. data → random split;
    • Time series → split by time (train on past, test on future);
    • Grouped data (same user / same patient) → group-level split (GroupSplit), otherwise samples from the same user crossing splits = leakage.

Six Sources of Data Leakage ​

Data leakage is the #1 accident in ML engineering: test information leaks into training, offline metrics are inflated, and production fails. Common sources:

Leakage SourceExample
Global statisticsStandardization/imputation computed on full data (including test set)
Future informationUsing "full-month statistics" to predict the beginning of the month (temporal leakage)
Target leakageFeatures containing the "result" (e.g., using "was refunded" to predict "will churn")
Duplicate samplesThe same sample appears in both train and test (dedup wasn't thorough enough)
Group leakageSamples from the same user randomly split to both sides
Data augmentation leakageAugmented samples and original samples cross splits

Self-check question: what could I actually see at this moment during prediction?

Before every feature design and data split, ask yourself: "if I were going live to predict right now, would I actually have this value at this moment?" If the answer is no, it's leakage. This is the most important mental model for data engineers.

Data Versioning and Reproducibility ​

For models to be reproducible, data must be versioned: "same code + different data = different results," and vice versa. Three levels:

  1. Raw data is immutable: raw collected data is read-only, never modified in place (if changed, keep a version);
  2. Processing logic is versioned: cleaning/feature code goes into Git, bound to data versions (data_v1 + code_v2 → dataset_v3);
  3. Dataset manifests: use DVC (Data Version Control), lakeFS, or a "manifest file + object storage" to record data hashes, sources, and processing parameters.

The simplest practice: data directory (dataset card) + hash + processing script version number. The "unreproducibility" of ML experiments is mostly due to poor data version management, not code issues.

Tradeoffs ​

  • Data volume vs. data quality: more dirty data is worse than less clean data — but "rough first, then refine" (use a small amount of clean data to run the pipeline, then scale up) is a common rhythm;
  • Labeling cost vs. model return: when labeling is expensive, first use unsupervised / semi-supervised (clustering, pre-training) for preliminary value, then precisely supplement with targeted labeling (active learning);
  • Cleaning thoroughness vs. iteration speed: cleaning perfectionism slows things down — do "good enough" cleaning first to establish a baseline, then iterate the cleaning logic;
  • Feature store vs. simple pipeline: small teams with few features can start with "scripts + conventions"; once features multiply or span teams, a Feature Store becomes essential.

Further Reading ​

References ​