Theme
Data and Data Engineering
Concept Definition: Data Sets the Model's Ceiling
A saying circulates in the machine learning community: "garbage in, garbage out." No matter how strong the model is, it can't fill the gaps left by bad data. The real-world time allocation of ML projects also confirms this: data preparation and cleaning typically take 60–80% of the time, while modeling itself is only a small part.
In the ML context, data engineering refers to all the work of transforming raw data into high-quality, reproducible, versioned training data and inference data — collection, cleaning, labeling, splitting, pipelines, storage.
The Data Lifecycle Panorama
text
Business systems / logs / external data
│ 1. Collect (logs, DBs, APIs, crawlers, purchased)
▼
Raw data storage (data lake / warehouse)
│ 2. Clean (missing/outliers/duplicates/formats)
▼
Clean data (versionable)
│ 3. Label (needed for supervised learning)
▼
Train/validation/test splits (leakage-proof)
│ 4. Feature engineering (see Feature Engineering chapter)
▼
Feature store (unified for training/inference)
│ 5. Train → Evaluate → Deploy (see MLOps chapter)
▼
Monitoring → Data drift → Re-collect/re-label (closed loop)Collection and Storage
Three Sources of Data
| Source | Examples | Caveats |
|---|---|---|
| Business systems | Transactions, clicks, behavior logs | Data definitions change as products evolve (schema drift) |
| External data | Public datasets, APIs, crawlers | Copyright, compliance (GDPR), stability |
| Labeled data | Human annotation, crowdsourcing, remote supervision | Labeling cost and quality are the core bottlenecks |
Storage Selection
- Data lake: store in raw format (object storage), flexible but messy;
- Data warehouse: structured, queried fast after ETL (BigQuery, Snowflake, ClickHouse);
- Feature store: unified training/inference features (Feast, Tecton, ByteDance FeatureStore) — solving the #1 engineering incident: "training features differ from online features."
Why feature stores matter
During training, you used the feature "user's average spending in the last 30 days." How do you guarantee the same definition and same time point during online inference? A feature store unified the definition, computation, and freshness of features; both training and inference pull from it — this is the watershed of MLOps maturity. See MLOps and Model Deployment.
Data Cleaning: Eight Forms of Dirty Data
| Dirty Data | Example | Handling |
|---|---|---|
| Missing values | User didn't fill in age | Impute/delete/flag; see Feature Engineering |
| Outliers | A house price of 100 million | Domain judgment + percentile truncation |
| Duplicates | Two records for the same order | Deduplicate (watch for legitimate "natural duplicates") |
| Format inconsistency | Dates: 2023/1/1 vs 01-01-2023 | Unified parsing and standardization |
| Unit inconsistency | Meters vs. feet, yuan vs. dollars | Unit normalization (easy to get wrong!) |
| Encoding chaos | Chinese garbled text, emojis | Unified encoding (UTF-8) |
| Logical contradictions | Age 5 but has a driver's license | Rule validation |
| Data drift | Purchasing distribution changes dramatically before/after Double 11 | Monitoring + per-group modeling |
The principle of cleaning is "auditability": every cleaning decision (what was deleted, what was imputed) must be traceable and re-runnable, otherwise you'll never know why the model "suddenly got better/worse" — it might be that the cleaning logic changed, not that the model improved.
Data Labeling: The Bottleneck of Supervised Learning
Labeling quality directly determines the model's ceiling. Three common approaches:
| Approach | Cost | Quality | Suitable For |
|---|---|---|---|
| Human labeling (in-house / crowdsourcing) | High | High | Core business data |
| Remote supervision (rule-based / knowledge base auto-labeling) | Low | Medium (noisy) | Large-scale initial labeling |
| Weak supervision / pre-labeling + human review | Medium | Medium-high | The mainstream approach of labeling tools |
Three disciplines of labeling:
- Labeling guidelines first: ambiguous cases ("does this image count as a cat?") must have rules defined beforehand, otherwise different annotators do their own thing;
- Consistency checks: have multiple annotators label the same batch of samples and compute labeling consistency (Cohen's Kappa); low consistency means the guidelines are unclear;
- Annotator bias: differences in standards among annotators become "noisy labels" that the model learns — consider majority voting and annotator review.
Data Splitting and Data Leakage
The Iron Rules of Splitting
- Never touch the test set during iteration (see Model Evaluation and Validation);
- Split criteria must match the real prediction scenario:
- i.i.d. data → random split;
- Time series → split by time (train on past, test on future);
- Grouped data (same user / same patient) → group-level split (GroupSplit), otherwise samples from the same user crossing splits = leakage.
Six Sources of Data Leakage
Data leakage is the #1 accident in ML engineering: test information leaks into training, offline metrics are inflated, and production fails. Common sources:
| Leakage Source | Example |
|---|---|
| Global statistics | Standardization/imputation computed on full data (including test set) |
| Future information | Using "full-month statistics" to predict the beginning of the month (temporal leakage) |
| Target leakage | Features containing the "result" (e.g., using "was refunded" to predict "will churn") |
| Duplicate samples | The same sample appears in both train and test (dedup wasn't thorough enough) |
| Group leakage | Samples from the same user randomly split to both sides |
| Data augmentation leakage | Augmented samples and original samples cross splits |
Self-check question: what could I actually see at this moment during prediction?
Before every feature design and data split, ask yourself: "if I were going live to predict right now, would I actually have this value at this moment?" If the answer is no, it's leakage. This is the most important mental model for data engineers.
Data Versioning and Reproducibility
For models to be reproducible, data must be versioned: "same code + different data = different results," and vice versa. Three levels:
- Raw data is immutable: raw collected data is read-only, never modified in place (if changed, keep a version);
- Processing logic is versioned: cleaning/feature code goes into Git, bound to data versions (
data_v1 + code_v2 → dataset_v3); - Dataset manifests: use DVC (Data Version Control), lakeFS, or a "manifest file + object storage" to record data hashes, sources, and processing parameters.
The simplest practice: data directory (dataset card) + hash + processing script version number. The "unreproducibility" of ML experiments is mostly due to poor data version management, not code issues.
Tradeoffs
- Data volume vs. data quality: more dirty data is worse than less clean data — but "rough first, then refine" (use a small amount of clean data to run the pipeline, then scale up) is a common rhythm;
- Labeling cost vs. model return: when labeling is expensive, first use unsupervised / semi-supervised (clustering, pre-training) for preliminary value, then precisely supplement with targeted labeling (active learning);
- Cleaning thoroughness vs. iteration speed: cleaning perfectionism slows things down — do "good enough" cleaning first to establish a baseline, then iterate the cleaning logic;
- Feature store vs. simple pipeline: small teams with few features can start with "scripts + conventions"; once features multiply or span teams, a Feature Store becomes essential.
Further Reading
- Feature Engineering — the downstream of cleaning
- Model Evaluation and Validation — the evaluation perspective on splitting and leakage
- MLOps and Model Deployment — connecting data pipelines and deployment monitoring
- Building an ML Project from Scratch — end-to-end data flow in practice
- Common Pitfalls and Anti-Patterns — data leakage failure stories
- Datasets and Tools — public datasets and tool lists
References
- DVC (Data Version Control) official documentation — data versioning tool
- Feature Store overview: Feast documentation — open-source feature store implementation
- Chen et al. Data Management in Machine Learning Systems (SIGMOD 2019) — classic paper on ML data management
- Polyzotis et al. Data Management Challenges in Production Machine Learning (SIGMOD 2017) — data challenges in production ML
- Kaggle: Data leakage discussions and competition experiences