Skip to content

Data and Data Engineering

Quick overview "Models are the cups, data is the water" — data quality directly determines the ceiling of models. This article covers data pipelines, augmentation methods for three modalities, annotation and active learning, data quality governance (deduplication/debiasing/noise), data leakage defenses, distributed loading, and data version management.

Data and Data Engineering ​

One-sentence definition: Data engineering is the end-to-end process of turning "raw materials" into "trainable data" — collection, cleaning, annotation, validation, augmentation, loading, and version management. There's a widely circulated saying in the deep learning community: "Garbage in, garbage out" — the ceiling of a model is determined by its data; algorithms are just tools for approaching that ceiling. This is also the prerequisite for the credibility of all conclusions in Deep Learning Evaluation and Experiments.

1. The Primary Importance of Data in DL ​

Deep learning is "data-driven function fitting" (see Neural Network Fundamentals). Data determines four things:

  1. The achievable ceiling: no matter how good the model is, it can't learn information not present in the data.
  2. The boundaries of generalization: the coverage of data (distribution) determines where the model can generalize (out-of-distribution topics covered in MLOps and Model Deployment).
  3. Whether the problem is solvable: poor annotation quality means no amount of loss function optimization will help.
  4. Engineering cost: the data pipeline often consumes 50–80% of project time — this is a universal consensus across all industry projects.

A golden rule: look at the data before building models. Intuitive judgment through data visualization (class distributions, bad samples, annotation errors) can save massive amounts of blind hyperparameter tuning time and is the first step in Debugging and Diagnostics.

2. Data Pipelines: Collection, Cleaning, Annotation, Validation ​

A complete data pipeline contains four stages:

  • Collection: sources include public datasets, logs, web crawlers, sensors, user feedback, and synthetic data. Be mindful of copyright and privacy compliance.
  • Cleaning: deduplicate, remove invalid samples (empty values / corrupted files), unify formats, and fix obvious errors. A classic cautionary tale: an image dataset contaminated with "screenshot images of text" caused the model to learn to classify based on "whether there is text" rather than the actual content.
  • Annotation: humans or automated systems tag samples. Annotation guidelines must clearly specify edge cases; otherwise, standard drift between annotators becomes model noise.
  • Validation: spot-check annotation quality (e.g., have two annotators label 5% of samples and calculate agreement), check distributions (class counts, length distributions, outliers), and set quality thresholds (reject batches that don't meet the threshold).

Every step should be traceable — "who produced this batch, when, and under what guidelines" — this is a hard requirement in the reproducibility section of MLOps and Model Deployment.

3. Image / Text / Audio Augmentation ​

Data augmentation creates transformed samples on-the-fly during training, teaching the model invariance — it serves both as regularization (see Overfitting and Regularization) and as part of the data pipeline.

  • Images: geometric transforms (flip/rotate/crop/scale), color transforms (brightness/contrast/hue), noise injection, blurring; advanced methods like Mixup/CutMix (see Overfitting and Regularization); domain-specific constraints — medical images shouldn't be arbitrarily flipped (left/right asymmetry has anatomical meaning), document OCR shouldn't be rotated.
  • Text: synonym replacement, back-translation (translate to another language and back), random deletion/swap, EDA; using language models to generate paraphrases is the latest trend. Text augmentation carries the highest risk — changing one word can flip the semantics (the word "not"), so semantic preservation checks are essential.
  • Audio: time/speed perturbation, volume/pitch changes, noise injection, SpecAugment (time/frequency masking on spectrograms). See the Speech and Audio chapter for speech task guidelines.

Principles of augmentation

Augmentation must align with the task's "invariances": augmentation is essentially declaring "these transformations shouldn't change semantics." Get it wrong, and you inject wrong labels into the training set. Set intensity based on the validation set (Deep Learning Evaluation and Experiments).

4. Annotation Strategies and Active Learning ​

Annotation is the most expensive part of data engineering; strategy choices affect both cost and quality:

  • Budget allocation: the highest ROI comes from labeling the "hardest" samples first — this is active learning: train an initial model, select its most uncertain samples (e.g., highest prediction entropy, closest to the decision boundary) for human annotation, iterate, and reach target accuracy with fewer annotations.
  • Pre-labeling + human review: use an existing model to generate pseudo-labels, with humans correcting only errors — cost can drop by an order of magnitude.
  • Annotation consistency: multiple annotators require clear guidelines and disagreement arbitration mechanisms; calculate annotation agreement (e.g., Cohen's kappa) as a quality gate.
  • Weak supervision / programmatic annotation: use rules, heuristics, and distant supervision for automatic labeling (e.g., "if a brand name is present, label it as a brand mention"), suitable for coarse-grained tasks but watch the noise rate.

5. Data Quality: Deduplication, Debiasing, Noise ​

Deduplication: duplicate/near-duplicate samples in the training set amplify their statistical weight, causing the model to bias toward high-frequency content. Deduplicating at scale is especially critical in LLM training corpora (using sentence n-gram similarity, embedding similarity); otherwise, generated content will clearly favor repetitive text. Details in Large Language Models (LLM).

Debiasing: statistical biases in data will be faithfully learned into representations (see the bias discussion in Representation Learning and Pretraining). Debiasing methods: ensure representation across groups, oversampling/reweighting, removing proxy features for sensitive attributes. Evaluation and correction are covered in Interpretability and Fairness.

Noise: label noise (wrong labels) is more lethal than feature noise. Remedies: confident learning to identify mislabeled samples, label smoothing (see the losses chapter) to increase tolerance to noise, and sample-weighted learning by difficulty.

6. Types and Defenses Against Data Leakage ​

Data leakage is "future/external information leaking into training" — the top source of inflated evaluation numbers (see the evaluation chapter). Typical types and defenses:

TypeExampleDefense
Pre-processing leakageFitting normalization statistics on the entire datasetFit only on the training set, then apply to validation/test
Temporal leakageUsing future data to predict the past (in finance)Split by time order; no random splitting
Same-source leakageSamples from the same user/scenario split across setsSplit by user/session/video groups
Duplicate leakageSplitting before deduplication, with duplicate samples across setsGlobal deduplication first, then split
Human leakageFeatures containing label information (e.g., "whether the patient was cured")Feature-source auditing, feature dictionary management

A self-check checklist for defenses: split order (preprocess before or after splitting?; split before or after preprocessing?), grouping dimension (split by entity?), temporal dimension (for time-series tasks, is the future forbidden?).

7. Distributed Data Loading: DataLoader, Caching, Prefetching ​

Data loading is a hidden bottleneck in training throughput — waiting GPUs for CPU data is pure waste. Key points:

  • DataLoader parameters: num_workers (multi-process reading), pin_memory (pinned memory for faster GPU transfer), prefetch_factor (prefetching), persistent_workers (worker reuse).
  • Caching strategy: small data → cache everything in memory (the MemoryDataset pattern in torch.utils.data); large data → preprocess into efficient formats (TFRecord, WebDataset, Safetensors) with sequential reading + multi-process parallel decompression.
  • Training/inference throughput accounting: is GPU compute matched with data IO? Watch for the classic symptom of "low GPU utilization + high CPU load."
  • At scale: in distributed training, each rank reads only its own shard; use torch.utils.data.distributed.DistributedSampler to guarantee no global duplicates and correct shuffling.

Common pitfall

Out-of-sync random seeds across processes cause "each worker to see the same augmentations"; after setting the global random seed, remember to derive different seeds for each worker. These symptoms are subtle — see Debugging and Diagnostics.

8. Data Version Management: DVC and HF Datasets ​

For reproducible models, you need versioned data. Two key tools:

  • DVC (Data Version Control): stores metadata of data files in git (content-addressed), while data itself is stored remotely (S3/MinIO/local). dvc repro tracks the full lineage of "data version → processing script → model artifacts." Suitable for custom pipelines.
  • Hugging Face Datasets: load_dataset + caching + Arrow format, with built-in dataset cards and versioning. The de facto standard in the open-source community, also supports streaming processing for massive corpora.

Data version management pairs with experiment tracking (see the MLOps chapter): a complete record should be "data v2.1 + code commit + hyperparameters + seed → metrics" — no component is optional.

9. Trade-offs ​

Trade-offs

Data scale vs. data quality: no absolute answer to "more noisy data vs. less clean data." A common approach: "get the pipeline working with a cleaner, smaller dataset first, then decide whether to scale based on returns" — the cost-effectiveness of fixing data quality far exceeds blindly adding volume.

Online vs. offline augmentation: online augmentation doesn't take disk space, produces different transforms each epoch, and runs in parallel with training; offline augmentation (pre-generated) is reproducible and auditable but consumes storage, doubling disk usage. Large projects often use "online as primary + key transforms backed up offline."

Annotation cost vs. model returns: how much does each additional 10k annotated samples improve accuracy? Use learning curves (see the evaluation chapter) early in the project to calculate the "data return diminishing point," avoiding wasteful spending during the plateau.

Full pre-processing vs. streaming: full pre-processing (convert all at once) loads fast, but changing the pre-processing logic requires a full re-run; streaming is flexible and cheap to modify, but each epoch costs extra time for decompression/transformations. When data is so large that "pre-processed results can't be stored," streaming is the only option.

Data engineering has no glamorous algorithms, but it determines the ceiling of any algorithm. Making every step of the data pipeline "reproducible, auditable, and versioned" is half the battle for a successful deep learning project. For an archive and selection guide of commonly used public datasets, see Datasets and Tools; for a hands-on guide to building a small data pipeline from scratch, see the practice section.

Further Reading ​

References ​