Theme
Pretraining: Data and Objectives
PrPretraining is Phase 1 of the LLM lifecycle: on massive, diverse, high-quality text corpora, training a Transformer with the single objective of "predicting the next token," encoding language patterns and world knowledge into parameters. It is the foundation of why large language models are "large" — post-training (SFT, alignment) can only fine-tune behavior; the vast majority of a model's knowledge forms during pretraining. Data determines the knowledge ceiling, and training dynamics determine whether the model can stably approach that ceiling.
One-line summary: pretraining = objective (next-token prediction) + data (full pipeline engineering) + dynamics (stably driving loss down), with all three jointly determining the base model's intellectual foundation. The "foundation model" produced by pretraining still needs fine-tuning and alignment before it can truly serve users.
1. Training Objectives and Loss
The pretraining objective is the autoregressive next-token prediction of language modeling (see tokenization & vocabulary for token definitions and choices): given a prefix, predict the next token, using cross-entropy to minimize negative log-likelihood.
text
Training corpus: massive text, tokenized into sequences (see tokenization & vocab)
For each training sample (fixed-length segment, e.g., 2048/404 tokens):
Input x1 x2 ... xn
Model outputs per-position next-token probability distribution q_t
Loss L = -(1/n) Σ_t log q_t(x(t+1)) # only computed at positions with true tokens
Optimization: AdamW + cosine LR schedule + gradient clipping, running in parallel across multi-node multi-GPUKey engineering points:
- Uniform length within batches: samples are bucketed by length or padded to fixed length to keep tensor shapes regular.
- Loss only on real positions: masked padding positions in the loss to prevent padding from affecting gradients.
- Validation loss = PPL: "health checks" in pretraining are observing training/validation loss curves.
2. The Full Data Pipeline: From Internet to Training Corpus
The main source of pretraining corpus is web crawlers, but "the internet" ≠ "trainable corpus" — raw HTML is full of navigation bars, ads, duplicated content, and junk. The full pipeline is a production line:
text
Collection → Cleaning → Language filtering → Deduplication → Quality filtering → Mixing → Training batches
Collection: Web crawler snapshots like Common Crawl, books (eBooks/Gutenberg),
papers (arXiv), code (GitHub), encyclopedias (Wikipedia), forums (Reddit, etc.)
Cleaning: HTML parsing to extract body text, remove ads/navigation, encoding normalization (UTF-8),
paragraph splitting, tail noise removal
Language filtering: fastText language classifier, routing/splitting by language
Deduplication: URL dedup → document-level dedup (MinHash/SimHash approximate dedup)
→ line/paragraph-level dedup (suffix arrays), cleaning PII/toxic content
Quality filtering: quality classifier scoring (or perplexity filtering) to remove low-quality pages
Mixing: multi-language, code, math, books, dialogue, etc. mixed by target ratios1. Corpus Sources & "Data is the Model"
| Corpus Source | Characteristics | Typical Use |
|---|---|---|
| Common Crawl subsets (C4, RefinedWeb, etc.) | Largest scale, noisy, requires heavy cleaning | Main data source |
| Wikipedia/Encyclopedias | Clean, structured, high knowledge density | Factual backbone |
| Books/E-books | Long coherent text, world knowledge | Long-range dependencies & narrative |
| GitHub Code | Program syntax and logic | Coding capability |
| arXiv/Papers | Rigorous academic language | Science/math |
| Forums (Reddit/StackExchange) | Dialogue and Q&A style | Instruction-like expressions |
| Multilingual corpus (CulturaX, SkyPile, etc.) | Coverage for Chinese/Japanese/European languages | Multilingual capability |
The industry consensus is "data is the model": the quality and mixing of corpus affects final capabilities at least as much as model scale. Gopher (2022)'s systematic analysis, GPT-3's discussion of data source weights, all point to the same conclusion — cleaning and mixing are the engineering decisions that determine the ceiling, while model architecture simply extracts patterns from the data.
A few public cases quantify this impact: GPT-3's paper gives weight allocations for different data sources (Common Crawl ~60%, WebText2 ~22%, books & wiki ~18%) and corresponding oversampling multipliers; Gopher's analysis of an 800GB corpus showed that "document-level dedup + quality filtering" stably improved downstream benchmarks; Llama 2's team publicly shared their 2T-token mixing philosophy — more books and wiki, higher proportion of multilingual and code. The common conclusion: there's no universal optimal mixing ratio — only engineering solutions "aligned with target capabilities," and they must be calibrated with ablation studies.
2. Deduplication: Preventing "Cramming"
Duplicated data brings dual harm: first, training efficiency drops (many tokens just repeat already-known information, causing loss to saturate too early); second, evaluation data contamination (heavy overlap between training and eval sets, inflating scores). Mainstream approaches:
| Level | Method | Note |
|---|---|---|
| URL dedup | Keep only one copy per URL | The cheapest first gate |
| Document-level approximate dedup | MinHash signatures + LSH for near-duplicate detection | Effective even for "rewritten reprints" |
| Line/paragraph-level dedup | Suffix arrays (e.g., CCNet pipeline) | Tackles "many documents sharing the same paragraph" (e.g., template text) |
| Eval set dedup | n-gram overlap removal against benchmarks (e.g., MMLU/GSM8K, etc.) | Protects evaluation credibility |
The classic experiment from Deduplicating Training Data Makes Language Models Better (Lee et al., 2022) shows: simply removing duplicated data stably improves performance across multiple downstream tasks while training faster.
Another practical lesson on data contamination: when popular benchmarks (like MMLU, GSM8K) get widely discussed, they easily leak into web crawler corpus in various forms. Detection methods include n-gram overlap checks against benchmarks, timestamp comparison (if the corpus timestamp post-dates benchmark release, it's highly suspicious), and monitoring anomalous high scores for specific benchmarks during training. Contamination makes "model progress" measurements distorted — this is the most insidious hidden injury in evaluation practice (see Evaluation & Benchmarks).
3. Quality Filtering and Toxicity Handling
- Quality classifiers: train a fastText classifier on "quality anchors" (like Wikipedia text) to score crawler pages; discard below threshold; or use language model perplexity filtering — low-PPL "more linguistically regular" text tends to be cleaner.
- Toxicity/NSFW filtering: use classifiers and keyword filtering for adult content and violent text, both for safety and quality.
- Privacy (PII): remove ID numbers, emails, phone numbers, etc. (Gopher's team has dedicated pipelines for this).
Over-cleaning has costs too
Filtering isn't "the more aggressive, the better": excessive dedup can delete legitimate long-tail knowledge; aggressive toxicity filtering can make models "over-evade" on related topics. The cleaning goal isn't "cleanest possible" but "distribution best supporting target capabilities." Modern teams use tiered multi-level filtering with small-scale ablation studies to set levels.
4. Data Mixing: The Science of Blend Ratios
How different sources are mixed directly shapes capability structure:
| Decision | Typical Practice | Considerations |
|---|---|---|
| Multilingual mix | Weighted by population/corpus size (English dominant + several percentage points for Chinese/European) | Language drift: underrepresented languages get diluted by English if ratio is too low |
| Code vs natural language | Code often 10%–30% | Code improves reasoning and structured thinking, but too much degrades "natural speech" quality |
| Books/long-text oversampling | 2–3× resampling for books | Compensates for their naturally low representation in crawler corpus |
| Math/science | Targeted corpus (math web pages, arXiv) mixed in small amounts | Targeted improvement of reasoning and symbolic ability |
| Time cutoff | Use "recent data" as validation set (time-based split) | More realistic than random splits for real deployment scenarios |
Mix optimization relies on small-scale ablations: train multiple mix variants on 1/1000 of the token budget, compare validation loss and target capabilities (e.g., code, math), then extrapolate to full scale. Mixing ratios also keep evolving — Datasets & Benchmarks Archive documents the composition and scale of public corpora like RedPajama, RefinedWeb, CulturaX, etc.
3. Training Dynamics: The Art of Driving Loss Down
With data and model in place, training itself is a whole engineering of dynamics:
1. Learning Rate Scheduling
Modern large models almost universally use warmup + cosine decay (or linear decay + very low tail LR):
text
Phase 1 (warmup, ~1%–3% of steps): LR ramps linearly from 0 to peak (e.g., 3e-4)
Phase 2 (main training): LR decays along cosine curve to near 0
Optional: last 10% of steps LR linearly drops to very low ("cool-down")
empirically stabilizes convergence and slightly lowers lossPeak learning rate decreases with model scale (roughly inversely proportional to the square root of parameters — an empirical rule, e.g., ~3e-4 for 7B, ~1.2e-4 for 70B), supported by μP/hyperparameter-transfer research (see Scaling Laws).
2. Batch Size and Gradient Accumulation
- Global batch size is typically very large (hundreds of thousands to millions of tokens/step); if training is unstable, adjust batch first rather than LR.
- Gradient accumulation: when GPU memory is insufficient, split large batches into micro-batches, each doing forward/backward pass, accumulating gradients before a unified update — mathematically equivalent to large batches, only affecting speed.
- Theoretical basis: gradient noise scale (McCandlish et al., 2018) shows that once batch size reaches a critical threshold, training steps can be proportionally reduced, but too-large batches make the learning signal "too smooth," slowing convergence.
3. Training Monitoring and Mid-Run Evaluation
Long pretraining must be "monitored while running." Common practices: periodically run fixed small test sets (a few representative tasks, like code, math, multilingual samples) to observe capability improvement curves across training steps; save checkpoints and retain historical versions for rollback and post-hoc comparison; monitor gradient norm and loss variance — anomalous fluctuations often appear before explicit failures. In-training evaluation doesn't need full benchmarks; a few hundred representative samples suffice for "trend judgment."
4. Reading Loss Curves
| Curve Pattern | Meaning | Action |
|---|---|---|
| Training loss continuously drops, validation drops in sync | Healthy training | Continue |
| Training drops, validation stalls | Early overfitting (common with duplicated data) | Increase data, heavier dedup |
| Training loss suddenly spikes (loss spike) | Numerical instability or anomalous samples in data | Rollback checkpoint, check recent data |
| Validation loss rebounds at some point | LR decay too fast or overfitting | Adjust schedule, finish early |
A critical rule of thumb
70% of pretraining effort goes to data, 20% to stability, 10% to architecture. When encountering "can't train," first suspect the data pipeline (duplicates, dirty samples, tokenizer mismatch), then numerical stability (LR, mixed precision, gradient clipping), and only then model structure.
5. A Typical Pretraining Configuration (Magnitude Reference)
| Hyperparameter | 7B Scale | 70B Scale |
|---|---|---|
| Global batch (tokens/step) | ~0.5M–1M | ~2M–4M |
| Peak learning rate | ~3e-4 | ~1.2e-4 |
| Warmup step ratio | ~1%–3% | ~1% |
| Sequence length | 2048–4096 | 4096–8192 |
| Precision | bf16 + fp32 master | bf16 + fp32 master |
| Gradient clipping | ~1.0 | ~1.0 |
(Magnitude from open-source model tech reports; specifics depend on target model and data; larger batches typically pair with lower LR.)
6. Data Pipeline Engineering Checklist
The pretraining data chain is long-term infrastructure, worth managing like software engineering: corpus versioning (snapshot every cleaning/mixing change for reproducibility), dedup and filtering scripts in CI (regression tests on changes to prevent accidentally deleting a corpus type), data distribution monitoring (token-level language distribution, dedup rate, quality scores over time), isolation from eval (continuously updating benchmark exclusion lists). Treating data like code is the first step of "data is the model" landing in engineering.
7. Parallelism and Stability
Pretraining runs on thousands of GPUs for months, relying on data parallelism / tensor parallelism / pipeline parallelism (see Frameworks & Tool Selection's DeepSpeed, Megatron-LM), along with mixed precision (bf16), gradient clipping, loss scaling, and periodic checkpoint saving — one instability can ruin an entire training run, making "stability" a core metric of training engineering.
Common instability sources and countermeasures: gradient explosion (gradient clipping, lower LR), loss spikes (rollback to checkpoint before spike, investigate whether an anomalous data batch entered), mixed-precision overflow (bf16 is typically more stable than fp16; retain fp32 master weights when needed), validation loss rebound from data duplication (strengthen dedup). Making "rollback + investigation" a standard process is more realistic than trying to avoid all problems at once.
4. Pretraining vs Post-Training
| Dimension | Pretraining | Post-Training (SFT/Alignment) |
|---|---|---|
| Objective | General language modeling (next-token prediction) | Instruction following, helpfulness/honesty/safety |
| Data | Trillions of tokens of internet corpus | Hundreds of thousands to millions of high-quality instruction/dialogue samples |
| Cost | Thousands of GPU·months | Hundreds of GPU·hours to GPU·days |
| Output | Base model ("can speak, has knowledge") | Chat/instruction model ("can follow instructions") |
| Evaluation | Validation PPL, small downstream probes | Instruction evals, human comparison, safety tests |
| Key risks | Data quality and stability | Overfitting, capability regression (catastrophic forgetting) |
The two post-training routes are detailed in Fine-Tuning: SFT and Parameter-Efficient Fine-Tuning and Alignment: RLHF and DPO.
A common misconception is "pretraining handles knowledge, post-training handles instruction following" — more accurately: language structure, world knowledge, and reasoning foundations all form during pretraining; post-training changes "how behavior's conditional reflex works" (how to organize knowledge into instruction-aligned responses, how to refuse inappropriate requests). This is also why "post-training can't inject new knowledge out of thin air": if the base model hasn't seen a certain content type during pretraining, no amount of fine-tuning examples will stably generalize. Understanding this boundary is key to correctly allocating data and budget.
5. Scaling Laws Primer: Why "Pretrain Bigger"
Pretraining value follows scaling laws: for a fixed compute budget, validation loss decreases as a power law with parameter scale, data volume, and compute, with an optimal ratio among the three (~1 parameter per ~20 training tokens). This means:
- "More data + bigger model" isn't magic; it's an engineering decision with a clear curve;
- When data is insufficient, models are "over-parameterized and undertrained" — adding data is often more cost-effective than adding parameters (Chinchilla's conclusion);
- Pretraining budget decisions (add data or add parameters first?) should be based on scaling laws, not intuition.
6. Trade-offs and Boundaries
- Knowledge cutoff: pretraining data has a cutoff time; the model is ignorant of the world after that — this is one of the core motivations for RAG.
- Long tail and rare languages: crawler corpus naturally favors English and popular domains; underrepresented languages and specialized long-tail knowledge have low coverage.
- Data contamination is a long campaign: once a benchmark becomes popular, it can enter crawler corpus; contamination monitoring (n-gram overlap, timestamp comparison) should be normalized.
- Cost and return: prPretraining is a "one-time purchase"; model architecture and tokenizer must be locked in before starting; subsequent capabilities can only be patched via post-training or incremental pretraining.
Three common pitfalls
- "Bigger corpus = better": stacking volume without controlling quality causes loss to saturate early and capability skew;
- "Loss won't go down → add parameters": rule out data and stability issues first, then discuss scale (see Scaling Laws);
- "One-step pretraining": data mix, tokenizer, and architecture decisions must be frozen before launch; mid-course changes are prohibitively expensive.
One-line summary
Pretraining answers "where does the model's intellectual foundation come from": the objective is simple (next-token prediction), data carries all the complexity, dynamics determine stability. The foundation's height determines how far post-training can go.
Further Reading
- Language Modeling: The Next-Token Prediction Paradigm — the probability-theoretic essence of the pretraining objective
- Tokenization & Vocab — the input unit directly shapes corpus and vocab morphology
- Scaling Laws — optimal ratio of data, parameters, compute
- Fine-Tuning: SFT and Parameter-Efficient Fine-Tuning — the first stop after pretraining
- Alignment: RLHF and DPO — turning the base model into an instruction-following assistant
- Datasets & Benchmarks Archive — true composition of public corpora and benchmarks
References
- Kaplan et al. Scaling Laws for Neural Language Models (2020) — the classic conclusion that pretraining loss decreases as a power law with scale
- Hoffmann et al. Training Compute-Optimal Large Language Models (Chinchilla, 2022) — optimal ratio of parameters and tokens
- Rae et al. Scaling Language Models: Methods, Analysis & Insights from Training Gopher (2021) — industry-level analysis of the full pretraining data pipeline
- Lee et al. Deduplicating Training Data Makes Language Models Better (2022) — empirical value of MinHash deduplication
- Penedo et al. The RefinedWeb Dataset (2023) — representative open-source high-quality web corpus pipeline
- Gao et al. The Pile: An 800GB Dataset of Diverse Text for Language Modeling (2020) — multi-source corpus benchmark
- McCandlish et al. An Empirical Model of Large-Batch Training (2018) — theory of batch size and training efficiency
- Common Crawl website — the largest public source for pretraining corpus