Skip to content

Pretraining: Data and Objectives

At a glance PrPretraining is Phase 1 of an LLM: next-token prediction on trillions of tokens of corpus, writing "language and world knowledge" into parameters. This article covers training objectives and loss, the full data pipeline (collection/cleaning/deduplication/filtering/mixing), data quality equals model quality, training dynamics like learning rate and batch size, and sets up scaling laws and post-training.

Pretraining: Data and Objectives ​

PrPretraining is Phase 1 of the LLM lifecycle: on massive, diverse, high-quality text corpora, training a Transformer with the single objective of "predicting the next token," encoding language patterns and world knowledge into parameters. It is the foundation of why large language models are "large" — post-training (SFT, alignment) can only fine-tune behavior; the vast majority of a model's knowledge forms during pretraining. Data determines the knowledge ceiling, and training dynamics determine whether the model can stably approach that ceiling.

One-line summary: pretraining = objective (next-token prediction) + data (full pipeline engineering) + dynamics (stably driving loss down), with all three jointly determining the base model's intellectual foundation. The "foundation model" produced by pretraining still needs fine-tuning and alignment before it can truly serve users.

1. Training Objectives and Loss ​

The pretraining objective is the autoregressive next-token prediction of language modeling (see tokenization & vocabulary for token definitions and choices): given a prefix, predict the next token, using cross-entropy to minimize negative log-likelihood.

text
Training corpus: massive text, tokenized into sequences (see tokenization & vocab)
For each training sample (fixed-length segment, e.g., 2048/404 tokens):
  Input  x1 x2 ... xn
  Model outputs per-position next-token probability distribution q_t
  Loss   L = -(1/n) Σ_t log q_t(x(t+1))     # only computed at positions with true tokens
Optimization: AdamW + cosine LR schedule + gradient clipping, running in parallel across multi-node multi-GPU

Key engineering points:

  • Uniform length within batches: samples are bucketed by length or padded to fixed length to keep tensor shapes regular.
  • Loss only on real positions: masked padding positions in the loss to prevent padding from affecting gradients.
  • Validation loss = PPL: "health checks" in pretraining are observing training/validation loss curves.

2. The Full Data Pipeline: From Internet to Training Corpus ​

The main source of pretraining corpus is web crawlers, but "the internet" ≠ "trainable corpus" — raw HTML is full of navigation bars, ads, duplicated content, and junk. The full pipeline is a production line:

text
Collection → Cleaning → Language filtering → Deduplication → Quality filtering → Mixing → Training batches

Collection: Web crawler snapshots like Common Crawl, books (eBooks/Gutenberg),
            papers (arXiv), code (GitHub), encyclopedias (Wikipedia), forums (Reddit, etc.)
Cleaning: HTML parsing to extract body text, remove ads/navigation, encoding normalization (UTF-8),
          paragraph splitting, tail noise removal
Language filtering: fastText language classifier, routing/splitting by language
Deduplication: URL dedup → document-level dedup (MinHash/SimHash approximate dedup)
              → line/paragraph-level dedup (suffix arrays), cleaning PII/toxic content
Quality filtering: quality classifier scoring (or perplexity filtering) to remove low-quality pages
Mixing: multi-language, code, math, books, dialogue, etc. mixed by target ratios

1. Corpus Sources & "Data is the Model" ​

Corpus SourceCharacteristicsTypical Use
Common Crawl subsets (C4, RefinedWeb, etc.)Largest scale, noisy, requires heavy cleaningMain data source
Wikipedia/EncyclopediasClean, structured, high knowledge densityFactual backbone
Books/E-booksLong coherent text, world knowledgeLong-range dependencies & narrative
GitHub CodeProgram syntax and logicCoding capability
arXiv/PapersRigorous academic languageScience/math
Forums (Reddit/StackExchange)Dialogue and Q&A styleInstruction-like expressions
Multilingual corpus (CulturaX, SkyPile, etc.)Coverage for Chinese/Japanese/European languagesMultilingual capability

The industry consensus is "data is the model": the quality and mixing of corpus affects final capabilities at least as much as model scale. Gopher (2022)'s systematic analysis, GPT-3's discussion of data source weights, all point to the same conclusion — cleaning and mixing are the engineering decisions that determine the ceiling, while model architecture simply extracts patterns from the data.

A few public cases quantify this impact: GPT-3's paper gives weight allocations for different data sources (Common Crawl ~60%, WebText2 ~22%, books & wiki ~18%) and corresponding oversampling multipliers; Gopher's analysis of an 800GB corpus showed that "document-level dedup + quality filtering" stably improved downstream benchmarks; Llama 2's team publicly shared their 2T-token mixing philosophy — more books and wiki, higher proportion of multilingual and code. The common conclusion: there's no universal optimal mixing ratio — only engineering solutions "aligned with target capabilities," and they must be calibrated with ablation studies.

2. Deduplication: Preventing "Cramming" ​

Duplicated data brings dual harm: first, training efficiency drops (many tokens just repeat already-known information, causing loss to saturate too early); second, evaluation data contamination (heavy overlap between training and eval sets, inflating scores). Mainstream approaches:

LevelMethodNote
URL dedupKeep only one copy per URLThe cheapest first gate
Document-level approximate dedupMinHash signatures + LSH for near-duplicate detectionEffective even for "rewritten reprints"
Line/paragraph-level dedupSuffix arrays (e.g., CCNet pipeline)Tackles "many documents sharing the same paragraph" (e.g., template text)
Eval set dedupn-gram overlap removal against benchmarks (e.g., MMLU/GSM8K, etc.)Protects evaluation credibility

The classic experiment from Deduplicating Training Data Makes Language Models Better (Lee et al., 2022) shows: simply removing duplicated data stably improves performance across multiple downstream tasks while training faster.

Another practical lesson on data contamination: when popular benchmarks (like MMLU, GSM8K) get widely discussed, they easily leak into web crawler corpus in various forms. Detection methods include n-gram overlap checks against benchmarks, timestamp comparison (if the corpus timestamp post-dates benchmark release, it's highly suspicious), and monitoring anomalous high scores for specific benchmarks during training. Contamination makes "model progress" measurements distorted — this is the most insidious hidden injury in evaluation practice (see Evaluation & Benchmarks).

3. Quality Filtering and Toxicity Handling ​

  • Quality classifiers: train a fastText classifier on "quality anchors" (like Wikipedia text) to score crawler pages; discard below threshold; or use language model perplexity filtering — low-PPL "more linguistically regular" text tends to be cleaner.
  • Toxicity/NSFW filtering: use classifiers and keyword filtering for adult content and violent text, both for safety and quality.
  • Privacy (PII): remove ID numbers, emails, phone numbers, etc. (Gopher's team has dedicated pipelines for this).

Over-cleaning has costs too

Filtering isn't "the more aggressive, the better": excessive dedup can delete legitimate long-tail knowledge; aggressive toxicity filtering can make models "over-evade" on related topics. The cleaning goal isn't "cleanest possible" but "distribution best supporting target capabilities." Modern teams use tiered multi-level filtering with small-scale ablation studies to set levels.

4. Data Mixing: The Science of Blend Ratios ​

How different sources are mixed directly shapes capability structure:

DecisionTypical PracticeConsiderations
Multilingual mixWeighted by population/corpus size (English dominant + several percentage points for Chinese/European)Language drift: underrepresented languages get diluted by English if ratio is too low
Code vs natural languageCode often 10%–30%Code improves reasoning and structured thinking, but too much degrades "natural speech" quality
Books/long-text oversampling2–3× resampling for booksCompensates for their naturally low representation in crawler corpus
Math/scienceTargeted corpus (math web pages, arXiv) mixed in small amountsTargeted improvement of reasoning and symbolic ability
Time cutoffUse "recent data" as validation set (time-based split)More realistic than random splits for real deployment scenarios

Mix optimization relies on small-scale ablations: train multiple mix variants on 1/1000 of the token budget, compare validation loss and target capabilities (e.g., code, math), then extrapolate to full scale. Mixing ratios also keep evolving — Datasets & Benchmarks Archive documents the composition and scale of public corpora like RedPajama, RefinedWeb, CulturaX, etc.

3. Training Dynamics: The Art of Driving Loss Down ​

With data and model in place, training itself is a whole engineering of dynamics:

1. Learning Rate Scheduling ​

Modern large models almost universally use warmup + cosine decay (or linear decay + very low tail LR):

text
Phase 1 (warmup, ~1%–3% of steps): LR ramps linearly from 0 to peak (e.g., 3e-4)
Phase 2 (main training): LR decays along cosine curve to near 0
Optional: last 10% of steps LR linearly drops to very low ("cool-down")
          empirically stabilizes convergence and slightly lowers loss

Peak learning rate decreases with model scale (roughly inversely proportional to the square root of parameters — an empirical rule, e.g., ~3e-4 for 7B, ~1.2e-4 for 70B), supported by μP/hyperparameter-transfer research (see Scaling Laws).

2. Batch Size and Gradient Accumulation ​

  • Global batch size is typically very large (hundreds of thousands to millions of tokens/step); if training is unstable, adjust batch first rather than LR.
  • Gradient accumulation: when GPU memory is insufficient, split large batches into micro-batches, each doing forward/backward pass, accumulating gradients before a unified update — mathematically equivalent to large batches, only affecting speed.
  • Theoretical basis: gradient noise scale (McCandlish et al., 2018) shows that once batch size reaches a critical threshold, training steps can be proportionally reduced, but too-large batches make the learning signal "too smooth," slowing convergence.

3. Training Monitoring and Mid-Run Evaluation ​

Long pretraining must be "monitored while running." Common practices: periodically run fixed small test sets (a few representative tasks, like code, math, multilingual samples) to observe capability improvement curves across training steps; save checkpoints and retain historical versions for rollback and post-hoc comparison; monitor gradient norm and loss variance — anomalous fluctuations often appear before explicit failures. In-training evaluation doesn't need full benchmarks; a few hundred representative samples suffice for "trend judgment."

4. Reading Loss Curves ​

Curve PatternMeaningAction
Training loss continuously drops, validation drops in syncHealthy trainingContinue
Training drops, validation stallsEarly overfitting (common with duplicated data)Increase data, heavier dedup
Training loss suddenly spikes (loss spike)Numerical instability or anomalous samples in dataRollback checkpoint, check recent data
Validation loss rebounds at some pointLR decay too fast or overfittingAdjust schedule, finish early

A critical rule of thumb

70% of pretraining effort goes to data, 20% to stability, 10% to architecture. When encountering "can't train," first suspect the data pipeline (duplicates, dirty samples, tokenizer mismatch), then numerical stability (LR, mixed precision, gradient clipping), and only then model structure.

5. A Typical Pretraining Configuration (Magnitude Reference) ​

Hyperparameter7B Scale70B Scale
Global batch (tokens/step)~0.5M–1M~2M–4M
Peak learning rate~3e-4~1.2e-4
Warmup step ratio~1%–3%~1%
Sequence length2048–40964096–8192
Precisionbf16 + fp32 masterbf16 + fp32 master
Gradient clipping~1.0~1.0

(Magnitude from open-source model tech reports; specifics depend on target model and data; larger batches typically pair with lower LR.)

6. Data Pipeline Engineering Checklist ​

The pretraining data chain is long-term infrastructure, worth managing like software engineering: corpus versioning (snapshot every cleaning/mixing change for reproducibility), dedup and filtering scripts in CI (regression tests on changes to prevent accidentally deleting a corpus type), data distribution monitoring (token-level language distribution, dedup rate, quality scores over time), isolation from eval (continuously updating benchmark exclusion lists). Treating data like code is the first step of "data is the model" landing in engineering.

7. Parallelism and Stability ​

Pretraining runs on thousands of GPUs for months, relying on data parallelism / tensor parallelism / pipeline parallelism (see Frameworks & Tool Selection's DeepSpeed, Megatron-LM), along with mixed precision (bf16), gradient clipping, loss scaling, and periodic checkpoint saving — one instability can ruin an entire training run, making "stability" a core metric of training engineering.

Common instability sources and countermeasures: gradient explosion (gradient clipping, lower LR), loss spikes (rollback to checkpoint before spike, investigate whether an anomalous data batch entered), mixed-precision overflow (bf16 is typically more stable than fp16; retain fp32 master weights when needed), validation loss rebound from data duplication (strengthen dedup). Making "rollback + investigation" a standard process is more realistic than trying to avoid all problems at once.

4. Pretraining vs Post-Training ​

DimensionPretrainingPost-Training (SFT/Alignment)
ObjectiveGeneral language modeling (next-token prediction)Instruction following, helpfulness/honesty/safety
DataTrillions of tokens of internet corpusHundreds of thousands to millions of high-quality instruction/dialogue samples
CostThousands of GPU·monthsHundreds of GPU·hours to GPU·days
OutputBase model ("can speak, has knowledge")Chat/instruction model ("can follow instructions")
EvaluationValidation PPL, small downstream probesInstruction evals, human comparison, safety tests
Key risksData quality and stabilityOverfitting, capability regression (catastrophic forgetting)

The two post-training routes are detailed in Fine-Tuning: SFT and Parameter-Efficient Fine-Tuning and Alignment: RLHF and DPO.

A common misconception is "pretraining handles knowledge, post-training handles instruction following" — more accurately: language structure, world knowledge, and reasoning foundations all form during pretraining; post-training changes "how behavior's conditional reflex works" (how to organize knowledge into instruction-aligned responses, how to refuse inappropriate requests). This is also why "post-training can't inject new knowledge out of thin air": if the base model hasn't seen a certain content type during pretraining, no amount of fine-tuning examples will stably generalize. Understanding this boundary is key to correctly allocating data and budget.

5. Scaling Laws Primer: Why "Pretrain Bigger" ​

Pretraining value follows scaling laws: for a fixed compute budget, validation loss decreases as a power law with parameter scale, data volume, and compute, with an optimal ratio among the three (~1 parameter per ~20 training tokens). This means:

  • "More data + bigger model" isn't magic; it's an engineering decision with a clear curve;
  • When data is insufficient, models are "over-parameterized and undertrained" — adding data is often more cost-effective than adding parameters (Chinchilla's conclusion);
  • Pretraining budget decisions (add data or add parameters first?) should be based on scaling laws, not intuition.

6. Trade-offs and Boundaries ​

  • Knowledge cutoff: pretraining data has a cutoff time; the model is ignorant of the world after that — this is one of the core motivations for RAG.
  • Long tail and rare languages: crawler corpus naturally favors English and popular domains; underrepresented languages and specialized long-tail knowledge have low coverage.
  • Data contamination is a long campaign: once a benchmark becomes popular, it can enter crawler corpus; contamination monitoring (n-gram overlap, timestamp comparison) should be normalized.
  • Cost and return: prPretraining is a "one-time purchase"; model architecture and tokenizer must be locked in before starting; subsequent capabilities can only be patched via post-training or incremental pretraining.

Three common pitfalls

  • "Bigger corpus = better": stacking volume without controlling quality causes loss to saturate early and capability skew;
  • "Loss won't go down → add parameters": rule out data and stability issues first, then discuss scale (see Scaling Laws);
  • "One-step pretraining": data mix, tokenizer, and architecture decisions must be frozen before launch; mid-course changes are prohibitively expensive.

One-line summary

Pretraining answers "where does the model's intellectual foundation come from": the objective is simple (next-token prediction), data carries all the complexity, dynamics determine stability. The foundation's height determines how far post-training can go.

Further Reading ​

References ​