Skip to content

Building a Large Model from Scratch

At a glance llowing the nanoGPT approach, walk through the complete pipeline of "data → tokenization → model → training → sampling" using minimal runnable PyTorch code: hand-train a small GPT that can produce Shakespeare-like text, and get an extension checklist from 1M to 1B parameters with four acceptance criteria.

Building a Large Model from Scratch ​

Reading ten papers is not as good as training one model yourself. This article takes you through nanoGPT's minimal path, writing a small GPT that truly runs and truly generates text: data → tokenizer → model → training → sampling, five steps, none skipped.

Many people's first reaction to large models is "that's something only big companies can touch." But the fact is: you can train a small GPT that produces coherent text with an ordinary laptop GPU (or even pure CPU). This article compresses the process to the minimal runnable version — all code combined is under two hundred lines, and everything is provided. What you get is not a "toy," but a hands-on verifiable mental model for understanding how GPT models work: from then on, when you see the term "large model," you'll have in mind its data flow, loss curves, and sampling process, not just a black box.

One, Roadmap Overview: Five Steps + Four Acceptance Criteria ​

The complete roadmap and the core question each step answers:

Shakespeare's complete works (approx. 1MB of raw text)
    │
    ▼
① Data preparation ──────── Answers "what kind of text is the model learning"
    ▼
② Tokenizer ─────────────── Answers "how text gets chunked into model input units"
    ▼
③ Model definition ──────── Answers "what the network looks like for predicting the next token"
    ▼
④ Training loop ─────────── Answers "how loss decreases step by step and the model gets smarter"
    ▼
⑤ Sampling / generation ─── Answers "what a trained model can output"

This is the minimal closed loop of the language modeling paradigm: the model's only task is to predict the next token. Understanding this matters more than memorizing any formula.

Four Acceptance Criteria ​

Before you start, define what "success" looks like so you don't finish training and not know how to evaluate results:

#Acceptance ItemThreshold (with this config)How to Check
①Code runsTraining loop doesn't error; loss drops from 4~5Observe the printed train loss
②Loss drops noticeablyTrain loss goes from ~4.2 to under 1.8 within 3000 stepsCompare against training curves (see Section Five)
③Overfitting signal is reasonableval loss drops then rises, train loss continues to dropCompare train / val loss curves
④Generated text is readableSampled output is a sequence of "English / Shakespeare-like" wordsRun the sampling function and read it

What you're accepting is not "looking like ChatGPT"

A small GPT (~1M parameters) cannot produce logically coherent dialogue. Its success standard is: it has learned Shakespeare's word choice habits, punctuation rhythm, and partial spelling patterns. This is actually the best teaching moment — you can clearly see how "statistical patterns" emerge from data. To verify a smarter model, go through the extension checklist in Section Seven first.

Local Resource Budget ​

ConfigVRAM / MemoryEstimated Training Time for 1M ModelWhat It's Good For
Pure CPU (e.g., laptop 8-core)Only ~2GB RAM needed~20~40 minutesEverything in this article, just be patient
Entry GPU (e.g., RTX 3060 12GB)~2GB VRAM~3~5 minutesThis article + parameter tweaking
Mid-range GPU (e.g., RTX 4090 24GB)~2GB VRAM~1~2 minutesFast iteration; continue toward 10M~50M parameters
Cloud GPU (e.g., A100 40GB)No pressureSecondsNot necessary for this article; save for the Section Seven extensions

A harsh comparison

GPT-2 (1.5B parameters) needs days of pre-training across 8 A100s, while the 1M model in this article can be trained on a CPU in half an hour — the gap brought by scale is orders of magnitude. This is a real-world footnote to scaling laws: a large model's "intelligence" is largely the result of piling compute and data, and this article lets you experience firsthand "why small models fall short."

Two, Data: Starting with Shakespeare ​

1. Why Shakespeare ​

Karpathy's nanoGPT tutorial chooses Shakespeare's complete plays (about 1MB of plain text) as the starting point for highly practical reasons:

  • Small enough: 1MB loads in seconds on a CPU; any machine can train it.
  • Patterned enough: Dramatic text has character, dialogue, and scene structure with clear statistical patterns that small models can learn to mimic.
  • Interesting enough: The generated "fake Shakespeare" is intuitively readable, giving a sense of accomplishment at verification time.

Download the data:

bash
# Directly download tiny Shakespeare (the data file from Karpathy's char-rnn repo)
mkdir -p data/shakespeare
wget https://raw.githubusercontent.com/karpathy/char-rnn/master/data/tinyshakespeare/input.txt -O data/shakespeare/input.txt
# If Windows lacks wget, use curl: curl -o data/shakespeare/input.txt <same URL>
python
# Load and quick health check
with open("data/shakespeare/input.txt", "r", encoding="utf-8") as f:
    text = f.read()

print(f"Total characters: {len(text)}")          # ~1.1M
print(text[:500])                                # First 500 characters, get a feel

If you want to level up to "more like real pre-training," you can swap in OpenWebText (the unofficial recreation of GPT-2's training set, see References), but that's a "from 1M to 1B" step — we'll leave that for later.

2. Data Split: Training / Validation ​

Consistent with the machine learning manual's discipline of "touch the test set only once at the end" (see Common Pitfalls and Anti-Patterns), LLM training also requires an independent validation set — but note: an LLM's validation set is "unseen text segments from the same document," not a separate document. This is a key difference between text data and tabular data.

python
# Use 90% of text for training, 10% for validation (don't shuffle randomly — maintaining text order matters for language models)
n = len(text)
train_data = text[: int(n * 0.9)]
val_data = text[int(n * 0.9):]

Why no shuffling here

Text sequences have an inherent order (what comes before determines what comes after). If you randomly shuffle characters like tabular data, the model will learn "the probability distribution of random strings," which is meaningless. During training, we cut sequential batches, but the positions within each batch are randomized (see Section Four), ensuring every step's gradient comes from different regions of the text.

Three, Tokenization: From Character-Level to BPE ​

1. Character-Level Tokenizer: Best for Teaching ​

The tokenization mechanism of large models is detailed in Tokenization & Vocabularies. We start with the simplest character-level tokenization: each character is one token. This has two advantages — minimal code and clearest concepts.

python
# Count all distinct characters in the text as the vocabulary
chars = sorted(list(set(train_data)))
vocab_size = len(chars)
print(f"Vocabulary size: {vocab_size}")   # Usually 65 (letters + punctuation + spaces, etc.)

# Build bidirectional character <-> integer mapping
stoi = {ch: i for i, ch in enumerate(chars)}   # char -> id
itos = {i: ch for i, ch in enumerate(chars)}   # id -> char

def encode(s):
    """Convert a string to a list of integers"""
    return [stoi[c] for c in s]

def decode(ids):
    """Convert a list of integers back to a string"""
    return "".join(itos[i] for i in ids)

# Self-check: encoding then decoding should restore the original text
sample = "To be, or not to be"
assert decode(encode(sample)) == sample
print("Tokenizer self-check passed")

2. Upgrading from Character-Level to BPE ​

A character-level vocabulary has only 65 tokens, but the cost is very low information per token — the model needs more steps to learn the "word" concept. Real large models all use subword tokenization like BPE, with vocabularies typically of 30K~100K tokens (Llama 3 uses 128K, GPT-4 uses about 100K). You can try it in one line with OpenAI's open-source tiktoken:

python
pip install tiktoken

import tiktoken
enc = tiktoken.get_encoding("cl100k_base")      # Tokenizer used by GPT-4
ids = enc.encode("To be, or not to be")
print(ids)                                       # [904, 471, ..., 471]
print(enc.decode(ids))                           # Restores the original text
print("Token count:", len(ids), "vs. char count:", len("To be, or not to be"))

This comparison (how many tokens a piece of English gets chunked into) directly determines API call costs — tokens are the pricing unit in the LLM world. Conversion factors are covered in Context Windows & Long Text.

How to choose for this article

All code in the next four sections is based on character-level tokenizers (to guarantee minimal runnable code). Swap the encode/decode from Section Three Part 1 with tiktoken's enc.encode/enc.decode, and the model can ingest BPE tokens — you only need to change two lines of code. This is a concrete demonstration of "tokenization is an external component of the model."

Four, Model: Minimal GPT (Complete Runnable Code) ​

1. Architecture Overview ​

The model itself is the decoder-only subset of Transformer Architecture Explained, with four core components:

ComponentRoleNotes
TokenEmbeddingMaps token IDs to dense vectorsLearnable lookup table
CausalSelfAttentionLet each position only see itself and previous tokensCausal mask is key to GPT
MLP (Feed-Forward)Non-linear transformation per positionEach token computed independently
LayerNorm + ResidualStabilize deep trainingPre-Norm layout

2. Complete Model Code (PyTorch) ​

python
import torch
import torch.nn as nn
from torch.nn import functional as F

# ── Hyperparameters (minimal config) ───────────────────────────
batch_size = 32      # Number of sequences per batch
block_size = 128     # Context length: each sequence can look at up to 128 previous tokens
max_iters = 3000     # Total training steps
eval_interval = 300  # Validate every 300 steps
learning_rate = 3e-4 # Learning rate
n_layer = 2          # Number of Transformer layers
n_head = 4           # Number of attention heads
n_embd = 64          # Embedding dimension (model width)
dropout = 0.0        # Small models don't benefit much from dropout

class CausalSelfAttention(nn.Module):
    """Single-layer multi-head causal self-attention: QKV + scaled dot-product + causal mask"""
    def __init__(self):
        super().__init__()
        assert n_embd % n_head == 0
        self.c_attn = nn.Linear(n_embd, 3 * n_embd)      # Compute Q, K, V in one pass
        self.c_proj = nn.Linear(n_embd, n_embd)          # Output projection
        self.n_head = n_head
        self.n_embd = n_embd

    def forward(self, x):
        B, T, C = x.shape                                # B batch, T seq len, C channels
        qkv = self.c_attn(x)                             # (B, T, 3C)
        q, k, v = qkv.split(self.n_embd, dim=2)
        # Split into heads: channels per head = C/n_head
        q = q.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
        k = k.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
        v = v.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)

        # Scaled dot-product attention; att = softmax(QK^T / sqrt(d_k)) V
        att = (q @ k.transpose(-2, -1)) * (1.0 / (k.shape[-1] ** 0.5))
        # Causal mask: lower triangular matrix, fill upper-right with -inf, softmax makes them 0
        mask = torch.tril(torch.ones(T, T, device=x.device)).view(1, 1, T, T)
        att = att.masked_fill(mask == 0, float("-inf"))
        att = F.softmax(att, dim=-1)
        y = att @ v                                     # (B, n_head, T, head_dim)
        y = y.transpose(1, 2).contiguous().view(B, T, C)  # Concatenate heads back
        return self.c_proj(y)


class MLP(nn.Module):
    """Per-position feed-forward network (simple GELU activation here)"""
    def __init__(self):
        super().__init__()
        self.c_fc = nn.Linear(n_embd, 4 * n_embd)
        self.c_proj = nn.Linear(4 * n_embd, n_embd)

    def forward(self, x):
        return self.c_proj(F.gelu(self.c_fc(x)))


class Block(nn.Module):
    """One Transformer layer = LayerNorm + Attention + LayerNorm + MLP (Pre-Norm)"""
    def __init__(self):
        super().__init__()
        self.ln1 = nn.LayerNorm(n_embd)
        self.attn = CausalSelfAttention()
        self.ln2 = nn.LayerNorm(n_embd)
        self.mlp = MLP()

    def forward(self, x):
        x = x + self.attn(self.ln1(x))   # Residual + Pre-Norm
        x = x + self.mlp(self.ln2(x))
        return x


class GPT(nn.Module):
    """Minimal GPT: embeddings + N Blocks + final linear layer + loss"""
    def __init__(self):
        super().__init__()
        self.token_embedding = nn.Embedding(vocab_size, n_embd)
        self.position_embedding = nn.Embedding(block_size, n_embd)  # Absolute positional encoding
        self.blocks = nn.Sequential(*[Block() for _ in range(n_layer)])
        self.ln_f = nn.LayerNorm(n_embd)
        self.lm_head = nn.Linear(n_embd, vocab_size)   # Score each token

    def forward(self, idx, targets=None):
        B, T = idx.shape
        tok_emb = self.token_embedding(idx)                          # (B, T, n_embd)
        pos = torch.arange(T, device=idx.device)                     # Positions 0..T-1
        pos_emb = self.position_embedding(pos)                       # (T, n_embd)
        x = tok_emb + pos_emb                                        # Broadcast add
        x = self.blocks(x)
        x = self.ln_f(x)
        logits = self.lm_head(x)                                     # (B, T, vocab_size)

        loss = None
        if targets is not None:
            # Cross-entropy: predict "next token" at each position
            B, T, V = logits.shape
            logits = logits.view(B * T, V)
            targets = targets.view(B * T)
            loss = F.cross_entropy(logits, targets)
        return logits, loss

    def generate(self, idx, max_new_tokens):
        """Autoregressive sampling: append generated token, predict next"""
        for _ in range(max_new_tokens):
            idx_cond = idx[:, -block_size:]                          # Only take the last block_size
            logits, _ = self.forward(idx_cond)
            logits = logits[:, -1, :]                                # Look only at last position
            probs = F.softmax(logits, dim=-1)
            idx_next = torch.multinomial(probs, num_samples=1)       # Sample by probability
            idx = torch.cat((idx, idx_next), dim=1)
        return idx

Reading this code line by line is worth copying ten times. Three details most worth pausing on:

  1. Causal mask (torch.tril): Ensures position T can only see itself and previous tokens. This is the architecturalization of the "next-word prediction" discipline — leaking the future ruins learning.
  2. Positional encoding (position_embedding): Attention itself doesn't perceive order; positional information must be explicitly injected. Real models mostly use RoPE or other relative positional encodings, see Transformer Architecture Explained.
  3. Autoregressive generate: Generate one token → append it → predict again. This is the code form of the "autoregressive generation loop" from Inference Fundamentals.

Five, Training Loop: Watching the Loss Drop ​

1. Data Loader ​

python
import torch

data = torch.tensor(encode(text), dtype=torch.long)
n_data = len(data)

def get_batch(split):
    """Randomly cut a fixed-length sequence from train/val data, construct (input, target) pairs"""
    src = train_data if split == "train" else val_data
    src = torch.tensor(encode(src), dtype=torch.long)
    ix = torch.randint(len(src) - block_size, (batch_size,))
    x = torch.stack([src[i: i + block_size] for i in ix])
    y = torch.stack([src[i + 1: i + 1 + block_size] for i in ix])   # target = input shifted right by one
    return x, y

The target is the input shifted right by one — this single line of code is the most condensed expression of the entire LLM training objective: given x[i], make the model predict x[i+1].

2. Main Training Loop ​

python
torch.manual_seed(1337)
model = GPT()
optimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate)

@torch.no_grad()
def estimate_loss():
    """Evaluate average loss once on train/val data"""
    out = {}
    model.eval()
    for split in ["train", "val"]:
        losses = torch.zeros(eval_interval)
        for k in range(eval_interval):
            x, y = get_batch(split)
            _, loss = model(x, y)
            losses[k] = loss.item()
        out[split] = losses.mean().item()
    model.train()
    return out

for step in range(max_iters):
    x, y = get_batch("train")
    _, loss = model(x, y)
    optimizer.zero_grad(set_to_none=True)
    loss.backward()
    optimizer.step()

    if step % eval_interval == 0 or step == max_iters - 1:
        losses = estimate_loss()
        print(f"step {step:5d} | train loss {losses['train']:.4f} | val loss {losses['val']:.4f}")

3. How to Read the Loss Curve ​

Typical output (illustrative; exact numbers vary by random seed):

step     0 | train loss 4.2236 | val loss 4.2176
step   300 | train loss 2.4113 | val loss 2.5021
step   600 | train loss 1.8928 | val loss 2.0815
step   900 | train loss 1.6229 | val loss 1.8931
step  1200 | train loss 1.4651 | val loss 1.7893
...
step  3000 | train loss 1.2310 | val loss 1.6492

Three readings you must master:

PhenomenonMeaningResponse
train loss and val loss both decreaseModel is truly learning, no obvious overfittingKeep training
train loss drops low, val loss stalls / risesOverfitting: model is memorizing training textEarly stop, add dropout, add data
Both can't go down (loss stuck high)Underfitting or wrong hyperparametersIncrease model size / learning rate, check data
Initial loss ≈ ln(65) ≈ 4.17Model starts as "uniformly guessing"This is a standard sanity check, not a bug

Why initial loss ≈ ln(vocab_size)

A randomly initialized model gives approximately uniform distribution over each token, and the cross-entropy loss is -ln(1/65) ≈ 4.17. If your initial loss is far from this value, there's likely a bug in your code (e.g., wrong vocabulary count, target not shifted right).

Train loss can drop to 1.2 — does this mean the model has "memorized" Shakespeare?

Not entirely. A val loss of ~1.6 means the model is still, on average, "confused" about about e^1.6 ≈ 5 candidate tokens — it has learned the statistical structure of the text (word order, grammar, character dialogue patterns), not memorized the original text character by character. True "memorization" becomes obvious only when parameters vastly exceed data size.

Training Optimization Toolkit: Making 1M Models Converge More Steadily ​

The minimal training loop in this article has no optimization tricks but already works. When you start "training seriously" (bigger model, longer data), add the toolkit in order:

python
# 1. Learning rate warmup: linearly ramp to target lr for the first N steps, stabilizing early training
# 2. Weight decay: apply L2 regularization to non-bias/non-norm parameters
# 3. Cosine decay: lr smoothly decays from peak to near-zero, more stable convergence in the final phase

import math
def get_lr(step, warmup_iters=200, lr_decay_iters=3000, max_lr=6e-4, min_lr=6e-5):
    """Learning rate schedule in nanoGPT style: warmup + cosine decay"""
    if step < warmup_iters:
        return max_lr * (step + 1) / warmup_iters
    if step > lr_decay_iters:
        return min_lr
    decay_ratio = (step - warmup_iters) / (lr_decay_iters - warmup_iters)
    coeff = 0.5 * (1.0 + math.cos(math.pi * decay_ratio))
    return min_lr + coeff * (max_lr - min_lr)

# Usage: in the training loop, optimizer.param_groups[0]["lr"] = get_lr(step)

This matches the schedule in the GPT-2 paper/training code, and is the first batch of engineering changes "from toy to real model" (detailed principles in the training dynamics section of Pretraining).

Six, Sampling Generation: Watch It Talk ​

python
# Start from a "newline" (equivalent to clearing context), generate 500 tokens
context = torch.tensor([[encode("\n")[0]]], dtype=torch.long)
print(decode(model.generate(context, max_new_tokens=500)[0].tolist()))

A typical (illustrative) output — note that spelling, punctuation, and dialogue structure are already "looking legit":

HENRY:
If you be to, the soul of my son,
The great and the gracious this sword
Hath brought my body to this gentle king.

GLOUCESTER:
What is he? What were his father's grief,
And let me speak of me to the people.

Sampling is an inference strategy

The torch.multinomial here samples randomly from the probability distribution (temperature=1). If you want more deterministic output, adjust temperature, top-k, top-p, and other decoding strategies — this is the code prototype of the sampling strategies in Inference Fundamentals. Divide the logits of probs by temperature before softmax to control the "randomness vs. determinism" of the output.

Seven, From 1M to 1B: Extension Checklist ​

After training a 1M small model, you understand all mechanisms. Every step of "getting bigger" requires solving new problems simultaneously:

ScaleParameter MagnitudeRequired UpgradesRelated Content
Toy~1M (this article)Nothing, runs on CPUThis article
Entry-level10M~50MSwap data to OpenWebText; add layers/width (e.g., 4 layers, 256-dim); GPU trainingPretraining
Advanced100M~300MLR warmup + cosine schedule; gradient clipping; VRAM optimization; use tiktoken/BPE tokenizationPretraining, Tokenization & Vocabularies
Research-grade1B+Multi-GPU data/tensor parallelism; mixed precision; larger data mix ratiosScaling Laws, Framework & Tool Selection

By the nanoGPT official repo: config/train_gpt2.py is a complete config for standard GPT-2 (124M parameters), training data is OpenWebText, using multi-GPU + mixed precision — it's the best springboard from this article to "real pretraining" (References).

Eight, Common Pitfalls ​

PitfallSymptomsDebug Direction
Target not shifted rightModel can "remember" input, loss won't dropCheck that alignment like y = x[:, 1:] is written correctly
Causal mask written wrongTraining loss abnormally low but generates garbagePrint the mask matrix separately to check if it's lower triangular
Context exceeds block_sizeShape errors during generationgenerate must slice the last block_size tokens
Data not splitval loss ≈ train lossConfirm validation text is independent from training text
Vocabulary counted wrongInconsistent character sets between train/valUse train_data to count the vocabulary; new chars in validation will crash
Initial loss wrongNot around ln(vocab_size)Check data loading and batch construction

To dig deeper into any step: the intuition behind loss and perplexity is in Language Modeling, attention details in Transformer Architecture Explained, data mixing and cleaning in Pretraining, and if you want to fine-tune into a specific style after training, see Fine-Tuning in Practice.

Further Reading ​

References ​