Theme
Building a Large Model from Scratch
Reading ten papers is not as good as training one model yourself. This article takes you through nanoGPT's minimal path, writing a small GPT that truly runs and truly generates text: data → tokenizer → model → training → sampling, five steps, none skipped.
Many people's first reaction to large models is "that's something only big companies can touch." But the fact is: you can train a small GPT that produces coherent text with an ordinary laptop GPU (or even pure CPU). This article compresses the process to the minimal runnable version — all code combined is under two hundred lines, and everything is provided. What you get is not a "toy," but a hands-on verifiable mental model for understanding how GPT models work: from then on, when you see the term "large model," you'll have in mind its data flow, loss curves, and sampling process, not just a black box.
One, Roadmap Overview: Five Steps + Four Acceptance Criteria
The complete roadmap and the core question each step answers:
Shakespeare's complete works (approx. 1MB of raw text)
│
▼
① Data preparation ──────── Answers "what kind of text is the model learning"
▼
② Tokenizer ─────────────── Answers "how text gets chunked into model input units"
▼
③ Model definition ──────── Answers "what the network looks like for predicting the next token"
▼
④ Training loop ─────────── Answers "how loss decreases step by step and the model gets smarter"
▼
⑤ Sampling / generation ─── Answers "what a trained model can output"This is the minimal closed loop of the language modeling paradigm: the model's only task is to predict the next token. Understanding this matters more than memorizing any formula.
Four Acceptance Criteria
Before you start, define what "success" looks like so you don't finish training and not know how to evaluate results:
| # | Acceptance Item | Threshold (with this config) | How to Check |
|---|---|---|---|
| ① | Code runs | Training loop doesn't error; loss drops from 4~5 | Observe the printed train loss |
| ② | Loss drops noticeably | Train loss goes from ~4.2 to under 1.8 within 3000 steps | Compare against training curves (see Section Five) |
| ③ | Overfitting signal is reasonable | val loss drops then rises, train loss continues to drop | Compare train / val loss curves |
| ④ | Generated text is readable | Sampled output is a sequence of "English / Shakespeare-like" words | Run the sampling function and read it |
What you're accepting is not "looking like ChatGPT"
A small GPT (~1M parameters) cannot produce logically coherent dialogue. Its success standard is: it has learned Shakespeare's word choice habits, punctuation rhythm, and partial spelling patterns. This is actually the best teaching moment — you can clearly see how "statistical patterns" emerge from data. To verify a smarter model, go through the extension checklist in Section Seven first.
Local Resource Budget
| Config | VRAM / Memory | Estimated Training Time for 1M Model | What It's Good For |
|---|---|---|---|
| Pure CPU (e.g., laptop 8-core) | Only ~2GB RAM needed | ~20~40 minutes | Everything in this article, just be patient |
| Entry GPU (e.g., RTX 3060 12GB) | ~2GB VRAM | ~3~5 minutes | This article + parameter tweaking |
| Mid-range GPU (e.g., RTX 4090 24GB) | ~2GB VRAM | ~1~2 minutes | Fast iteration; continue toward 10M~50M parameters |
| Cloud GPU (e.g., A100 40GB) | No pressure | Seconds | Not necessary for this article; save for the Section Seven extensions |
A harsh comparison
GPT-2 (1.5B parameters) needs days of pre-training across 8 A100s, while the 1M model in this article can be trained on a CPU in half an hour — the gap brought by scale is orders of magnitude. This is a real-world footnote to scaling laws: a large model's "intelligence" is largely the result of piling compute and data, and this article lets you experience firsthand "why small models fall short."
Two, Data: Starting with Shakespeare
1. Why Shakespeare
Karpathy's nanoGPT tutorial chooses Shakespeare's complete plays (about 1MB of plain text) as the starting point for highly practical reasons:
- Small enough: 1MB loads in seconds on a CPU; any machine can train it.
- Patterned enough: Dramatic text has character, dialogue, and scene structure with clear statistical patterns that small models can learn to mimic.
- Interesting enough: The generated "fake Shakespeare" is intuitively readable, giving a sense of accomplishment at verification time.
Download the data:
bash
# Directly download tiny Shakespeare (the data file from Karpathy's char-rnn repo)
mkdir -p data/shakespeare
wget https://raw.githubusercontent.com/karpathy/char-rnn/master/data/tinyshakespeare/input.txt -O data/shakespeare/input.txt
# If Windows lacks wget, use curl: curl -o data/shakespeare/input.txt <same URL>python
# Load and quick health check
with open("data/shakespeare/input.txt", "r", encoding="utf-8") as f:
text = f.read()
print(f"Total characters: {len(text)}") # ~1.1M
print(text[:500]) # First 500 characters, get a feelIf you want to level up to "more like real pre-training," you can swap in OpenWebText (the unofficial recreation of GPT-2's training set, see References), but that's a "from 1M to 1B" step — we'll leave that for later.
2. Data Split: Training / Validation
Consistent with the machine learning manual's discipline of "touch the test set only once at the end" (see Common Pitfalls and Anti-Patterns), LLM training also requires an independent validation set — but note: an LLM's validation set is "unseen text segments from the same document," not a separate document. This is a key difference between text data and tabular data.
python
# Use 90% of text for training, 10% for validation (don't shuffle randomly — maintaining text order matters for language models)
n = len(text)
train_data = text[: int(n * 0.9)]
val_data = text[int(n * 0.9):]Why no shuffling here
Text sequences have an inherent order (what comes before determines what comes after). If you randomly shuffle characters like tabular data, the model will learn "the probability distribution of random strings," which is meaningless. During training, we cut sequential batches, but the positions within each batch are randomized (see Section Four), ensuring every step's gradient comes from different regions of the text.
Three, Tokenization: From Character-Level to BPE
1. Character-Level Tokenizer: Best for Teaching
The tokenization mechanism of large models is detailed in Tokenization & Vocabularies. We start with the simplest character-level tokenization: each character is one token. This has two advantages — minimal code and clearest concepts.
python
# Count all distinct characters in the text as the vocabulary
chars = sorted(list(set(train_data)))
vocab_size = len(chars)
print(f"Vocabulary size: {vocab_size}") # Usually 65 (letters + punctuation + spaces, etc.)
# Build bidirectional character <-> integer mapping
stoi = {ch: i for i, ch in enumerate(chars)} # char -> id
itos = {i: ch for i, ch in enumerate(chars)} # id -> char
def encode(s):
"""Convert a string to a list of integers"""
return [stoi[c] for c in s]
def decode(ids):
"""Convert a list of integers back to a string"""
return "".join(itos[i] for i in ids)
# Self-check: encoding then decoding should restore the original text
sample = "To be, or not to be"
assert decode(encode(sample)) == sample
print("Tokenizer self-check passed")2. Upgrading from Character-Level to BPE
A character-level vocabulary has only 65 tokens, but the cost is very low information per token — the model needs more steps to learn the "word" concept. Real large models all use subword tokenization like BPE, with vocabularies typically of 30K~100K tokens (Llama 3 uses 128K, GPT-4 uses about 100K). You can try it in one line with OpenAI's open-source tiktoken:
python
pip install tiktoken
import tiktoken
enc = tiktoken.get_encoding("cl100k_base") # Tokenizer used by GPT-4
ids = enc.encode("To be, or not to be")
print(ids) # [904, 471, ..., 471]
print(enc.decode(ids)) # Restores the original text
print("Token count:", len(ids), "vs. char count:", len("To be, or not to be"))This comparison (how many tokens a piece of English gets chunked into) directly determines API call costs — tokens are the pricing unit in the LLM world. Conversion factors are covered in Context Windows & Long Text.
How to choose for this article
All code in the next four sections is based on character-level tokenizers (to guarantee minimal runnable code). Swap the encode/decode from Section Three Part 1 with tiktoken's enc.encode/enc.decode, and the model can ingest BPE tokens — you only need to change two lines of code. This is a concrete demonstration of "tokenization is an external component of the model."
Four, Model: Minimal GPT (Complete Runnable Code)
1. Architecture Overview
The model itself is the decoder-only subset of Transformer Architecture Explained, with four core components:
| Component | Role | Notes |
|---|---|---|
| TokenEmbedding | Maps token IDs to dense vectors | Learnable lookup table |
| CausalSelfAttention | Let each position only see itself and previous tokens | Causal mask is key to GPT |
| MLP (Feed-Forward) | Non-linear transformation per position | Each token computed independently |
| LayerNorm + Residual | Stabilize deep training | Pre-Norm layout |
2. Complete Model Code (PyTorch)
python
import torch
import torch.nn as nn
from torch.nn import functional as F
# ── Hyperparameters (minimal config) ───────────────────────────
batch_size = 32 # Number of sequences per batch
block_size = 128 # Context length: each sequence can look at up to 128 previous tokens
max_iters = 3000 # Total training steps
eval_interval = 300 # Validate every 300 steps
learning_rate = 3e-4 # Learning rate
n_layer = 2 # Number of Transformer layers
n_head = 4 # Number of attention heads
n_embd = 64 # Embedding dimension (model width)
dropout = 0.0 # Small models don't benefit much from dropout
class CausalSelfAttention(nn.Module):
"""Single-layer multi-head causal self-attention: QKV + scaled dot-product + causal mask"""
def __init__(self):
super().__init__()
assert n_embd % n_head == 0
self.c_attn = nn.Linear(n_embd, 3 * n_embd) # Compute Q, K, V in one pass
self.c_proj = nn.Linear(n_embd, n_embd) # Output projection
self.n_head = n_head
self.n_embd = n_embd
def forward(self, x):
B, T, C = x.shape # B batch, T seq len, C channels
qkv = self.c_attn(x) # (B, T, 3C)
q, k, v = qkv.split(self.n_embd, dim=2)
# Split into heads: channels per head = C/n_head
q = q.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
k = k.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
v = v.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
# Scaled dot-product attention; att = softmax(QK^T / sqrt(d_k)) V
att = (q @ k.transpose(-2, -1)) * (1.0 / (k.shape[-1] ** 0.5))
# Causal mask: lower triangular matrix, fill upper-right with -inf, softmax makes them 0
mask = torch.tril(torch.ones(T, T, device=x.device)).view(1, 1, T, T)
att = att.masked_fill(mask == 0, float("-inf"))
att = F.softmax(att, dim=-1)
y = att @ v # (B, n_head, T, head_dim)
y = y.transpose(1, 2).contiguous().view(B, T, C) # Concatenate heads back
return self.c_proj(y)
class MLP(nn.Module):
"""Per-position feed-forward network (simple GELU activation here)"""
def __init__(self):
super().__init__()
self.c_fc = nn.Linear(n_embd, 4 * n_embd)
self.c_proj = nn.Linear(4 * n_embd, n_embd)
def forward(self, x):
return self.c_proj(F.gelu(self.c_fc(x)))
class Block(nn.Module):
"""One Transformer layer = LayerNorm + Attention + LayerNorm + MLP (Pre-Norm)"""
def __init__(self):
super().__init__()
self.ln1 = nn.LayerNorm(n_embd)
self.attn = CausalSelfAttention()
self.ln2 = nn.LayerNorm(n_embd)
self.mlp = MLP()
def forward(self, x):
x = x + self.attn(self.ln1(x)) # Residual + Pre-Norm
x = x + self.mlp(self.ln2(x))
return x
class GPT(nn.Module):
"""Minimal GPT: embeddings + N Blocks + final linear layer + loss"""
def __init__(self):
super().__init__()
self.token_embedding = nn.Embedding(vocab_size, n_embd)
self.position_embedding = nn.Embedding(block_size, n_embd) # Absolute positional encoding
self.blocks = nn.Sequential(*[Block() for _ in range(n_layer)])
self.ln_f = nn.LayerNorm(n_embd)
self.lm_head = nn.Linear(n_embd, vocab_size) # Score each token
def forward(self, idx, targets=None):
B, T = idx.shape
tok_emb = self.token_embedding(idx) # (B, T, n_embd)
pos = torch.arange(T, device=idx.device) # Positions 0..T-1
pos_emb = self.position_embedding(pos) # (T, n_embd)
x = tok_emb + pos_emb # Broadcast add
x = self.blocks(x)
x = self.ln_f(x)
logits = self.lm_head(x) # (B, T, vocab_size)
loss = None
if targets is not None:
# Cross-entropy: predict "next token" at each position
B, T, V = logits.shape
logits = logits.view(B * T, V)
targets = targets.view(B * T)
loss = F.cross_entropy(logits, targets)
return logits, loss
def generate(self, idx, max_new_tokens):
"""Autoregressive sampling: append generated token, predict next"""
for _ in range(max_new_tokens):
idx_cond = idx[:, -block_size:] # Only take the last block_size
logits, _ = self.forward(idx_cond)
logits = logits[:, -1, :] # Look only at last position
probs = F.softmax(logits, dim=-1)
idx_next = torch.multinomial(probs, num_samples=1) # Sample by probability
idx = torch.cat((idx, idx_next), dim=1)
return idxReading this code line by line is worth copying ten times. Three details most worth pausing on:
- Causal mask (
torch.tril): Ensures position T can only see itself and previous tokens. This is the architecturalization of the "next-word prediction" discipline — leaking the future ruins learning. - Positional encoding (
position_embedding): Attention itself doesn't perceive order; positional information must be explicitly injected. Real models mostly use RoPE or other relative positional encodings, see Transformer Architecture Explained. - Autoregressive
generate: Generate one token → append it → predict again. This is the code form of the "autoregressive generation loop" from Inference Fundamentals.
Five, Training Loop: Watching the Loss Drop
1. Data Loader
python
import torch
data = torch.tensor(encode(text), dtype=torch.long)
n_data = len(data)
def get_batch(split):
"""Randomly cut a fixed-length sequence from train/val data, construct (input, target) pairs"""
src = train_data if split == "train" else val_data
src = torch.tensor(encode(src), dtype=torch.long)
ix = torch.randint(len(src) - block_size, (batch_size,))
x = torch.stack([src[i: i + block_size] for i in ix])
y = torch.stack([src[i + 1: i + 1 + block_size] for i in ix]) # target = input shifted right by one
return x, yThe target is the input shifted right by one — this single line of code is the most condensed expression of the entire LLM training objective: given x[i], make the model predict x[i+1].
2. Main Training Loop
python
torch.manual_seed(1337)
model = GPT()
optimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate)
@torch.no_grad()
def estimate_loss():
"""Evaluate average loss once on train/val data"""
out = {}
model.eval()
for split in ["train", "val"]:
losses = torch.zeros(eval_interval)
for k in range(eval_interval):
x, y = get_batch(split)
_, loss = model(x, y)
losses[k] = loss.item()
out[split] = losses.mean().item()
model.train()
return out
for step in range(max_iters):
x, y = get_batch("train")
_, loss = model(x, y)
optimizer.zero_grad(set_to_none=True)
loss.backward()
optimizer.step()
if step % eval_interval == 0 or step == max_iters - 1:
losses = estimate_loss()
print(f"step {step:5d} | train loss {losses['train']:.4f} | val loss {losses['val']:.4f}")3. How to Read the Loss Curve
Typical output (illustrative; exact numbers vary by random seed):
step 0 | train loss 4.2236 | val loss 4.2176
step 300 | train loss 2.4113 | val loss 2.5021
step 600 | train loss 1.8928 | val loss 2.0815
step 900 | train loss 1.6229 | val loss 1.8931
step 1200 | train loss 1.4651 | val loss 1.7893
...
step 3000 | train loss 1.2310 | val loss 1.6492Three readings you must master:
| Phenomenon | Meaning | Response |
|---|---|---|
| train loss and val loss both decrease | Model is truly learning, no obvious overfitting | Keep training |
| train loss drops low, val loss stalls / rises | Overfitting: model is memorizing training text | Early stop, add dropout, add data |
| Both can't go down (loss stuck high) | Underfitting or wrong hyperparameters | Increase model size / learning rate, check data |
| Initial loss ≈ ln(65) ≈ 4.17 | Model starts as "uniformly guessing" | This is a standard sanity check, not a bug |
Why initial loss ≈ ln(vocab_size)
A randomly initialized model gives approximately uniform distribution over each token, and the cross-entropy loss is -ln(1/65) ≈ 4.17. If your initial loss is far from this value, there's likely a bug in your code (e.g., wrong vocabulary count, target not shifted right).
Train loss can drop to 1.2 — does this mean the model has "memorized" Shakespeare?
Not entirely. A val loss of ~1.6 means the model is still, on average, "confused" about about e^1.6 ≈ 5 candidate tokens — it has learned the statistical structure of the text (word order, grammar, character dialogue patterns), not memorized the original text character by character. True "memorization" becomes obvious only when parameters vastly exceed data size.
Training Optimization Toolkit: Making 1M Models Converge More Steadily
The minimal training loop in this article has no optimization tricks but already works. When you start "training seriously" (bigger model, longer data), add the toolkit in order:
python
# 1. Learning rate warmup: linearly ramp to target lr for the first N steps, stabilizing early training
# 2. Weight decay: apply L2 regularization to non-bias/non-norm parameters
# 3. Cosine decay: lr smoothly decays from peak to near-zero, more stable convergence in the final phase
import math
def get_lr(step, warmup_iters=200, lr_decay_iters=3000, max_lr=6e-4, min_lr=6e-5):
"""Learning rate schedule in nanoGPT style: warmup + cosine decay"""
if step < warmup_iters:
return max_lr * (step + 1) / warmup_iters
if step > lr_decay_iters:
return min_lr
decay_ratio = (step - warmup_iters) / (lr_decay_iters - warmup_iters)
coeff = 0.5 * (1.0 + math.cos(math.pi * decay_ratio))
return min_lr + coeff * (max_lr - min_lr)
# Usage: in the training loop, optimizer.param_groups[0]["lr"] = get_lr(step)This matches the schedule in the GPT-2 paper/training code, and is the first batch of engineering changes "from toy to real model" (detailed principles in the training dynamics section of Pretraining).
Six, Sampling Generation: Watch It Talk
python
# Start from a "newline" (equivalent to clearing context), generate 500 tokens
context = torch.tensor([[encode("\n")[0]]], dtype=torch.long)
print(decode(model.generate(context, max_new_tokens=500)[0].tolist()))A typical (illustrative) output — note that spelling, punctuation, and dialogue structure are already "looking legit":
HENRY:
If you be to, the soul of my son,
The great and the gracious this sword
Hath brought my body to this gentle king.
GLOUCESTER:
What is he? What were his father's grief,
And let me speak of me to the people.Sampling is an inference strategy
The torch.multinomial here samples randomly from the probability distribution (temperature=1). If you want more deterministic output, adjust temperature, top-k, top-p, and other decoding strategies — this is the code prototype of the sampling strategies in Inference Fundamentals. Divide the logits of probs by temperature before softmax to control the "randomness vs. determinism" of the output.
Seven, From 1M to 1B: Extension Checklist
After training a 1M small model, you understand all mechanisms. Every step of "getting bigger" requires solving new problems simultaneously:
| Scale | Parameter Magnitude | Required Upgrades | Related Content |
|---|---|---|---|
| Toy | ~1M (this article) | Nothing, runs on CPU | This article |
| Entry-level | 10M~50M | Swap data to OpenWebText; add layers/width (e.g., 4 layers, 256-dim); GPU training | Pretraining |
| Advanced | 100M~300M | LR warmup + cosine schedule; gradient clipping; VRAM optimization; use tiktoken/BPE tokenization | Pretraining, Tokenization & Vocabularies |
| Research-grade | 1B+ | Multi-GPU data/tensor parallelism; mixed precision; larger data mix ratios | Scaling Laws, Framework & Tool Selection |
By the nanoGPT official repo: config/train_gpt2.py is a complete config for standard GPT-2 (124M parameters), training data is OpenWebText, using multi-GPU + mixed precision — it's the best springboard from this article to "real pretraining" (References).
Eight, Common Pitfalls
| Pitfall | Symptoms | Debug Direction |
|---|---|---|
| Target not shifted right | Model can "remember" input, loss won't drop | Check that alignment like y = x[:, 1:] is written correctly |
| Causal mask written wrong | Training loss abnormally low but generates garbage | Print the mask matrix separately to check if it's lower triangular |
Context exceeds block_size | Shape errors during generation | generate must slice the last block_size tokens |
| Data not split | val loss ≈ train loss | Confirm validation text is independent from training text |
| Vocabulary counted wrong | Inconsistent character sets between train/val | Use train_data to count the vocabulary; new chars in validation will crash |
| Initial loss wrong | Not around ln(vocab_size) | Check data loading and batch construction |
To dig deeper into any step: the intuition behind loss and perplexity is in Language Modeling, attention details in Transformer Architecture Explained, data mixing and cleaning in Pretraining, and if you want to fine-tune into a specific style after training, see Fine-Tuning in Practice.
Further Reading
- Language Modeling: Next-Word Prediction Paradigm —— The math and intuition behind this article's model: perplexity, cross-entropy, why predicting the next word can learn knowledge
- Transformer Architecture Explained —— Scaling this 200-line code into a full architecture: multi-head attention, RoPE, SwiGLU
- Pretraining: Data & Objectives —— From Shakespeare to TB-scale corpora: cleaning, deduplication, mixing
- Tokenization & Vocabularies —— The complete map from character-level to BPE to SentencePiece
- Scaling Laws —— This article's 1M model vs GPT-3 175B: how loss decreases as a power law with scale
- Classic Papers Deep Dive —— How to read Attention Is All You Need and GPT series papers paragraph by paragraph
References
- Karpathy/nanoGPT (GitHub) —— The official source of this article's code; repo contains both Shakespeare and GPT-2 training configs
- Let's build GPT: from scratch (Karpathy video course) —— The video version of this article's content, highly recommended to watch alongside
- tiny Shakespeare data file —— The training corpus used in this article
- OpenWebTextCorpus (GitHub) —— An unofficial recreation of GPT-2's training WebText, the next dataset for the extension phase
- Attention Is All You Need (arXiv:1706.03762) —— The original Transformer paper, source of this article's attention implementation
- Language Models are Unsupervised Multitask Learners (GPT-2 paper) —— Original GPT-2 and WebText documentation