Theme
Learning Paths: Three Routes
In one sentence: Learning Paths is the site's map — it doesn't answer "what is an LLM" (that's What Is a Large Language Model's job), nor does it dive into any specific technology. Instead, it answers three more fundamental questions: In what order should I read? How deep should I go at each stage? How do I know I've actually learned it?
This site has seven modules (Introduction / Core Knowledge / Case Studies / Paper Deep Dives / Practice Guide / Careers & JD / Resources) with roughly 40 pages in total. No one can realistically read everything from start to finish, nor should they. These three routes help you cut through the noise: the same mountain, three different paths up.
I. Overview of the Three Routes
| Dimension | Route A: Career Sprint | Route B: Systematic Deep Dive | Route C: Desk Reference |
|---|---|---|---|
| Target audience | Preparing for interviews, short on time | Want to build a systematic knowledge framework | Already working with LLMs |
| Core objective | Can answer interview questions, can talk about projects | Clear on concepts, can reproduce implementations | Quickly find answers when problems arise |
| Suggested duration | 4–6 weeks × 3–4 hrs/day | 8 weeks × 10–15 hrs/week | No fixed schedule, look up as needed |
| Core pages | Concept pages + interview question bank | Concept pages + case pages + practice pages | Practice pages + resource pages |
| Reading depth | Need conclusions, can retell them | Need derivations, can break things down | Need answers, can follow along |
| Verification criteria | Get ~80% of interview questions right | Can explain any core concept from scratch | No longer need to ask others for help |
text
The three routes are not three levels — they are three ways to invest:
Career Sprint: Concept ──→ Interview Qs ──→ Can answer
Systematic Deep Dive: Concept ──→ Case ──→ Practice ──→ Can build
Desk Reference: Problem ──→ Page ──→ Answer ──→ Can look upHow to choose?
Be clear about "what will I deliver in a month." Interviewing next month → pick A. Running a fine-tuning/RAG project next month → pick B. Shipping a service next month → start with C, and fill in B's concept pages when you hit a bottleneck. A and B are not mutually exclusive: most serious job seekers actually walk a simplified version of B plus A's interview intensification. This article presents B as the main thread with a complete 8-week plan, A as its trimmed version, and C as its reverse lookup index.
II. Route A: Career Sprint (4–6 Weeks)
1. Goals and Principles
Career Sprint has the best return on investment, but only if you focus on the big picture and skip the details: read only the high-frequency interview topics, skip the obscure stuff. High-frequency interview coverage for LLM positions spans four blocks — architecture, training, inference, and applications — see the interview question bank. The sprint principle is "be able to explain any topic in your own words for two minutes," not "have read the most pages."
2. Four-Week Schedule
| Week | Topic | Required Pages | Daily Output |
|---|---|---|---|
| W1 | Concepts and foundations | What Is a Large Language Model, Language Modeling, Tokenization and Vocabularies | Define "what is an LLM" in one paragraph; hand-draw the next-token prediction flow |
| W2 | Architecture and training | Transformer Architecture, Pretraining, Scaling Laws | Derive the QKV computation flow; explain why the scaling factor is √d_k |
| W3 | Post-training and alignment | Fine-tuning, Alignment, Inference Fundamentals | Recite LoRA's W=W₀+BA; draw the three steps of RLHF |
| W4 | Applications and sprint | RAG, Agents with LLMs, Interview Question Bank | Go through 10 interview questions daily; write project experience as "background → challenge → solution → result" |
3. Verification Criteria
- Without notes, you can draw the Transformer structure and label Q, K, V, residuals, and LayerNorm;
- You can explain the SFT → reward model → PPO three steps of RLHF, and say why DPO can skip the reward model;
- You can explain what problems RAG and fine-tuning each solve, and when to choose which (see RAG vs. long context trade-offs);
- Mock interviews: pick 10 random questions from the interview question bank and answer 8 correctly on the spot.
The most common pitfall of Career Sprint
Spending time on "memorizing more jargon" rather than "clearly explaining concepts you already know." Interviewers probe depth through follow-up questions, and follow-ups rely on mechanism understanding, not jargon volume. It's better to read one page on Transformers three times than to half-swallow ten pages.
4. Quick Reference: High-Frequency Interview Topics
The topic map corresponding to the four-week route (a compressed version of the interview question bank, for self-check):
| Block | High-frequency topics | Related pages |
|---|---|---|
| Architecture | QKV computation and scaling factor √d_k, motivation for multi-head attention, positional encoding extrapolation, causal masking | Transformer |
| Training | Next-token objective and cross-entropy, data mixing, scaling laws, LoRA's W=W₀+BA | Language Modeling, Scaling Laws, Fine-tuning |
| Alignment | RLHF three steps, reward model training, PPO and KL penalty, why DPO can skip the reward model | Alignment |
| Inference | Autoregression and KV Cache, temperature and top-p, quantization, continuous batching, memory estimation | Inference Fundamentals, Deployment |
| Applications | RAG three stages, Agent loops, prompt injection and security | RAG, Agent |
Usage: pick one row daily and "explain it in your own words for 2 minutes." If you can't, go back to the corresponding page. After four weeks, this table becomes your interview outline.
III. Route B: Systematic Deep Dive (8 Weeks)
This is the main thread of this handbook. It is organized in the order of "foundation → architecture → training → alignment → evaluation → applications → deployment → wrap-up," corresponding one-to-one with the lifecycle in Overall Architecture Anatomy. Each week covers three types of content: concept pages for foundations, case pages for intuition, and practice pages for hands-on work.
1. Eight-Week Plan
| Week | Topic | Concept pages | Case pages | Practice pages | Verification criteria |
|---|---|---|---|---|---|
| W1 | Introduction + language modeling | What Is a Large Language Model, Brief History, Language Modeling, Tokenization and Vocabularies | — | — | Can explain "why predicting the next word teaches knowledge" |
| W2 | Architecture | Transformer Architecture, Context and Long Contexts | GPT Series, BERT and the Encoder Family | — | Can derive self-attention complexity O(n²) by hand; can compare encoder vs. decoder |
| W3 | Training and scale | Pretraining, Scaling Laws, MoE | MoE and Ultra-Large Models | Building a Large Model from Scratch | Train a small model locally with a visibly declining loss curve |
| W4 | Post-training and alignment | Fine-tuning, Alignment | ChatGPT and Conversational Models | Fine-tuning in Practice: Full LoRA Workflow | Complete one LoRA fine-tuning run and compare before/after results |
| W5 | Evaluation and safety | Evaluation and Benchmarks, Hallucination, Safety and Risks | — | Evaluation in Practice | Can explain what MMLU/GSM8K/HumanEval each measure; can design a custom evaluation set |
| W6 | Prompting and applications | Prompting, Inference Fundamentals | RAG, Agents with LLMs | Prompting in Practice, RAG in Practice | Independently build a retrieval-augmented QA system |
| W7 | Deployment and services | — | — | Deployment and Servicing, Framework and Tool Selection, Common Pitfalls | Can estimate a model's GPU memory usage; can explain what continuous batching solves |
| W8 | Papers and wrap-up | — | — | Reading Path, Paper Map, Classic Paper Deep Dives | Deep-read 2–3 core papers, can explain their methods and point out limitations |
2. Weekly Rhythm
60 minutes/day reading concept pages (close reading, take notes)
One evening per week reading case pages (reading only, to get a feel for "what this technology looks like")
One afternoon per week doing hands-on work (running code, tweaking parameters, observing changes)
30 minutes every Sunday for self-check (against the "verification criteria" column)3. Eight Reusable Outputs
The real payoff from systematic deep dive is not "finishing" but eight things you can showcase:
- Your own "what is an LLM" explanation (including a one-sentence definition + a three-layer breakdown);
- A hand-drawn Transformer diagram (can recite QKV and multi-head attention);
- A small language model trained locally (even if it only has a few million parameters);
- A complete LoRA fine-tuning log (including loss curves and before/after comparison);
- A custom-built evaluation set (10–20 golden-set examples, with evaluation criteria);
- A working RAG demo (documents → chunking → vector retrieval → generation);
- An inference/deployment selection note (framework comparison + memory estimation);
- 2–3 paper deep-read notes (following the template in Classic Paper Deep Dives).
Why this order?
W1→W7 follows the order of the model lifecycle: first data and objectives, then architecture, then training, alignment, evaluation, and finally deployment and applications. Learning backwards (playing with APIs first, filling in theory later) is certainly possible, but "wrong theory but still playable, mastered but can't piece theory back together" — the former costs a one-time frustration, the latter costs systemic understanding. Additionally, the paper reading in W8 requires the conceptual foundation of the previous seven weeks, so placing it at the end is most efficient.
4. Selection Logic Behind the Weekly Mix
Every column in the eight-week table is intentionally placed. The logic behind those choices deserves to be explained clearly, so you can adjust them yourself:
| Week | Why this mix |
|---|---|
| W1 | Establish "definition and goals" first (what is an LLM, why predicting the next word works), then dive into mechanisms. Otherwise everything hangs in the air. |
| W2 | Architecture is the carrier of all mechanisms, and comparing encoder/decoder (on the case page) grounds the architectural abstraction. |
| W3 | Understand the architecture, then see "how to train it"; after training, see "how to scale it" (scaling laws). The logic flows naturally. |
| W4 | Pretraining gives capability; post-training gives form. SFT/alignment and LoRA are the most commonly used techniques in this phase from an engineering standpoint. |
| W5 | Establish evaluation and safety before "start tweaking the model," to avoid "blindly tuning parameters and going to production unprotected." |
| W6 | Prompting is a zero-code entry point; RAG/Agent are the two most dominant system patterns today. |
| W7 | Deployment turns a model into a service; concepts like continuous batching and quantization only become truly meaningful here. |
| W8 | Papers give "first-hand sources," interviews give "output verification." Both are wrap-ups, not new openings. |
Adjustability principle: The order within each week can be swapped (reading cases before concepts is fine), but dependencies between weeks (W2 before W3) are not recommended to break. If you fall behind on a week, it's better to delay by a week than to skip — skipped chapters will repeatedly block you in all subsequent weeks.
IV. Route C: Desk Reference
1. How to Use It
Desk reference doesn't require "reading." You just need to know what problem each page solves. Start by spending 30 minutes going through Overall Architecture Anatomy (it's the site's master index), then flip to the corresponding page directly when a problem comes up:
| When you encounter at work | Go directly to |
|---|---|
| Model output quality is poor, want to improve | Prompting in Practice → Common Pitfalls |
| Need private knowledge QA | RAG in Practice |
| Need the model to output JSON in a specific format | Prompting |
| Model generates repeats/garbage/hallucinations | Hallucination → Inference Fundamentals |
| Want the model smaller and faster | Deployment and Servicing |
| Choosing frameworks/models | Framework and Tool Selection, Mainstream Model Profiles |
| Looking up terms/datasets | Glossary, Datasets and Benchmarks Archive |
| Following the frontier | Frontier Progress |
2. Advanced Forms of Desk Reference
After six months on the job, desk reference naturally evolves into "check yourself first, then check the handbook": when a problem arises, first restate the problem itself (often just restating it tells you where the answer lives), then decide whether to check the glossary for terminology, check the practice pages for workflows, or check the papers for theory. This habit is worth more than memorizing any single page.
V. Five Universal Tips
Regardless of which route you take, five tips apply to all:
- Hands-on beats reading. Reading 80% of a concept page is not as effective as running the small model from Building a Large Model from Scratch. Code forces you to confront the illusion of "I thought I understood."
- Main thread + side threads. Push only one theme per week on the main thread; side threads (glossary, resource lists, model profiles) are only browsed when needed and don't consume main-thread time.
- Verify input with output. At the end of each week, try explaining that week's topic to a friend (or a wall). Where you get stuck is what you need to fill in next week.
- Mark knowledge with expiration dates. The LLM field eliminates one generation of conclusions every 18 months; "most recent" on a page gets old quickly. When it comes to model specs and leaderboards, rely on dataAsOf annotations in pages like Mainstream Model Profiles, and keep revisiting Frontier Progress.
- Set verification criteria before you start reading. Each stage of the three routes has a "verification criteria" column. Translate goals into verifiable behaviors — "can derive attention complexity by hand" is far more testable than "finished the Transformer page."
How to use verification criteria
Verification criteria are not "self-test exams after reading." They are traffic signals that tell you when to stop: hit the target, move on; miss it, go back and reread — don't power through. LLM knowledge is a web — if you can't get through one page, it's often because a prerequisite page wasn't digested (struggling through Fine-tuning usually means Transformer wasn't fully absorbed). Allow yourself to backtrack; that's how progress actually happens.
VI. Site Map: Overview of Seven Modules
| Module | In one sentence | Typical pages | When to read |
|---|---|---|---|
| Introduction (guide) | Definitions, history, map, and concept boundaries | What Is a Large Language Model, Brief History | Must-read in week one |
| Core knowledge (concepts) | 14 pages breaking LLM into independently understandable knowledge bits | Transformer, Alignment | Backbone of Systematic Deep Dive |
| Case studies (case-studies) | Ground knowledge bits into specific models and systems | GPT Series, Llama and the Open-Source Ecosystem | After learning each concept, pair it with a case |
| Paper deep dives (papers) | Go back to first-hand papers, establish "the source of principles" | Paper Map, Classic Paper Deep Dives | W8 and beyond |
| Practice guide (practice) | "Follow along and it works" at the code and workflow level | Fine-tuning in Practice, RAG in Practice | Consult alongside hands-on work |
| Careers & JD (career) | Job landscape, JD breakdown, and interview question bank | Interview Question Bank | Exclusive to Career Sprint |
| Resources (resources) | Terms, datasets, models, and curated resource archives | Glossary, Mainstream Model Profiles | Look up as needed |
VII. Common Learning Misconceptions and Countermeasures
Shared pitfalls across the three routes, addressed proactively:
| Misconception | Typical behavior | Countermeasure |
|---|---|---|
| Reading concepts without touching code | Can recite LoRA's principle, but can't run a single line of fine-tuning code | Force some hands-on time every week, starting from Building from Scratch |
| Chasing new models without building foundations | Scrolling through release news daily, but can't explain attention mechanisms | Treat new models as "cases" and ground them in concept pages like Transformer |
| Greedily reading fast | Finish 10 pages in a week, but none of them are digested | Slow down according to verification criteria; "finishing" does not equal "learning" |
| Memorizing conclusions without tracing causes | Remembers "data/params ≈ 1:20" but can't say why | Ask "where does this number come from, how is it derived," and go back to Scaling Laws |
| Neglecting evaluation and safety | Model works → ship it. No eval set, no risk assessment | Fold Evaluation in Practice and Safety and Risks into required reading, not optional |
One-sentence summary
Career Sprint turns high-frequency topics into your own words in 4 weeks, Systematic Deep Dive walks from foundations to papers and interviews in 8 weeks, and Desk Reference makes the handbook ahandbook a go-to reference tool. The three routes share the same pages; the difference is only in order, depth, and verification method.
Further Reading
- What Is a Large Language Model — The single most important page in the book: three-layer definition, capability inventory, and empirical benchmarks
- Overall Architecture Anatomy — Site-wide map: which pages correspond to each of the seven stages in the model lifecycle
- Brief History — Seven decades from information theory to ChatGPT, and the "three paradigm shifts"
- Building a Large Model from Scratch — The hands-on core of Route B, Week 3
- Interview Question Bank — The verification tool for Route A, Week 4
- Paper Map — Starting point for Route C's Week 8 to dive into first-hand literature
References
- Andrej Karpathy · nanoGPT (GitHub repository) — Training GPT in a few hundred lines of code; the standard hands-on project for Systematic Deep Dive W3
- Hugging Face · LLM Course (free course) — An open-source course that pairs with this handbook, with plenty of runnable notebooks
- DeepLearning.AI · Short Courses — Andrew Ng's team of short courses (prompting, RAG, Agent, etc.), great for Route C to fill gaps
- OpenAI · Official Blog — First-hand release notes for GPT series, ChatGPT, o1, and other milestones
- EleutherAI · Language Model Evaluation Harness — The standard tool for evaluation in practice; pairs with Evaluation in Practice