Skip to content

Learning Paths: Three Routes

At a glance This handbook offers three learning paths — Career Sprint, Systematic Deep Dive, and Desk Reference — each tailored to different audiences and time budgets, with weekly schedules, page checklists, and verification criteria. It serves as the master index map for the entire site.

Learning Paths: Three Routes ​

In one sentence: Learning Paths is the site's map — it doesn't answer "what is an LLM" (that's What Is a Large Language Model's job), nor does it dive into any specific technology. Instead, it answers three more fundamental questions: In what order should I read? How deep should I go at each stage? How do I know I've actually learned it?

This site has seven modules (Introduction / Core Knowledge / Case Studies / Paper Deep Dives / Practice Guide / Careers & JD / Resources) with roughly 40 pages in total. No one can realistically read everything from start to finish, nor should they. These three routes help you cut through the noise: the same mountain, three different paths up.

I. Overview of the Three Routes ​

DimensionRoute A: Career SprintRoute B: Systematic Deep DiveRoute C: Desk Reference
Target audiencePreparing for interviews, short on timeWant to build a systematic knowledge frameworkAlready working with LLMs
Core objectiveCan answer interview questions, can talk about projectsClear on concepts, can reproduce implementationsQuickly find answers when problems arise
Suggested duration4–6 weeks × 3–4 hrs/day8 weeks × 10–15 hrs/weekNo fixed schedule, look up as needed
Core pagesConcept pages + interview question bankConcept pages + case pages + practice pagesPractice pages + resource pages
Reading depthNeed conclusions, can retell themNeed derivations, can break things downNeed answers, can follow along
Verification criteriaGet ~80% of interview questions rightCan explain any core concept from scratchNo longer need to ask others for help
text
The three routes are not three levels — they are three ways to invest:

  Career Sprint:       Concept ──→ Interview Qs ──→ Can answer
  Systematic Deep Dive: Concept ──→ Case ──→ Practice ──→ Can build
  Desk Reference:       Problem ──→ Page ──→ Answer ──→ Can look up

How to choose?

Be clear about "what will I deliver in a month." Interviewing next month → pick A. Running a fine-tuning/RAG project next month → pick B. Shipping a service next month → start with C, and fill in B's concept pages when you hit a bottleneck. A and B are not mutually exclusive: most serious job seekers actually walk a simplified version of B plus A's interview intensification. This article presents B as the main thread with a complete 8-week plan, A as its trimmed version, and C as its reverse lookup index.

II. Route A: Career Sprint (4–6 Weeks) ​

1. Goals and Principles ​

Career Sprint has the best return on investment, but only if you focus on the big picture and skip the details: read only the high-frequency interview topics, skip the obscure stuff. High-frequency interview coverage for LLM positions spans four blocks — architecture, training, inference, and applications — see the interview question bank. The sprint principle is "be able to explain any topic in your own words for two minutes," not "have read the most pages."

2. Four-Week Schedule ​

WeekTopicRequired PagesDaily Output
W1Concepts and foundationsWhat Is a Large Language Model, Language Modeling, Tokenization and VocabulariesDefine "what is an LLM" in one paragraph; hand-draw the next-token prediction flow
W2Architecture and trainingTransformer Architecture, Pretraining, Scaling LawsDerive the QKV computation flow; explain why the scaling factor is √d_k
W3Post-training and alignmentFine-tuning, Alignment, Inference FundamentalsRecite LoRA's W=W₀+BA; draw the three steps of RLHF
W4Applications and sprintRAG, Agents with LLMs, Interview Question BankGo through 10 interview questions daily; write project experience as "background → challenge → solution → result"

3. Verification Criteria ​

  • Without notes, you can draw the Transformer structure and label Q, K, V, residuals, and LayerNorm;
  • You can explain the SFT → reward model → PPO three steps of RLHF, and say why DPO can skip the reward model;
  • You can explain what problems RAG and fine-tuning each solve, and when to choose which (see RAG vs. long context trade-offs);
  • Mock interviews: pick 10 random questions from the interview question bank and answer 8 correctly on the spot.

The most common pitfall of Career Sprint

Spending time on "memorizing more jargon" rather than "clearly explaining concepts you already know." Interviewers probe depth through follow-up questions, and follow-ups rely on mechanism understanding, not jargon volume. It's better to read one page on Transformers three times than to half-swallow ten pages.

4. Quick Reference: High-Frequency Interview Topics ​

The topic map corresponding to the four-week route (a compressed version of the interview question bank, for self-check):

BlockHigh-frequency topicsRelated pages
ArchitectureQKV computation and scaling factor √d_k, motivation for multi-head attention, positional encoding extrapolation, causal maskingTransformer
TrainingNext-token objective and cross-entropy, data mixing, scaling laws, LoRA's W=W₀+BALanguage Modeling, Scaling Laws, Fine-tuning
AlignmentRLHF three steps, reward model training, PPO and KL penalty, why DPO can skip the reward modelAlignment
InferenceAutoregression and KV Cache, temperature and top-p, quantization, continuous batching, memory estimationInference Fundamentals, Deployment
ApplicationsRAG three stages, Agent loops, prompt injection and securityRAG, Agent

Usage: pick one row daily and "explain it in your own words for 2 minutes." If you can't, go back to the corresponding page. After four weeks, this table becomes your interview outline.

III. Route B: Systematic Deep Dive (8 Weeks) ​

This is the main thread of this handbook. It is organized in the order of "foundation → architecture → training → alignment → evaluation → applications → deployment → wrap-up," corresponding one-to-one with the lifecycle in Overall Architecture Anatomy. Each week covers three types of content: concept pages for foundations, case pages for intuition, and practice pages for hands-on work.

1. Eight-Week Plan ​

WeekTopicConcept pagesCase pagesPractice pagesVerification criteria
W1Introduction + language modelingWhat Is a Large Language Model, Brief History, Language Modeling, Tokenization and Vocabularies——Can explain "why predicting the next word teaches knowledge"
W2ArchitectureTransformer Architecture, Context and Long ContextsGPT Series, BERT and the Encoder Family—Can derive self-attention complexity O(n²) by hand; can compare encoder vs. decoder
W3Training and scalePretraining, Scaling Laws, MoEMoE and Ultra-Large ModelsBuilding a Large Model from ScratchTrain a small model locally with a visibly declining loss curve
W4Post-training and alignmentFine-tuning, AlignmentChatGPT and Conversational ModelsFine-tuning in Practice: Full LoRA WorkflowComplete one LoRA fine-tuning run and compare before/after results
W5Evaluation and safetyEvaluation and Benchmarks, Hallucination, Safety and Risks—Evaluation in PracticeCan explain what MMLU/GSM8K/HumanEval each measure; can design a custom evaluation set
W6Prompting and applicationsPrompting, Inference FundamentalsRAG, Agents with LLMsPrompting in Practice, RAG in PracticeIndependently build a retrieval-augmented QA system
W7Deployment and services——Deployment and Servicing, Framework and Tool Selection, Common PitfallsCan estimate a model's GPU memory usage; can explain what continuous batching solves
W8Papers and wrap-up——Reading Path, Paper Map, Classic Paper Deep DivesDeep-read 2–3 core papers, can explain their methods and point out limitations

2. Weekly Rhythm ​

60 minutes/day reading concept pages (close reading, take notes)
One evening per week reading case pages (reading only, to get a feel for "what this technology looks like")
One afternoon per week doing hands-on work (running code, tweaking parameters, observing changes)
30 minutes every Sunday for self-check (against the "verification criteria" column)

3. Eight Reusable Outputs ​

The real payoff from systematic deep dive is not "finishing" but eight things you can showcase:

  1. Your own "what is an LLM" explanation (including a one-sentence definition + a three-layer breakdown);
  2. A hand-drawn Transformer diagram (can recite QKV and multi-head attention);
  3. A small language model trained locally (even if it only has a few million parameters);
  4. A complete LoRA fine-tuning log (including loss curves and before/after comparison);
  5. A custom-built evaluation set (10–20 golden-set examples, with evaluation criteria);
  6. A working RAG demo (documents → chunking → vector retrieval → generation);
  7. An inference/deployment selection note (framework comparison + memory estimation);
  8. 2–3 paper deep-read notes (following the template in Classic Paper Deep Dives).

Why this order?

W1→W7 follows the order of the model lifecycle: first data and objectives, then architecture, then training, alignment, evaluation, and finally deployment and applications. Learning backwards (playing with APIs first, filling in theory later) is certainly possible, but "wrong theory but still playable, mastered but can't piece theory back together" — the former costs a one-time frustration, the latter costs systemic understanding. Additionally, the paper reading in W8 requires the conceptual foundation of the previous seven weeks, so placing it at the end is most efficient.

4. Selection Logic Behind the Weekly Mix ​

Every column in the eight-week table is intentionally placed. The logic behind those choices deserves to be explained clearly, so you can adjust them yourself:

WeekWhy this mix
W1Establish "definition and goals" first (what is an LLM, why predicting the next word works), then dive into mechanisms. Otherwise everything hangs in the air.
W2Architecture is the carrier of all mechanisms, and comparing encoder/decoder (on the case page) grounds the architectural abstraction.
W3Understand the architecture, then see "how to train it"; after training, see "how to scale it" (scaling laws). The logic flows naturally.
W4Pretraining gives capability; post-training gives form. SFT/alignment and LoRA are the most commonly used techniques in this phase from an engineering standpoint.
W5Establish evaluation and safety before "start tweaking the model," to avoid "blindly tuning parameters and going to production unprotected."
W6Prompting is a zero-code entry point; RAG/Agent are the two most dominant system patterns today.
W7Deployment turns a model into a service; concepts like continuous batching and quantization only become truly meaningful here.
W8Papers give "first-hand sources," interviews give "output verification." Both are wrap-ups, not new openings.

Adjustability principle: The order within each week can be swapped (reading cases before concepts is fine), but dependencies between weeks (W2 before W3) are not recommended to break. If you fall behind on a week, it's better to delay by a week than to skip — skipped chapters will repeatedly block you in all subsequent weeks.

IV. Route C: Desk Reference ​

1. How to Use It ​

Desk reference doesn't require "reading." You just need to know what problem each page solves. Start by spending 30 minutes going through Overall Architecture Anatomy (it's the site's master index), then flip to the corresponding page directly when a problem comes up:

When you encounter at workGo directly to
Model output quality is poor, want to improvePrompting in Practice → Common Pitfalls
Need private knowledge QARAG in Practice
Need the model to output JSON in a specific formatPrompting
Model generates repeats/garbage/hallucinationsHallucination → Inference Fundamentals
Want the model smaller and fasterDeployment and Servicing
Choosing frameworks/modelsFramework and Tool Selection, Mainstream Model Profiles
Looking up terms/datasetsGlossary, Datasets and Benchmarks Archive
Following the frontierFrontier Progress

2. Advanced Forms of Desk Reference ​

After six months on the job, desk reference naturally evolves into "check yourself first, then check the handbook": when a problem arises, first restate the problem itself (often just restating it tells you where the answer lives), then decide whether to check the glossary for terminology, check the practice pages for workflows, or check the papers for theory. This habit is worth more than memorizing any single page.

V. Five Universal Tips ​

Regardless of which route you take, five tips apply to all:

  1. Hands-on beats reading. Reading 80% of a concept page is not as effective as running the small model from Building a Large Model from Scratch. Code forces you to confront the illusion of "I thought I understood."
  2. Main thread + side threads. Push only one theme per week on the main thread; side threads (glossary, resource lists, model profiles) are only browsed when needed and don't consume main-thread time.
  3. Verify input with output. At the end of each week, try explaining that week's topic to a friend (or a wall). Where you get stuck is what you need to fill in next week.
  4. Mark knowledge with expiration dates. The LLM field eliminates one generation of conclusions every 18 months; "most recent" on a page gets old quickly. When it comes to model specs and leaderboards, rely on dataAsOf annotations in pages like Mainstream Model Profiles, and keep revisiting Frontier Progress.
  5. Set verification criteria before you start reading. Each stage of the three routes has a "verification criteria" column. Translate goals into verifiable behaviors — "can derive attention complexity by hand" is far more testable than "finished the Transformer page."

How to use verification criteria

Verification criteria are not "self-test exams after reading." They are traffic signals that tell you when to stop: hit the target, move on; miss it, go back and reread — don't power through. LLM knowledge is a web — if you can't get through one page, it's often because a prerequisite page wasn't digested (struggling through Fine-tuning usually means Transformer wasn't fully absorbed). Allow yourself to backtrack; that's how progress actually happens.

VI. Site Map: Overview of Seven Modules ​

ModuleIn one sentenceTypical pagesWhen to read
Introduction (guide)Definitions, history, map, and concept boundariesWhat Is a Large Language Model, Brief HistoryMust-read in week one
Core knowledge (concepts)14 pages breaking LLM into independently understandable knowledge bitsTransformer, AlignmentBackbone of Systematic Deep Dive
Case studies (case-studies)Ground knowledge bits into specific models and systemsGPT Series, Llama and the Open-Source EcosystemAfter learning each concept, pair it with a case
Paper deep dives (papers)Go back to first-hand papers, establish "the source of principles"Paper Map, Classic Paper Deep DivesW8 and beyond
Practice guide (practice)"Follow along and it works" at the code and workflow levelFine-tuning in Practice, RAG in PracticeConsult alongside hands-on work
Careers & JD (career)Job landscape, JD breakdown, and interview question bankInterview Question BankExclusive to Career Sprint
Resources (resources)Terms, datasets, models, and curated resource archivesGlossary, Mainstream Model ProfilesLook up as needed

VII. Common Learning Misconceptions and Countermeasures ​

Shared pitfalls across the three routes, addressed proactively:

MisconceptionTypical behaviorCountermeasure
Reading concepts without touching codeCan recite LoRA's principle, but can't run a single line of fine-tuning codeForce some hands-on time every week, starting from Building from Scratch
Chasing new models without building foundationsScrolling through release news daily, but can't explain attention mechanismsTreat new models as "cases" and ground them in concept pages like Transformer
Greedily reading fastFinish 10 pages in a week, but none of them are digestedSlow down according to verification criteria; "finishing" does not equal "learning"
Memorizing conclusions without tracing causesRemembers "data/params ≈ 1:20" but can't say whyAsk "where does this number come from, how is it derived," and go back to Scaling Laws
Neglecting evaluation and safetyModel works → ship it. No eval set, no risk assessmentFold Evaluation in Practice and Safety and Risks into required reading, not optional

One-sentence summary

Career Sprint turns high-frequency topics into your own words in 4 weeks, Systematic Deep Dive walks from foundations to papers and interviews in 8 weeks, and Desk Reference makes the handbook ahandbook a go-to reference tool. The three routes share the same pages; the difference is only in order, depth, and verification method.

Further Reading ​

References ​