page" for practicing deployment engineers.">
Skip to content

Learning Paths: Three Routes

At a glance The 50+ pages on this site are not a library to read cover to cover — they are roadmaps you combine as needed. There is a job-sprint track for anyone interviewing for an inference/deployment role within three months, an eight-week systematic track for those who want a solid foundation, and a desk-reference index that maps "problem -> page" for practicing deployment engineers.

Learning Paths: Three Routes ​

This site has more than 50 pages. If you start at the first page and read straight through, you will probably lose focus around section three of The Roofline Model and Compute Analysis — not because the content is bad, but because reading cover to cover is not the right way to learn inference acceleration. What you need is not a library; it is a roadmap.

In fact, everything on this site serves only four purposes: building concepts (the guides), explaining mechanisms (the core knowledge), providing evidence (the case studies and papers), and changing your behavior (the practice and career sections). People in different situations should consume these in completely different orders and at different depths. This page gives you three ready-made routes — just pick the one that fits.

text
                  What is your situation?
                             |
 ┌──────────────────────────┼───────────────────────────┐
         ▼                   ▼                   ▼
   Interview in       Time to spare,          Already
    1-3 months        want to switch         deploying

┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
| 1 Job Sprint    | | 2 Systematic    | | 3 Desk          |
|   ~2 weeks      | |   ~8 weeks      | | Reference       |
├─────────────────┤ ├─────────────────┤ ├─────────────────┤
| career core     | | Guide -> Core   | | Hit a problem   |
| + hot concepts  | | -> Cases ->     | | -> check the    |
| + hands-on      | | Practice ->     | |    index        |
|    demo         | | weekly checks   | | -> back to work |
├─────────────────┤ ├─────────────────┤ ├─────────────────┤
| Output:         | | Output:         | | Output:         |
| a polished      | | can explain     | | the problem at  |
| resume + 10     | | the whole stack | | hand gets       |
| Q&A answers     | | + a portfolio   | | solved          |
└─────────────────┘ └─────────────────┘ └─────────────────┘

The three routes are not mutually exclusive

The real trajectory of many practicing deployment engineers looks like this: they use the site as a desk reference for months, decide to fill in the foundations and run the eight-week track, then re-run the second half of the job-sprint track before an interview. The routes are navigation, not rails.

Route 1: Job-Sprint Track (~2 weeks) ​

Who it is for: You have an interview for an inference acceleration / model deployment / LLM serving role in 1–3 months. You already know Python and basic ML (you can run transformers), and you have no large blocks of time — only evenings and weekends.

Core idea: An inference/deployment interview really tests two things — whether you can explain the full engineering pipeline "from model weights to production serving" (from quantization to batching to KV cache), and whether your resume and communication survive screening. So this track starts from the end and works backward from job descriptions and interview questions. It does not aim for completeness; it aims to cover every high-frequency exam point and to let you speak to the "why" and the "trade-off" behind each one.

Main line: the four career-module pages ​

OrderPageWhat you take away
1JD List: Open Inference & Deployment Roles at Major CompaniesWhat target roles actually require — copy down the skill words that keep recurring (vLLM / TensorRT / quantization / KV cache appear in almost every one)
2Deconstructing JD Knowledge PointsMap the skill words from JDs to pages on this site and generate your personal gap-closure list
3Skills Benchmarking: What to Highlight on Your ResumeRewrite your project experience to its standard, foregrounding concrete latency / throughput / cost numbers
4Interview Question BankUse it for self-testing last, not on day one

Side dishes: the highest-frequency concepts + one hands-on demo ​

Judging from the distribution of JDs and interview write-ups at major companies at home and abroad, the following four pages cover the vast majority of "engineering-side" exam points. Read them in this order:

  1. Latency, Throughput, and Concurrency — almost guaranteed in every deployment interview; you should be able to rattle off TTFT / TPOT / E2E latency, tokens/s, and QPS, and explain why prefill and decode differ in arithmetic intensity by 100x.
  2. Model Quantization Fundamentals — the trade-offs among PTQ/QAT, symmetric/asymmetric, per-channel/per-tensor, and INT8/INT4/FP8.
  3. Model Serving and Orchestration — the three deployment shapes (online / batch / streaming), replica strategy, canary release, and SLO design.
  4. GPU Architecture and Optimization — SM/warp/HBM/SRAM, kernel launch, and the roofline model; at minimum, be able to explain "why LLM decode is memory-bound."

If you can only add one more page, add Batching and Request Scheduling — continuous batching is the core mechanism of vLLM/TGI/SGLang and is asked in almost every interview.

A resume built on reading alone is not convincing. Spend a weekend working through Deploy an Inference Service from Scratch: the repo's v1 minimal forward is only a few dozen lines; v2 adds KV cache + batching; v3 plugs into vLLM/TensorRT-LLM. Run it layer by layer and tweak it, and you will own a hands-on project you can talk about for five straight minutes in an interview — and it happens to cover interviewers' favorite question, "what have you tuned yourself, and what pitfalls have you hit?"

The two-week rhythm ​

Week 1 (understand + resume):

  • Monday to Wednesday: JD List -> Deconstructing JD Knowledge Points, producing your gap-closure list.
  • Thursday to Friday: read the four concept pages in the order above. After each one, close the page and record yourself on your phone explaining the core mechanism ("why is continuous batching faster than static batching", "why must INT4 quantization be aware of activation outliers"). Re-read whatever you cannot explain smoothly.
  • Weekend: get build-your-own's v1->v3 running; finish the first draft of your resume to the resume-analysis page's standard.

Week 2 (communicate + self-test):

  • Monday to Wednesday: work through the Interview Question Bank — answer each question yourself before checking, and mark the ones you missed in red to revisit the corresponding concept pages.
  • Thursday to Friday: run mock interviews with a friend or by voice-Q&A, focusing on "why" questions (Why is the KV cache paged? Why does speculative decoding not change the output distribution? Why is FP8 more popular than INT8 on H100?).
  • Weekend: finalize the resume and turn the demo project into a presentable repo with benchmark numbers attached.

Acceptance criteria after two weeks

You should have three things in hand: a resume rewritten to the deployment-role standard (every project carries latency/throughput/cost numbers), a runnable, explainable inference-service repo (with KV cache + batching + at least one form of quantization), and answers to 10 interview questions that each name the trade-offs and costs. Missing any one means you cut corners in that step.

Route 2: Systematic Track (~8 weeks) ​

Who it is for: You have relatively ample time (8–10 hours a week) and want to complete a transition or build a solid foundation within 1–2 quarters. You value a system over quick wins.

Core idea: Knowledge in inference acceleration has a dependency graph — you cannot understand why quantization works without understanding memory access, you cannot grasp the essence of vLLM without understanding batching, and you cannot write your own scheduling strategy without having dissected vLLM. This route is arranged in dependency order, with a checkpoint every week; do not move to the next week until you pass the current one.

WeekContentCheckpoint
Week 1The five guides: What Is Inference Acceleration? -> Inference vs. Training vs. Fine-Tuning -> A Brief History -> Anatomy of the Overall Architecture -> Learning Paths: Three RoutesExplain to a newcomer "why the same model delivers 100x the throughput on vLLM as on transformers," and draw the five-layer architecture of an inference system
Week 2The three performance pages: Latency, Throughput, and Concurrency -> The GPU Memory Hierarchy and the Bandwidth Wall -> The Roofline Model and Compute AnalysisCompute from scratch the single-request decode token/s ceiling of Llama-70B FP16 on an A100 (roughly bandwidth / model size ≈ 28 tokens/s), and explain the arithmetic-intensity gap between prefill and decode
Week 3The four compression pages: Model Quantization Fundamentals -> Weight-Only Quantization and Mixed Precision -> Pruning and Sparsification -> Knowledge DistillationExplain the difference between GPTQ and AWQ, why INT4 weight quantization often degrades accuracy on LLMs, and when to choose distillation over quantization
Week 4The two operator/graph pages: Kernel Fusion and Custom Kernels -> Computation Graph Optimization; optional case readings TensorRT and GPU Inference, ONNX Runtime: Cross-PlatformSketch FlashAttention's tiling scheme from memory and explain why it saves memory access; list at least five graph-level fusion patterns
Week 5The three system/hardware pages: GPU Architecture and Optimization -> Batching and Request Scheduling -> Model Serving and Orchestration; hands-on: Deploy an Inference Service from Scratch v1->v3Explain how requests join and leave continuous batching and the paging idea of PagedAttention; get the demo running and explain what each step does
Week 6Engine case readings: vLLM and PagedAttention + TensorRT-LLM are required; SGLang / TGI / llama.cpp and GGUF / Triton Inference Server are optionalCompare at least three engines across the four dimensions of "latency / throughput / usability / hardware support," and say which workload should run on which
Week 7Practice wrap-up: Inference Benchmarking in Practice -> Tuning and Performance Optimization -> Deployment Design Principles -> Common Pitfalls and Anti-Patterns; optional Speculative Decoding and Medusa/EAGLE, Distributed Inference (TP/PP), Mobile DeploymentYour project runs on someone else's machine, and you can state at what batch/seq_len it will fail and why
Week 8Papers + career-material wrap-up: read at least 3 of Classic Papers in Depth (PagedAttention, FlashAttention, and one of GPTQ/AWQ); optional Frontier Advances; then go through the career moduleExplain the motivation and contribution of at least two classic papers; bring your resume and 10 question answers up to Route 1's acceptance bar

The hard requirement for Week 5

The only non-compressible week in the eight is the hands-on week. After reading concept pages you will feel an illusion of "I already get it" — and that illusion shatters the first time you face a real OOM, KV cache fragmentation, or batching jitter. A transition does not count as done until you have been personally taught a lesson by a CUDA OOM, by padding waste, and by quantization accuracy loss.

If time is really tight, you can cut the Week 6 engine readings and the Week 8 papers in half, but do not cut the four blocks — guides, performance, quantization, and hands-on — they are the load-bearing walls of this route.

Route 3: Desk-Reference Track (practicing deployment engineers) ​

Who it is for: You already do inference/deployment development, and your problems arrive in units of "stuck today." You have neither the need nor the patience to read cover to cover.

The way to use this route: when you hit a problem, look up the entry page in the table below, read it, solve the problem, and leave. Do not follow every vine. Save the vine-following for the weekend.

Problem you are hittingGo to
Online latency spikes, TPOT climbingLatency, Throughput, and Concurrency, Batching and Request Scheduling
GPU utilization stuck low (<30%)The Roofline Model and Compute Analysis, GPU Architecture and Optimization, Batching and Request Scheduling
GPU memory OOM, KV cache exhaustedThe GPU Memory Hierarchy and the Bandwidth Wall, vLLM and PagedAttention
Want to quantize but unsure between INT8 and INT4Model Quantization Fundamentals, Weight-Only Quantization and Mixed Precision
Severe accuracy drop after quantizationWeight-Only Quantization and Mixed Precision (the AWQ/SmoothQuant part), Common Pitfalls and Anti-Patterns
Migrating from transformers to vLLMvLLM and PagedAttention, Deploy an Inference Service from Scratch
Migrating from vLLM to TensorRT-LLMTensorRT-LLM, Inference Engine Comparison
Need to deploy on CPU / at the edgeOpenVINO and CPU Inference, llama.cpp and GGUF, Mobile Deployment
Want multi-model routing, A/B testingModel Serving and Orchestration, Triton Inference Server
A single GPU cannot fit 70B; need TP/PPDistributed Inference (TP/PP)
Want to add speculative decoding for speedSpeculative Decoding and Medusa/EAGLE
Attention dominates; want to optimize itKernel Fusion and Custom Kernels (the FlashAttention part)
Graph optimization not taking effect, TensorRT errorsComputation Graph Optimization, Common Pitfalls and Anti-Patterns
Poor streaming experience, slow first tokenLatency, Throughput, and Concurrency (the TTFT part), Batching and Request Scheduling
Want load testing / a benchmark reportInference Benchmarking in Practice, Benchmark Data & Tool Profiles
Boss/client asks "why not a newer model"Deployment Design Principles, Inference Engine Comparison
Interviewing or being interviewed, need to probe enginesInterview Question Bank
Stuck on terminology (KV cache, continuous batching, speculative decoding)Glossary
Want a GPU's compute / bandwidth / memory specsHardware Primer
Looking for open-source engines or datasetsCurated Resources

Bookmark this page

The value of the desk-reference track is its reuse rate. Bookmark this page (not any single concept page) — it is the switchboard of the entire site.

Two Caveats ​

The paths are not rigid rules. The week counts and ordering of all three routes are estimated for a typical background: strong in algorithms? fast-forward through the compression weeks. Strong in engineering? fast-forward through model serving. Strong in hardware? fast-forward through GPU optimization. The standard is always whether you can pass that week's checkpoint, not which page the calendar has turned to.

Mind the shelf life of the content. Inference acceleration is a field with an extremely short half-life — engine versions, benchmark scores, and hardware generations all expire. Some pages in the career module and the case-studies chapter carry a dataAsOf tag (the month the data was current) in their frontmatter, shown at the top of the page. Before citing a specific speedup, memory figure, or engine version from them, check that date: anything older than a quarter should be treated as "directional reference" rather than "current fact," and re-verified against official benchmarks. Conceptual content (what KV cache is, why decode is memory-bound) decays far more slowly and can be relied on for the long term.

Further Reading ​