Appearance
Learning Paths: Three Routes
This site has more than 50 pages. If you start at the first page and read straight through, you will probably lose focus around section three of The Roofline Model and Compute Analysis — not because the content is bad, but because reading cover to cover is not the right way to learn inference acceleration. What you need is not a library; it is a roadmap.
In fact, everything on this site serves only four purposes: building concepts (the guides), explaining mechanisms (the core knowledge), providing evidence (the case studies and papers), and changing your behavior (the practice and career sections). People in different situations should consume these in completely different orders and at different depths. This page gives you three ready-made routes — just pick the one that fits.
text
What is your situation?
|
┌──────────────────────────┼───────────────────────────┐
▼ ▼ ▼
Interview in Time to spare, Already
1-3 months want to switch deploying
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
| 1 Job Sprint | | 2 Systematic | | 3 Desk |
| ~2 weeks | | ~8 weeks | | Reference |
├─────────────────┤ ├─────────────────┤ ├─────────────────┤
| career core | | Guide -> Core | | Hit a problem |
| + hot concepts | | -> Cases -> | | -> check the |
| + hands-on | | Practice -> | | index |
| demo | | weekly checks | | -> back to work |
├─────────────────┤ ├─────────────────┤ ├─────────────────┤
| Output: | | Output: | | Output: |
| a polished | | can explain | | the problem at |
| resume + 10 | | the whole stack | | hand gets |
| Q&A answers | | + a portfolio | | solved |
└─────────────────┘ └─────────────────┘ └─────────────────┘The three routes are not mutually exclusive
The real trajectory of many practicing deployment engineers looks like this: they use the site as a desk reference for months, decide to fill in the foundations and run the eight-week track, then re-run the second half of the job-sprint track before an interview. The routes are navigation, not rails.
Route 1: Job-Sprint Track (~2 weeks)
Who it is for: You have an interview for an inference acceleration / model deployment / LLM serving role in 1–3 months. You already know Python and basic ML (you can run transformers), and you have no large blocks of time — only evenings and weekends.
Core idea: An inference/deployment interview really tests two things — whether you can explain the full engineering pipeline "from model weights to production serving" (from quantization to batching to KV cache), and whether your resume and communication survive screening. So this track starts from the end and works backward from job descriptions and interview questions. It does not aim for completeness; it aims to cover every high-frequency exam point and to let you speak to the "why" and the "trade-off" behind each one.
Main line: the four career-module pages
| Order | Page | What you take away |
|---|---|---|
| 1 | JD List: Open Inference & Deployment Roles at Major Companies | What target roles actually require — copy down the skill words that keep recurring (vLLM / TensorRT / quantization / KV cache appear in almost every one) |
| 2 | Deconstructing JD Knowledge Points | Map the skill words from JDs to pages on this site and generate your personal gap-closure list |
| 3 | Skills Benchmarking: What to Highlight on Your Resume | Rewrite your project experience to its standard, foregrounding concrete latency / throughput / cost numbers |
| 4 | Interview Question Bank | Use it for self-testing last, not on day one |
Side dishes: the highest-frequency concepts + one hands-on demo
Judging from the distribution of JDs and interview write-ups at major companies at home and abroad, the following four pages cover the vast majority of "engineering-side" exam points. Read them in this order:
- Latency, Throughput, and Concurrency — almost guaranteed in every deployment interview; you should be able to rattle off TTFT / TPOT / E2E latency, tokens/s, and QPS, and explain why prefill and decode differ in arithmetic intensity by 100x.
- Model Quantization Fundamentals — the trade-offs among PTQ/QAT, symmetric/asymmetric, per-channel/per-tensor, and INT8/INT4/FP8.
- Model Serving and Orchestration — the three deployment shapes (online / batch / streaming), replica strategy, canary release, and SLO design.
- GPU Architecture and Optimization — SM/warp/HBM/SRAM, kernel launch, and the roofline model; at minimum, be able to explain "why LLM decode is memory-bound."
If you can only add one more page, add Batching and Request Scheduling — continuous batching is the core mechanism of vLLM/TGI/SGLang and is asked in almost every interview.
A resume built on reading alone is not convincing. Spend a weekend working through Deploy an Inference Service from Scratch: the repo's v1 minimal forward is only a few dozen lines; v2 adds KV cache + batching; v3 plugs into vLLM/TensorRT-LLM. Run it layer by layer and tweak it, and you will own a hands-on project you can talk about for five straight minutes in an interview — and it happens to cover interviewers' favorite question, "what have you tuned yourself, and what pitfalls have you hit?"
The two-week rhythm
Week 1 (understand + resume):
- Monday to Wednesday: JD List -> Deconstructing JD Knowledge Points, producing your gap-closure list.
- Thursday to Friday: read the four concept pages in the order above. After each one, close the page and record yourself on your phone explaining the core mechanism ("why is continuous batching faster than static batching", "why must INT4 quantization be aware of activation outliers"). Re-read whatever you cannot explain smoothly.
- Weekend: get build-your-own's v1->v3 running; finish the first draft of your resume to the resume-analysis page's standard.
Week 2 (communicate + self-test):
- Monday to Wednesday: work through the Interview Question Bank — answer each question yourself before checking, and mark the ones you missed in red to revisit the corresponding concept pages.
- Thursday to Friday: run mock interviews with a friend or by voice-Q&A, focusing on "why" questions (Why is the KV cache paged? Why does speculative decoding not change the output distribution? Why is FP8 more popular than INT8 on H100?).
- Weekend: finalize the resume and turn the demo project into a presentable repo with benchmark numbers attached.
Acceptance criteria after two weeks
You should have three things in hand: a resume rewritten to the deployment-role standard (every project carries latency/throughput/cost numbers), a runnable, explainable inference-service repo (with KV cache + batching + at least one form of quantization), and answers to 10 interview questions that each name the trade-offs and costs. Missing any one means you cut corners in that step.
Route 2: Systematic Track (~8 weeks)
Who it is for: You have relatively ample time (8–10 hours a week) and want to complete a transition or build a solid foundation within 1–2 quarters. You value a system over quick wins.
Core idea: Knowledge in inference acceleration has a dependency graph — you cannot understand why quantization works without understanding memory access, you cannot grasp the essence of vLLM without understanding batching, and you cannot write your own scheduling strategy without having dissected vLLM. This route is arranged in dependency order, with a checkpoint every week; do not move to the next week until you pass the current one.
| Week | Content | Checkpoint |
|---|---|---|
| Week 1 | The five guides: What Is Inference Acceleration? -> Inference vs. Training vs. Fine-Tuning -> A Brief History -> Anatomy of the Overall Architecture -> Learning Paths: Three Routes | Explain to a newcomer "why the same model delivers 100x the throughput on vLLM as on transformers," and draw the five-layer architecture of an inference system |
| Week 2 | The three performance pages: Latency, Throughput, and Concurrency -> The GPU Memory Hierarchy and the Bandwidth Wall -> The Roofline Model and Compute Analysis | Compute from scratch the single-request decode token/s ceiling of Llama-70B FP16 on an A100 (roughly bandwidth / model size ≈ 28 tokens/s), and explain the arithmetic-intensity gap between prefill and decode |
| Week 3 | The four compression pages: Model Quantization Fundamentals -> Weight-Only Quantization and Mixed Precision -> Pruning and Sparsification -> Knowledge Distillation | Explain the difference between GPTQ and AWQ, why INT4 weight quantization often degrades accuracy on LLMs, and when to choose distillation over quantization |
| Week 4 | The two operator/graph pages: Kernel Fusion and Custom Kernels -> Computation Graph Optimization; optional case readings TensorRT and GPU Inference, ONNX Runtime: Cross-Platform | Sketch FlashAttention's tiling scheme from memory and explain why it saves memory access; list at least five graph-level fusion patterns |
| Week 5 | The three system/hardware pages: GPU Architecture and Optimization -> Batching and Request Scheduling -> Model Serving and Orchestration; hands-on: Deploy an Inference Service from Scratch v1->v3 | Explain how requests join and leave continuous batching and the paging idea of PagedAttention; get the demo running and explain what each step does |
| Week 6 | Engine case readings: vLLM and PagedAttention + TensorRT-LLM are required; SGLang / TGI / llama.cpp and GGUF / Triton Inference Server are optional | Compare at least three engines across the four dimensions of "latency / throughput / usability / hardware support," and say which workload should run on which |
| Week 7 | Practice wrap-up: Inference Benchmarking in Practice -> Tuning and Performance Optimization -> Deployment Design Principles -> Common Pitfalls and Anti-Patterns; optional Speculative Decoding and Medusa/EAGLE, Distributed Inference (TP/PP), Mobile Deployment | Your project runs on someone else's machine, and you can state at what batch/seq_len it will fail and why |
| Week 8 | Papers + career-material wrap-up: read at least 3 of Classic Papers in Depth (PagedAttention, FlashAttention, and one of GPTQ/AWQ); optional Frontier Advances; then go through the career module | Explain the motivation and contribution of at least two classic papers; bring your resume and 10 question answers up to Route 1's acceptance bar |
The hard requirement for Week 5
The only non-compressible week in the eight is the hands-on week. After reading concept pages you will feel an illusion of "I already get it" — and that illusion shatters the first time you face a real OOM, KV cache fragmentation, or batching jitter. A transition does not count as done until you have been personally taught a lesson by a CUDA OOM, by padding waste, and by quantization accuracy loss.
If time is really tight, you can cut the Week 6 engine readings and the Week 8 papers in half, but do not cut the four blocks — guides, performance, quantization, and hands-on — they are the load-bearing walls of this route.
Route 3: Desk-Reference Track (practicing deployment engineers)
Who it is for: You already do inference/deployment development, and your problems arrive in units of "stuck today." You have neither the need nor the patience to read cover to cover.
The way to use this route: when you hit a problem, look up the entry page in the table below, read it, solve the problem, and leave. Do not follow every vine. Save the vine-following for the weekend.
Bookmark this page
The value of the desk-reference track is its reuse rate. Bookmark this page (not any single concept page) — it is the switchboard of the entire site.
Two Caveats
The paths are not rigid rules. The week counts and ordering of all three routes are estimated for a typical background: strong in algorithms? fast-forward through the compression weeks. Strong in engineering? fast-forward through model serving. Strong in hardware? fast-forward through GPU optimization. The standard is always whether you can pass that week's checkpoint, not which page the calendar has turned to.
Mind the shelf life of the content. Inference acceleration is a field with an extremely short half-life — engine versions, benchmark scores, and hardware generations all expire. Some pages in the career module and the case-studies chapter carry a dataAsOf tag (the month the data was current) in their frontmatter, shown at the top of the page. Before citing a specific speedup, memory figure, or engine version from them, check that date: anything older than a quarter should be treated as "directional reference" rather than "current fact," and re-verified against official benchmarks. Conceptual content (what KV cache is, why decode is memory-bound) decays far more slowly and can be relied on for the long term.
Further Reading
- What Is Inference Acceleration? — the common starting point of all three routes
- Anatomy of the Overall Architecture — the content map of the whole site; this page and that one index each other
- Module Guide and Job Landscape — the entry to the career module that anchors Route 1's main line
- Classic Papers in Depth — the primary literature for Week 8 of Route 2
- Deployment Design Principles — a page worth coming back to after finishing any route