Appearance
Deconstructing JD Knowledge Points
You pick up a JD from a company you like, and it says: "proficient in Python / C++," "familiar with vLLM or TensorRT-LLM," "understanding of CUDA and kernel optimization," "quantization experience a plus"... Every line reads clearly, yet every line is shrouded in fog: how good is "good enough"? And how will the interviewer test it?
What this page does is translate the "skill words" in JDs into a knowledge-point checklist you can review item by item — not a vague "learn CUDA well," but precisely "can write a GEMM kernel and explain bank conflicts," "can explain what problem PagedAttention solves and what data structure it uses." Combined with the five-level self-assessment scale and the L1→L3 catch-up route at the end, you can complete a "skills physical exam" of yourself in one afternoon and walk out with a learning roadmap of your own.
Get a JD ──① extract skill words──② map to exam points──③ self-assess──④ build catch-up list──⑤ retest weekly
│ │ │ │
proficient/familiar/ Sections 2–6 of Section 7 of Section 8 of
basic/bonus this page this page this pageThis page suits: candidates for inference engine engineer, LLM systems engineer, inference optimization engineer, ML Infra / MLOps, kernel engineer, and edge inference engineer roles — plus any learner who wants to turn "I think I know this" into "I can explain it."
1. What JD Skill Words Actually Test
A JD is not a syllabus — it is a company's minimum expectation of "what you can independently deliver." The hiring side's subtext is:
Writing "familiar with vLLM" doesn't require you to memorize the entire vLLM API; it requires you to independently judge, when a new model goes live, whether to enable PagedAttention, what to set the batch cap to, why, and how to troubleshoot when things go wrong.
So the first step in deconstructing a JD is mapping skill words to "the abilities an interviewer will actually test."
1. The "real exam points" behind high-frequency skill words
| JD wording | Literal meaning | What the interview actually tests (what you must prove in 5 minutes) |
|---|---|---|
| Proficient in Python / C++ | Can write both languages | Can explain the difference between Python's GIL and C++ multithreading; can write a leak-free RAII class in C++; can call a custom operator via a PyTorch C++ extension |
| Familiar with PyTorch | Can train with the framework | Can explain autograd / dispatch / eager vs compile; can explain the difference between torch.compile and torch.export; can sketch the PyTorch forward computation graph by hand |
| Familiar with vLLM or TensorRT-LLM | Has used an inference engine | Can explain PagedAttention's block table data structure; can explain the difference between continuous batching and in-flight batching; can locate the bottleneck when TTFT is high |
| Understanding of CUDA and kernel optimization | Has heard of CUDA | Can write a simple reduction kernel; can explain warp divergence and bank conflicts; can read an Nsight Compute report |
| Familiar with quantization | Has used the quantize API | Can explain the PTQ vs QAT paths; can explain the mechanistic differences among GPTQ / AWQ / SmoothQuant; can compute the memory savings of INT4 weight-only quantization |
| Familiar with distributed inference | Has used TP / PP | Can compute TP's communication overhead; can explain the difference between EP (expert parallelism) and DP; can state PP's bubble ratio |
| MLOps / platform experience | Has shipped to production | Can explain model version management, canary rollout, monitoring & alerting, and rollback plans; can compute cost and GPU utilization |
2. Why you can't read JDs literally
Three common misconceptions:
- Misconception 1: treating the JD as a full checklist. Of the 10 requirements in a JD, the company usually means only 3 as hard bars; the rest is an "ideal candidate portrait." Before applying, separate threshold items from bonus items and prioritize the thresholds.
- Misconception 2: reading "familiar" as "aware of." When an interviewer writes "familiar with vLLM," they test whether you can explain the principles, modify the source, and articulate trade-offs — not whether you've "heard of vLLM." When self-assessing, use "can explain it thoroughly" as the passing line for "familiar."
- Misconception 3: reviewing single points but not combinations. Interviews almost never test a single point in isolation: they hand you a scenario ("the P99 of your online Llama-70B service just spiked to 8 seconds — how do you investigate?") that simultaneously demands systems (check monitoring) + engine (check vLLM parameters) + kernels (check the attention kernel) + business (check the prompt distribution). So the three blocks below must be read together.
3. The full pipeline from JD to catch-up list
Get a JD
│
▼
① Extract skill words (proficient / familiar / basic / bonus)
│
▼
② Map to the three-layer knowledge structure (L1→L3, Sections 2–6)
│
▼
③ Self-assess each item on the five-level scale (Section 7)
│
▼
④ Generate your personal catch-up list (Section 8)
│
▼
⑤ Execute weekly → retest → update the listThe next five sections are the three blocks of this map (L1 / L2 / L3, each split into subsections).
2. L1 Fundamentals: Programming, Systems, Hardware Basics
JD words: "proficient in Python / C++," "solid engineering skills," "understanding of GPU architecture." L1 is the foundation of all inference & deployment roles — fail this layer and your resume is screened out at the first pass.
1. Production-grade Python
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| GIL and multithreading | "Can Python multithreading speed up CPU inference?" | Glossary |
| asyncio | "Why does vLLM's API server use asyncio?" | Model Serving and Orchestration |
| Type annotations / mypy | "How do you write type-safe interfaces for an inference engine?" | — |
| Memory management | "Why should LLM inference avoid frequent Python object creation?" | The GPU Memory Hierarchy and the Bandwidth Wall |
| C extensions / PyBind11 | "How do you call a C++ operator from Python?" | Kernel Fusion and Custom Kernels |
2. Systems-level C++
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| RAII / smart pointers | "How does an inference engine manage device-buffer lifetimes?" | vLLM |
| Templates and compile-time | "Have you looked at CUTLASS's template gymnastics?" | GPU Architecture and Optimization |
| Multithreading / mutexes | "How does Triton Server implement concurrent model instances?" | Triton Inference Server |
| Move semantics / rvalue refs | "Why does vLLM's Request object need move instead of copy?" | — |
| Build systems: CMake / Bazel | "How do you build a C++ project with CUDA?" | — |
3. Linux and performance analysis
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| Processes / threads / coroutines | "What does vLLM use for worker–driver communication?" | vLLM |
| strace / ltrace / perf | "Online inference is slow — how do you locate the syscall bottleneck?" | Inference Benchmarking in Practice |
| Nsight Systems | "Have you used nsys? Can you tell a trace story?" | GPU Architecture and Optimization |
| Nsight Compute | "Which metrics do you look at in an ncu report?" | Kernel Fusion and Custom Kernels |
| PyTorch Profiler | "How do you use torch.profiler to find forward bottlenecks?" | Tuning and Performance Optimization |
4. Networking and protocols
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| TCP / HTTP / gRPC | "Why does Triton Server use gRPC rather than HTTP?" | Triton Inference Server |
| RDMA / NCCL | "Does TP communication go over PCIe or NVLink — and how do you measure it?" | Distributed Inference (TP/PP) |
| WebSocket / SSE | "What protocol does LLM streaming output use?" | Model Serving and Orchestration |
5. GPU hardware basics
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| SIMT model | "How do GPU threads / blocks / grids map?" | Hardware Primer |
| Memory hierarchy | "How do shared memory / L1 / L2 / HBM differ in latency?" | The GPU Memory Hierarchy and the Bandwidth Wall |
| Tensor Core | "Tensor Core vs CUDA Core?" | GPU Architecture and Optimization |
| H100 vs A100 | "What are H100's key upgrades? How do FP8 / TMA affect inference?" | Hardware Primer |
| NVLink / NVSwitch | "Why is TP fast within an 8-GPU node but slow across nodes?" | Distributed Inference (TP/PP) |
The "false familiarity" trap at L1
"Can use Python" and "can explain the GIL" are two different things. L1 is where "can use it but can't explain it" false familiarity most often hides, because daily library calls are enough to get work done. Interviews strike at your weakest point: you think you know Python, then get asked about the relationship between the GIL and asyncio and freeze. Grade yourself by "can explain it thoroughly," not by "have used it."
3. L2 Topic Layer: Inference Engines, Quantization, Scheduling
JD words: "familiar with vLLM or TensorRT-LLM," "quantization experience," "familiar with KV cache and batching." L2 is the must-test main line of interviews, and the biggest difference between inference & deployment roles and traditional ML roles.
1. Inference engine mechanics
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| vLLM overall architecture | "Draw the three-layer relationship of LLMEngine / Scheduler / Executor" | vLLM |
| PagedAttention | "What data structure is the block table? Why does it solve fragmentation?" | Batching and Request Scheduling |
| Continuous batching | "Where does the throughput gain over static batching come from?" | Batching and Request Scheduling |
| TensorRT-LLM In-Flight Batching | "How does it differ from vLLM's continuous batching?" | TensorRT-LLM |
| Chunked Prefill | "Why split long prompts into chunks?" | vLLM |
| Speculative decoding | "Differences between Medusa and EAGLE? When should you not use it?" | Speculative Decoding and Medusa/EAGLE |
| Prefix caching | "How do repeated prompts save KV cache recomputation?" | vLLM |
2. KV cache and memory management
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| KV cache size calculation | "Llama-70B, batch=8, seq=2048 — how big is the KV cache?" | The GPU Memory Hierarchy and the Bandwidth Wall |
| KV cache quantization | "How much memory does INT8 KV cache save? What's the accuracy cost?" | Model Quantization Fundamentals |
| PagedAttention data structures | "How is block_size chosen? How are physical blocks allocated?" | Batching and Request Scheduling |
| Multi-Query / Grouped-Query Attention | "How do MQA / GQA save KV cache?" | vLLM |
| Ring / Paged Attention across nodes | "How do PP and KV cache coordinate?" | Distributed Inference (TP/PP) |
3. Quantization and compression
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| PTQ vs QAT | "What do you do when PTQ loses accuracy? When do you go QAT?" | Model Quantization Fundamentals |
| GPTQ / AWQ / SmoothQuant | "Mechanistic differences? Which suits which model?" | Model Quantization Fundamentals |
| Weight-only vs weight+activation | "Why does LLM inference mainly push weight-only?" | Weight-Only Quantization and Mixed Precision |
| INT4 / INT8 / FP8 | "On H100, which is faster, FP8 or INT8 — and why?" | Model Quantization Fundamentals |
| KV cache quantization | "Where does the accuracy loss of INT8 KV cache quantization come from?" | Model Quantization Fundamentals |
| Pruning and distillation | "Why do LLM inference roles rarely talk about pruning?" | Pruning and Sparsification, Knowledge Distillation |
4. Batching and scheduling
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| Static batching | "Why is static batching low-throughput for LLMs?" | Batching and Request Scheduling |
| Continuous batching | "How does the scheduler decide when to swap requests in and out?" | Batching and Request Scheduling |
| In-Flight Batching | "TensorRT-LLM's in-flight vs vLLM's continuous — differences?" | TensorRT-LLM |
| Prefill / decode separation | "Why separate them? What's the communication overhead after separation?" | Distributed Inference (TP/PP) |
| Priority scheduling | "How do you schedule VIP user requests without hurting regular users?" | Model Serving and Orchestration |
5. Serving and orchestration
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| Triton Server model repository | "How do you configure a model repository with multi-version canary?" | Triton Inference Server |
| FastAPI / asyncio serving | "How does vLLM's OpenAI-compatible server implement streaming?" | Model Serving and Orchestration |
| K8s deployment and autoscaling | "How do you autoscale vLLM replicas based on QPS?" | Model Serving and Orchestration |
| Canary and rollback | "How do you run a 5% canary for a model rollout?" | Model Serving and Orchestration |
| Monitoring metrics | "How do you instrument TTFT / TPOT / queue length?" | Inference Benchmarking in Practice |
The review strategy for L2
Every L2 exam point maps to a trio of "open-source project source + paper + measurement." The most effective review: pick one engine (vLLM recommended), read its source closely + run one benchmark + change one parameter and observe the effect. Finish these three steps and you can answer every L2 question with "what I did" instead of "what I read." See Inference Benchmarking in Practice and Tuning and Performance Optimization.
4. L3 Advanced Layer: Kernels, Graphs, Distributed, Edge
JD words: "understanding of kernel optimization," "familiar with distributed inference," "edge deployment experience a plus." L3 is the "bonus / differentiator" layer of interviews — reaching this layer means you've passed the L1/L2 screen and they're using it to separate "can use" from "can modify."
1. Kernel development
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| FlashAttention principles | "What did FlashAttention-1/2/3 each do?" | Kernel Fusion and Custom Kernels |
| CUDA kernel writing | "Hand-write a reduction kernel" | GPU Architecture and Optimization |
| Triton DSL | "Triton's advantages over CUDA? When should you use it?" | Kernel Fusion and Custom Kernels |
| Kernel fusion strategies | "How do you compute the launch-overhead savings of fusing Linear + ReLU?" | Kernel Fusion and Custom Kernels |
| Tensor Core / WGMMA | "On Hopper, how do wgmma instructions differ from mma?" | GPU Architecture and Optimization |
| TMA (Tensor Memory Accelerator) | "What problem does TMA solve? Why is H100 inference fast?" | Hardware Primer |
| Bank conflict / coalesced access | "How do you avoid shared-memory bank conflicts?" | GPU Architecture and Optimization |
| CUDA Graph | "Why does LLM inference capture CUDA Graphs?" | Computation Graph Optimization |
2. Graph optimization and compilation
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| Kernel fusion | "What does fusing Conv + BN + ReLU save?" | Computation Graph Optimization |
| Constant folding / dead-code elimination | "What does torch.compile's constant folding do?" | Computation Graph Optimization |
| torch.compile / Inductor | "How much does torch.compile beat eager in inference, and why?" | Computation Graph Optimization |
| AOT vs JIT | "TensorRT is AOT and PyTorch eager is JIT — what does that mean?" | Computation Graph Optimization |
| ONNX / TorchScript / Export | "Why is torch.export better than TorchScript?" | Computation Graph Optimization |
| MLIR / XLA / TVM | "Where do compiler backends sit inside an inference engine?" | Computation Graph Optimization |
3. Distributed inference
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| Tensor Parallel | "TP shards weights — how do you compute the communication overhead?" | Distributed Inference (TP/PP) |
| Pipeline Parallel | "How do you compute PP's bubble ratio?" | Distributed Inference (TP/PP) |
| Expert Parallel (MoE) | "EP vs DP? Why does MoE inference use EP?" | Distributed Inference (TP/PP) |
| Disaggregated Prefill / Decode | "Why separate them? How does the KV cache cross nodes?" | Distributed Inference (TP/PP) |
| Multi-node inference | "Why is cross-node TP slow? NVSwitch vs Ethernet?" | Distributed Inference (TP/PP) |
| Ring Attention | "What problem does Ring Attention solve for long context?" | Distributed Inference (TP/PP) |
4. Edge inference
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| llama.cpp / ggml | "Why can llama.cpp run 70B on an M1?" | llama.cpp and GGUF |
| MLC-LLM / Apache TVM | "How does MLC's compilation path differ from llama.cpp?" | Mobile Deployment |
| ARM NEON / SVE | "What drives ARM CPU inference optimization?" | Mobile Deployment |
| Apple Metal / ANE | "On iPhone, does an LLM run on Metal or the ANE?" | Mobile Deployment |
| WebGPU / Vulkan | "Where's the bottleneck when running an LLM in the browser?" | Mobile Deployment |
| INT4 group-wise quantization | "Why does the edge mainly use INT4 group-wise?" | Model Quantization Fundamentals |
5. Performance metrics and benchmarks
| Exam point | How it's tested in interviews | Page on this site |
|---|---|---|
| TTFT / TPOT definitions | "The physical meaning of TTFT and TPOT?" | Latency, Throughput, and Concurrency |
| Compute-bound vs memory-bound | "How do you tell which regime an LLM inference is in?" | The Roofline Model and Compute Analysis |
| Roofline model | "How do you use the Roofline model to compute the theoretical ceiling?" | The Roofline Model and Compute Analysis |
| HBM bandwidth | "Why is Llama-70B inference memory-bound?" | The GPU Memory Hierarchy and the Bandwidth Wall |
| Benchmark methodology | "How do you write a fair LLM inference benchmark?" | Inference Benchmarking in Practice |
| Warmup and statistics | "Why warm up? How do you take P50/P99?" | Inference Benchmarking in Practice |
5. Master Cross-Reference: JD Keywords → Site Pages
The four sections above are consolidated into a printable cross-reference table — your entry point for "read a JD → find a page":
| JD keyword | Real exam point | Page on this site | L1 / L2 / L3 |
|---|---|---|---|
| Python / C++ | GIL, RAII, templates, PyBind11 | Glossary, vLLM | L1 |
| Linux / performance analysis | nsys / ncu / perf / PyTorch Profiler | Inference Benchmarking in Practice | L1 |
| Networking / gRPC | TCP / HTTP / gRPC / RDMA / NCCL | Triton Inference Server | L1 |
| GPU architecture | SIMT / Tensor Core / HBM / NVLink | Hardware Primer | L1 |
| vLLM | Architecture / PagedAttention / continuous batching | vLLM | L2 |
| TensorRT-LLM | In-Flight Batching / Plugin / FP8 | TensorRT-LLM | L2 |
| Quantization | GPTQ / AWQ / SmoothQuant / INT4 / INT8 / FP8 | Model Quantization Fundamentals, Weight-Only Quantization and Mixed Precision | L2 |
| KV cache | Size calculation / PagedAttention / MQA / GQA | Batching and Request Scheduling, The GPU Memory Hierarchy and the Bandwidth Wall | L2 |
| Batching | Continuous / In-Flight / Prefill-Decode separation | Batching and Request Scheduling | L2 |
| Serving | Triton / FastAPI / K8s / canary | Triton Inference Server, Model Serving and Orchestration | L2 |
| CUDA | Kernel / SIMT / bank conflict / Tensor Core | GPU Architecture and Optimization | L3 |
| Triton DSL | Write attention / quantization kernels | Kernel Fusion and Custom Kernels | L3 |
| FlashAttention | FA1/2/3 / tiling / recompute | Kernel Fusion and Custom Kernels | L3 |
| CUDA Graph | Capture & replay / launch overhead | Computation Graph Optimization | L3 |
| Graph optimization | Kernel fusion / constant folding / torch.compile | Computation Graph Optimization | L3 |
| Distributed inference | TP / PP / EP / disaggregated | Distributed Inference (TP/PP) | L3 |
| Speculative decoding | Medusa / EAGLE / acceptance rate | Speculative Decoding and Medusa/EAGLE | L3 |
| Edge inference | llama.cpp / MLC / Metal / WebGPU | llama.cpp and GGUF, Mobile Deployment | L3 |
| Performance metrics | TTFT / TPOT / Roofline / HBM bandwidth | Latency, Throughput, and Concurrency, The Roofline Model and Compute Analysis | L3 |
How to use this table
Look up the "Page on this site" column row by row for each JD keyword, then mark "can do / can't / half-can" next to each page name — that gives you the entry point of your catch-up list. Don't try to fill the whole table — only fill the parts covered by "hard requirements + bonus items where you're one or two levels short."
6. Suggested Learning Priorities
Priorities are ranked by "return on investment + role fit":
Priority 1 L1 Python / C++ / Linux / performance analysis
The foundation and the hard bar for all inference roles
Investment: ongoing — 5–10 hours/week to maintain
Priority 2 L2 Inference engines (vLLM / TensorRT-LLM) + KV cache + batching
The main interview line for LLM inference roles
Investment: 2–4 weeks, close source reading + measurement
Priority 3 L2 Quantization + serving
Standard equipment for all deployment roles
Investment: 1–2 weeks, run comparison experiments
Priority 4 L3 Kernel development (CUDA / Triton) + graph optimization
Differentiators for engine / optimization / kernel roles
Investment: 3–6 weeks ongoing, starting from reduction
Priority 5 L3 Distributed inference + edge inference + speculative decoding
Role-specific differentiators; fill per target role
Investment: 1–2 weeks, papers + demosThe hard constraint on time allocation
L1 is always priority 1 — every role's hard bar includes Python + Linux, and it can't be skipped. L2 is priority 2 because LLM inference roles account for 70%+ of inference & deployment openings in 2026. L3 is customized to the target role: engine roles → kernels + graph optimization; systems roles → serving + K8s; edge roles → llama.cpp + quantization. See the "role type → learning path" mapping in Learning Paths: Three Routes.
7. Self-Assessment: The Five-Level Scale (Printable)
The exam points from the previous sections are consolidated into a printable self-assessment table. Scoring criteria:
| Level | Meaning | Criterion (grade yourself in the mirror) |
|---|---|---|
| 1 Never heard of it | Completely unfamiliar | Haven't even seen the term |
| 2 Heard of it | Vague impression | Seen or heard it, but can't give a definition |
| 3 Aware | Know the concept | Can give a definition and one example, but not the principles or trade-offs |
| 4 Familiar | Can use it hands-on | Can implement / tune / troubleshoot independently; knows pros, cons, and applicable scenarios |
| 5 Can teach it | Can teach others | Can explain principles + derivation + trade-offs + counterexamples in 5 minutes, and survive follow-ups |
How to print and check
Print this section, check one box per row for your current, honest level. Honesty is the premise — the goal is to expose every row below "can teach it," not to make the table look good. Check only one level per row; when in doubt, check the lower one.
Block A: L1 Fundamentals
| Exam point | Never heard | Heard | Aware | Familiar | Can teach |
|---|---|---|---|---|---|
| Python GIL / asyncio | □ | □ | □ | □ | □ |
| C++ RAII / smart pointers / templates | □ | □ | □ | □ | □ |
| Linux performance analysis (perf / nsys) | □ | □ | □ | □ | □ |
| Reading Nsight Compute reports | □ | □ | □ | □ | □ |
| PyTorch Profiler | □ | □ | □ | □ | □ |
| TCP vs gRPC vs HTTP | □ | □ | □ | □ | □ |
| RDMA / NCCL communication | □ | □ | □ | □ | □ |
| GPU SIMT / memory hierarchy | □ | □ | □ | □ | □ |
| Tensor Core / TMA | □ | □ | □ | □ | □ |
| H100 vs A100 key differences | □ | □ | □ | □ | □ |
Block B: L2 Engines and Scheduling
| Exam point | Never heard | Heard | Aware | Familiar | Can teach |
|---|---|---|---|---|---|
| vLLM architecture / Scheduler | □ | □ | □ | □ | □ |
| PagedAttention block table | □ | □ | □ | □ | □ |
| Continuous batching | □ | □ | □ | □ | □ |
| In-Flight Batching (TRT-LLM) | □ | □ | □ | □ | □ |
| Chunked Prefill | □ | □ | □ | □ | □ |
| Prefix caching | □ | □ | □ | □ | □ |
| KV cache size calculation | □ | □ | □ | □ | □ |
| KV cache quantization | □ | □ | □ | □ | □ |
| MQA / GQA | □ | □ | □ | □ | □ |
| Triton Server model repository | □ | □ | □ | □ | □ |
| K8s deployment and autoscaling | □ | □ | □ | □ | □ |
| Canary rollout and rollback | □ | □ | □ | □ | □ |
| TTFT / TPOT monitoring | □ | □ | □ | □ | □ |
Block C: L2 Quantization and Compression
| Exam point | Never heard | Heard | Aware | Familiar | Can teach |
|---|---|---|---|---|---|
| PTQ vs QAT | □ | □ | □ | □ | □ |
| GPTQ / AWQ / SmoothQuant | □ | □ | □ | □ | □ |
| Weight-only vs weight+activation | □ | □ | □ | □ | □ |
| INT4 / INT8 / FP8 selection | □ | □ | □ | □ | □ |
| Kernel fusion / graph optimization concepts | □ | □ | □ | □ | □ |
| Limits of pruning / distillation for LLMs | □ | □ | □ | □ | □ |
Block D: L3 Kernels and Graphs
| Exam point | Never heard | Heard | Aware | Familiar | Can teach |
|---|---|---|---|---|---|
| FlashAttention-1/2/3 principles | □ | □ | □ | □ | □ |
| Hand-writing a CUDA reduction kernel | □ | □ | □ | □ | □ |
| Triton DSL attention | □ | □ | □ | □ | □ |
| Kernel fusion strategies | □ | □ | □ | □ | □ |
| Tensor Core / WGMMA / TMA | □ | □ | □ | □ | □ |
| Bank conflict / coalesced access | □ | □ | □ | □ | □ |
| CUDA Graph capture | □ | □ | □ | □ | □ |
| torch.compile / Inductor | □ | □ | □ | □ | □ |
| AOT vs JIT compilation | □ | □ | □ | □ | □ |
Block E: L3 Distributed and Edge
| Exam point | Never heard | Heard | Aware | Familiar | Can teach |
|---|---|---|---|---|---|
| Tensor Parallel communication overhead | □ | □ | □ | □ | □ |
| Pipeline Parallel bubble | □ | □ | □ | □ | □ |
| Expert Parallel (MoE) | □ | □ | □ | □ | □ |
| Disaggregated Prefill/Decode | □ | □ | □ | □ | □ |
| Ring Attention | □ | □ | □ | □ | □ |
| llama.cpp / ggml | □ | □ | □ | □ | □ |
| MLC-LLM / TVM | □ | □ | □ | □ | □ |
| ARM NEON / Apple Metal | □ | □ | □ | □ | □ |
| WebGPU / Vulkan | □ | □ | □ | □ | □ |
| Speculative decoding (Medusa / EAGLE) | □ | □ | □ | □ | □ |
| Roofline model | □ | □ | □ | □ | □ |
| HBM bandwidth and memory-bound | □ | □ | □ | □ | □ |
8. How to Turn the Self-Assessment into a Personal Catch-Up List
After finishing the self-assessment, turn it into an executable weekly plan in four steps.
Step 1: Group by level
Classify each exam point into three buckets:
- Class A (levels 1–2, completely unfamiliar): highest priority — first solve "what is it." Goal: reach level 3 within a week.
- Class B (level 3, aware): what's missing is principles and trade-offs. Goal: reach level 4 within two weeks — the standard is "can implement independently + state applicable scenarios."
- Class C (level 4, familiar): what's missing is "teach-it" ability. Goal: turn each point into a 5-minute spoken talk. This maps directly to interview performance.
Level-5 points need no more time unless you're targeting a teaching- or theory-heavy role.
Step 2: Rank by block and total the debt
Aggregate the three classes by block to get three "debt tables." Suggested catch-up order:
Priority 1 Block A L1 fundamentals — the foundation under all other blocks
Priority 2 Block B L2 engines & scheduling — the main interview line for LLM
inference roles; highest ROI
Priority 3 Block C L2 quantization — required by nearly all deployment roles
Priority 4 Block D L3 kernels & graphs — if the target role is engine /
optimization / kernel
Priority 5 Block E L3 distributed & edge — fill per target roleStep 3: Apply the "catch-up list" template
| Week | Focus block | Class A tasks (to level 3) | Class B tasks (to level 4) | Class C tasks (write talks) | Deliverable |
|---|---|---|---|---|---|
| Week 1 | L1 fundamentals | 3 unfamiliar points | 2 points: add principles | 1 talk (e.g., GIL) | Notes + talk |
| Week 2 | L2 engines + KV cache | 2 unfamiliar points | 3 points: add principles | 2 talks (e.g., PagedAttention, continuous batching) | Source notes + talks |
| Week 3 | L2 quantization + serving | 1 unfamiliar point | 2 points: add principles | 2 talks (e.g., GPTQ vs AWQ, TTFT/TPOT) | Comparison experiment + talks |
| Week 4 | L3 kernels / distributed | 1 unfamiliar point | 2 points: add principles | 2 talks (e.g., FlashAttention, TP) | Hand-written kernel + talks |
| Week 5+ | Loop until A/B are cleared | Adjust dynamically | Adjust dynamically | Push toward "combination" questions | Mock-interview recordings |
Step 4: Three execution principles
- Principle 1: it doesn't count until you can say it. After reviewing each point, record a 5-minute self-talk on your phone and listen back — wherever you stumble is what you didn't actually understand.
- Principle 2: source + measurement beats documentation. For engine-type points, you must run vLLM yourself once, change one parameter once, and read one Nsight report — "read the docs" and "ran the measurement" are two different things in an interview.
- Principle 3: retest weekly. On Sunday night, rescore with the self-assessment table, cross out the points that leveled up, and flag the ones that didn't move — then ask why. It's usually not lack of effort but the wrong review method (e.g., only watching, never practicing).
One more thing
The goal of a catch-up list is not "coloring the whole table in" — it's getting every threshold item in your target role's JD to level 4 or above. Blocks your role doesn't need (e.g., edge inference for an engine role) can sit at level 4; put your time on the blade's edge.
9. Further Reading
On this site:
- Career module home · JD List · Skills Benchmarking · Interview Question Bank
- Deep dives for L1: What Is Inference Acceleration?, Inference vs. Training vs. Fine-Tuning, GPU Architecture and Optimization, Hardware Primer
- Deep dives for L2: vLLM, TensorRT-LLM, Model Quantization Fundamentals, Weight-Only Quantization and Mixed Precision, Batching and Request Scheduling, Model Serving and Orchestration, Triton Inference Server
- Deep dives for L3: Kernel Fusion and Custom Kernels, Computation Graph Optimization, Distributed Inference (TP/PP), Speculative Decoding and Medusa/EAGLE, llama.cpp and GGUF, Mobile Deployment, Latency, Throughput, and Concurrency, The Roofline Model and Compute Analysis, The GPU Memory Hierarchy and the Bandwidth Wall
- Practice pages: Inference Benchmarking in Practice, Tuning and Performance Optimization, Inference Engine Comparison, Common Pitfalls and Anti-Patterns
- Always at hand: the Glossary is the pocketbook for filling gaps
References (real resources):
- vLLM official docs and source — the best textbook for L2 engine exam points
- TensorRT-LLM official docs — the other pole, for learning in contrast with vLLM
- NVIDIA CUDA C++ Programming Guide — the official textbook for L3 kernel development
- OpenAI Triton Language Tutorials — getting started with the Triton DSL
- NVIDIA Nsight Systems Documentation — official docs for the performance-analysis toolchain
- Dao-Beck-Puzzar-Kolter. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (NeurIPS 2022) — the original FlashAttention paper
- Kwon et al. Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP 2023) — the original vLLM/PagedAttention paper
- Frantar et al. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (ICLR 2023) — the original GPTQ paper
- Lin et al. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (MLSys 2024) — the original AWQ paper
- Cai et al. Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads (ICML 2024) — the original Medusa paper
- Li et al. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty (ICML 2024) — the original EAGLE paper