Skip to content

Deconstructing JD Knowledge Points

At a glance Turn skill phrases from inference & deployment JDs — "proficient in vLLM," "familiar with CUDA," "quantization experience" — into a knowledge-point checklist you can review item by item. A three-layer competency structure (L1 fundamentals / L2 topics / L3 advanced) plus a five-level self-assessment scale, from "never heard of it" to "can teach it," generates your personal catch-up roadmap.

This page contains time-sensitive content. Data is current as of 2026-08; engine versions, benchmark rankings, and product features may have changed since — verify against primary sources before citing.

Deconstructing JD Knowledge Points ​

You pick up a JD from a company you like, and it says: "proficient in Python / C++," "familiar with vLLM or TensorRT-LLM," "understanding of CUDA and kernel optimization," "quantization experience a plus"... Every line reads clearly, yet every line is shrouded in fog: how good is "good enough"? And how will the interviewer test it?

What this page does is translate the "skill words" in JDs into a knowledge-point checklist you can review item by item — not a vague "learn CUDA well," but precisely "can write a GEMM kernel and explain bank conflicts," "can explain what problem PagedAttention solves and what data structure it uses." Combined with the five-level self-assessment scale and the L1→L3 catch-up route at the end, you can complete a "skills physical exam" of yourself in one afternoon and walk out with a learning roadmap of your own.

Get a JD ──① extract skill words──② map to exam points──③ self-assess──④ build catch-up list──⑤ retest weekly
               │                    │                  │              │
         proficient/familiar/   Sections 2–6 of    Section 7 of    Section 8 of
         basic/bonus            this page          this page       this page

This page suits: candidates for inference engine engineer, LLM systems engineer, inference optimization engineer, ML Infra / MLOps, kernel engineer, and edge inference engineer roles — plus any learner who wants to turn "I think I know this" into "I can explain it."

1. What JD Skill Words Actually Test ​

A JD is not a syllabus — it is a company's minimum expectation of "what you can independently deliver." The hiring side's subtext is:

Writing "familiar with vLLM" doesn't require you to memorize the entire vLLM API; it requires you to independently judge, when a new model goes live, whether to enable PagedAttention, what to set the batch cap to, why, and how to troubleshoot when things go wrong.

So the first step in deconstructing a JD is mapping skill words to "the abilities an interviewer will actually test."

1. The "real exam points" behind high-frequency skill words ​

JD wordingLiteral meaningWhat the interview actually tests (what you must prove in 5 minutes)
Proficient in Python / C++Can write both languagesCan explain the difference between Python's GIL and C++ multithreading; can write a leak-free RAII class in C++; can call a custom operator via a PyTorch C++ extension
Familiar with PyTorchCan train with the frameworkCan explain autograd / dispatch / eager vs compile; can explain the difference between torch.compile and torch.export; can sketch the PyTorch forward computation graph by hand
Familiar with vLLM or TensorRT-LLMHas used an inference engineCan explain PagedAttention's block table data structure; can explain the difference between continuous batching and in-flight batching; can locate the bottleneck when TTFT is high
Understanding of CUDA and kernel optimizationHas heard of CUDACan write a simple reduction kernel; can explain warp divergence and bank conflicts; can read an Nsight Compute report
Familiar with quantizationHas used the quantize APICan explain the PTQ vs QAT paths; can explain the mechanistic differences among GPTQ / AWQ / SmoothQuant; can compute the memory savings of INT4 weight-only quantization
Familiar with distributed inferenceHas used TP / PPCan compute TP's communication overhead; can explain the difference between EP (expert parallelism) and DP; can state PP's bubble ratio
MLOps / platform experienceHas shipped to productionCan explain model version management, canary rollout, monitoring & alerting, and rollback plans; can compute cost and GPU utilization

2. Why you can't read JDs literally ​

Three common misconceptions:

  • Misconception 1: treating the JD as a full checklist. Of the 10 requirements in a JD, the company usually means only 3 as hard bars; the rest is an "ideal candidate portrait." Before applying, separate threshold items from bonus items and prioritize the thresholds.
  • Misconception 2: reading "familiar" as "aware of." When an interviewer writes "familiar with vLLM," they test whether you can explain the principles, modify the source, and articulate trade-offs — not whether you've "heard of vLLM." When self-assessing, use "can explain it thoroughly" as the passing line for "familiar."
  • Misconception 3: reviewing single points but not combinations. Interviews almost never test a single point in isolation: they hand you a scenario ("the P99 of your online Llama-70B service just spiked to 8 seconds — how do you investigate?") that simultaneously demands systems (check monitoring) + engine (check vLLM parameters) + kernels (check the attention kernel) + business (check the prompt distribution). So the three blocks below must be read together.

3. The full pipeline from JD to catch-up list ​

  Get a JD
    │
    ▼
  ① Extract skill words (proficient / familiar / basic / bonus)
    │
    ▼
  ② Map to the three-layer knowledge structure (L1→L3, Sections 2–6)
    │
    ▼
  ③ Self-assess each item on the five-level scale (Section 7)
    │
    ▼
  ④ Generate your personal catch-up list (Section 8)
    │
    ▼
  ⑤ Execute weekly → retest → update the list

The next five sections are the three blocks of this map (L1 / L2 / L3, each split into subsections).

2. L1 Fundamentals: Programming, Systems, Hardware Basics ​

JD words: "proficient in Python / C++," "solid engineering skills," "understanding of GPU architecture." L1 is the foundation of all inference & deployment roles — fail this layer and your resume is screened out at the first pass.

1. Production-grade Python ​

Exam pointHow it's tested in interviewsPage on this site
GIL and multithreading"Can Python multithreading speed up CPU inference?"Glossary
asyncio"Why does vLLM's API server use asyncio?"Model Serving and Orchestration
Type annotations / mypy"How do you write type-safe interfaces for an inference engine?"—
Memory management"Why should LLM inference avoid frequent Python object creation?"The GPU Memory Hierarchy and the Bandwidth Wall
C extensions / PyBind11"How do you call a C++ operator from Python?"Kernel Fusion and Custom Kernels

2. Systems-level C++ ​

Exam pointHow it's tested in interviewsPage on this site
RAII / smart pointers"How does an inference engine manage device-buffer lifetimes?"vLLM
Templates and compile-time"Have you looked at CUTLASS's template gymnastics?"GPU Architecture and Optimization
Multithreading / mutexes"How does Triton Server implement concurrent model instances?"Triton Inference Server
Move semantics / rvalue refs"Why does vLLM's Request object need move instead of copy?"—
Build systems: CMake / Bazel"How do you build a C++ project with CUDA?"—

3. Linux and performance analysis ​

Exam pointHow it's tested in interviewsPage on this site
Processes / threads / coroutines"What does vLLM use for worker–driver communication?"vLLM
strace / ltrace / perf"Online inference is slow — how do you locate the syscall bottleneck?"Inference Benchmarking in Practice
Nsight Systems"Have you used nsys? Can you tell a trace story?"GPU Architecture and Optimization
Nsight Compute"Which metrics do you look at in an ncu report?"Kernel Fusion and Custom Kernels
PyTorch Profiler"How do you use torch.profiler to find forward bottlenecks?"Tuning and Performance Optimization

4. Networking and protocols ​

Exam pointHow it's tested in interviewsPage on this site
TCP / HTTP / gRPC"Why does Triton Server use gRPC rather than HTTP?"Triton Inference Server
RDMA / NCCL"Does TP communication go over PCIe or NVLink — and how do you measure it?"Distributed Inference (TP/PP)
WebSocket / SSE"What protocol does LLM streaming output use?"Model Serving and Orchestration

5. GPU hardware basics ​

Exam pointHow it's tested in interviewsPage on this site
SIMT model"How do GPU threads / blocks / grids map?"Hardware Primer
Memory hierarchy"How do shared memory / L1 / L2 / HBM differ in latency?"The GPU Memory Hierarchy and the Bandwidth Wall
Tensor Core"Tensor Core vs CUDA Core?"GPU Architecture and Optimization
H100 vs A100"What are H100's key upgrades? How do FP8 / TMA affect inference?"Hardware Primer
NVLink / NVSwitch"Why is TP fast within an 8-GPU node but slow across nodes?"Distributed Inference (TP/PP)

The "false familiarity" trap at L1

"Can use Python" and "can explain the GIL" are two different things. L1 is where "can use it but can't explain it" false familiarity most often hides, because daily library calls are enough to get work done. Interviews strike at your weakest point: you think you know Python, then get asked about the relationship between the GIL and asyncio and freeze. Grade yourself by "can explain it thoroughly," not by "have used it."

3. L2 Topic Layer: Inference Engines, Quantization, Scheduling ​

JD words: "familiar with vLLM or TensorRT-LLM," "quantization experience," "familiar with KV cache and batching." L2 is the must-test main line of interviews, and the biggest difference between inference & deployment roles and traditional ML roles.

1. Inference engine mechanics ​

Exam pointHow it's tested in interviewsPage on this site
vLLM overall architecture"Draw the three-layer relationship of LLMEngine / Scheduler / Executor"vLLM
PagedAttention"What data structure is the block table? Why does it solve fragmentation?"Batching and Request Scheduling
Continuous batching"Where does the throughput gain over static batching come from?"Batching and Request Scheduling
TensorRT-LLM In-Flight Batching"How does it differ from vLLM's continuous batching?"TensorRT-LLM
Chunked Prefill"Why split long prompts into chunks?"vLLM
Speculative decoding"Differences between Medusa and EAGLE? When should you not use it?"Speculative Decoding and Medusa/EAGLE
Prefix caching"How do repeated prompts save KV cache recomputation?"vLLM

2. KV cache and memory management ​

Exam pointHow it's tested in interviewsPage on this site
KV cache size calculation"Llama-70B, batch=8, seq=2048 — how big is the KV cache?"The GPU Memory Hierarchy and the Bandwidth Wall
KV cache quantization"How much memory does INT8 KV cache save? What's the accuracy cost?"Model Quantization Fundamentals
PagedAttention data structures"How is block_size chosen? How are physical blocks allocated?"Batching and Request Scheduling
Multi-Query / Grouped-Query Attention"How do MQA / GQA save KV cache?"vLLM
Ring / Paged Attention across nodes"How do PP and KV cache coordinate?"Distributed Inference (TP/PP)

3. Quantization and compression ​

Exam pointHow it's tested in interviewsPage on this site
PTQ vs QAT"What do you do when PTQ loses accuracy? When do you go QAT?"Model Quantization Fundamentals
GPTQ / AWQ / SmoothQuant"Mechanistic differences? Which suits which model?"Model Quantization Fundamentals
Weight-only vs weight+activation"Why does LLM inference mainly push weight-only?"Weight-Only Quantization and Mixed Precision
INT4 / INT8 / FP8"On H100, which is faster, FP8 or INT8 — and why?"Model Quantization Fundamentals
KV cache quantization"Where does the accuracy loss of INT8 KV cache quantization come from?"Model Quantization Fundamentals
Pruning and distillation"Why do LLM inference roles rarely talk about pruning?"Pruning and Sparsification, Knowledge Distillation

4. Batching and scheduling ​

Exam pointHow it's tested in interviewsPage on this site
Static batching"Why is static batching low-throughput for LLMs?"Batching and Request Scheduling
Continuous batching"How does the scheduler decide when to swap requests in and out?"Batching and Request Scheduling
In-Flight Batching"TensorRT-LLM's in-flight vs vLLM's continuous — differences?"TensorRT-LLM
Prefill / decode separation"Why separate them? What's the communication overhead after separation?"Distributed Inference (TP/PP)
Priority scheduling"How do you schedule VIP user requests without hurting regular users?"Model Serving and Orchestration

5. Serving and orchestration ​

Exam pointHow it's tested in interviewsPage on this site
Triton Server model repository"How do you configure a model repository with multi-version canary?"Triton Inference Server
FastAPI / asyncio serving"How does vLLM's OpenAI-compatible server implement streaming?"Model Serving and Orchestration
K8s deployment and autoscaling"How do you autoscale vLLM replicas based on QPS?"Model Serving and Orchestration
Canary and rollback"How do you run a 5% canary for a model rollout?"Model Serving and Orchestration
Monitoring metrics"How do you instrument TTFT / TPOT / queue length?"Inference Benchmarking in Practice

The review strategy for L2

Every L2 exam point maps to a trio of "open-source project source + paper + measurement." The most effective review: pick one engine (vLLM recommended), read its source closely + run one benchmark + change one parameter and observe the effect. Finish these three steps and you can answer every L2 question with "what I did" instead of "what I read." See Inference Benchmarking in Practice and Tuning and Performance Optimization.

4. L3 Advanced Layer: Kernels, Graphs, Distributed, Edge ​

JD words: "understanding of kernel optimization," "familiar with distributed inference," "edge deployment experience a plus." L3 is the "bonus / differentiator" layer of interviews — reaching this layer means you've passed the L1/L2 screen and they're using it to separate "can use" from "can modify."

1. Kernel development ​

Exam pointHow it's tested in interviewsPage on this site
FlashAttention principles"What did FlashAttention-1/2/3 each do?"Kernel Fusion and Custom Kernels
CUDA kernel writing"Hand-write a reduction kernel"GPU Architecture and Optimization
Triton DSL"Triton's advantages over CUDA? When should you use it?"Kernel Fusion and Custom Kernels
Kernel fusion strategies"How do you compute the launch-overhead savings of fusing Linear + ReLU?"Kernel Fusion and Custom Kernels
Tensor Core / WGMMA"On Hopper, how do wgmma instructions differ from mma?"GPU Architecture and Optimization
TMA (Tensor Memory Accelerator)"What problem does TMA solve? Why is H100 inference fast?"Hardware Primer
Bank conflict / coalesced access"How do you avoid shared-memory bank conflicts?"GPU Architecture and Optimization
CUDA Graph"Why does LLM inference capture CUDA Graphs?"Computation Graph Optimization

2. Graph optimization and compilation ​

Exam pointHow it's tested in interviewsPage on this site
Kernel fusion"What does fusing Conv + BN + ReLU save?"Computation Graph Optimization
Constant folding / dead-code elimination"What does torch.compile's constant folding do?"Computation Graph Optimization
torch.compile / Inductor"How much does torch.compile beat eager in inference, and why?"Computation Graph Optimization
AOT vs JIT"TensorRT is AOT and PyTorch eager is JIT — what does that mean?"Computation Graph Optimization
ONNX / TorchScript / Export"Why is torch.export better than TorchScript?"Computation Graph Optimization
MLIR / XLA / TVM"Where do compiler backends sit inside an inference engine?"Computation Graph Optimization

3. Distributed inference ​

Exam pointHow it's tested in interviewsPage on this site
Tensor Parallel"TP shards weights — how do you compute the communication overhead?"Distributed Inference (TP/PP)
Pipeline Parallel"How do you compute PP's bubble ratio?"Distributed Inference (TP/PP)
Expert Parallel (MoE)"EP vs DP? Why does MoE inference use EP?"Distributed Inference (TP/PP)
Disaggregated Prefill / Decode"Why separate them? How does the KV cache cross nodes?"Distributed Inference (TP/PP)
Multi-node inference"Why is cross-node TP slow? NVSwitch vs Ethernet?"Distributed Inference (TP/PP)
Ring Attention"What problem does Ring Attention solve for long context?"Distributed Inference (TP/PP)

4. Edge inference ​

Exam pointHow it's tested in interviewsPage on this site
llama.cpp / ggml"Why can llama.cpp run 70B on an M1?"llama.cpp and GGUF
MLC-LLM / Apache TVM"How does MLC's compilation path differ from llama.cpp?"Mobile Deployment
ARM NEON / SVE"What drives ARM CPU inference optimization?"Mobile Deployment
Apple Metal / ANE"On iPhone, does an LLM run on Metal or the ANE?"Mobile Deployment
WebGPU / Vulkan"Where's the bottleneck when running an LLM in the browser?"Mobile Deployment
INT4 group-wise quantization"Why does the edge mainly use INT4 group-wise?"Model Quantization Fundamentals

5. Performance metrics and benchmarks ​

Exam pointHow it's tested in interviewsPage on this site
TTFT / TPOT definitions"The physical meaning of TTFT and TPOT?"Latency, Throughput, and Concurrency
Compute-bound vs memory-bound"How do you tell which regime an LLM inference is in?"The Roofline Model and Compute Analysis
Roofline model"How do you use the Roofline model to compute the theoretical ceiling?"The Roofline Model and Compute Analysis
HBM bandwidth"Why is Llama-70B inference memory-bound?"The GPU Memory Hierarchy and the Bandwidth Wall
Benchmark methodology"How do you write a fair LLM inference benchmark?"Inference Benchmarking in Practice
Warmup and statistics"Why warm up? How do you take P50/P99?"Inference Benchmarking in Practice

5. Master Cross-Reference: JD Keywords → Site Pages ​

The four sections above are consolidated into a printable cross-reference table — your entry point for "read a JD → find a page":

JD keywordReal exam pointPage on this siteL1 / L2 / L3
Python / C++GIL, RAII, templates, PyBind11Glossary, vLLML1
Linux / performance analysisnsys / ncu / perf / PyTorch ProfilerInference Benchmarking in PracticeL1
Networking / gRPCTCP / HTTP / gRPC / RDMA / NCCLTriton Inference ServerL1
GPU architectureSIMT / Tensor Core / HBM / NVLinkHardware PrimerL1
vLLMArchitecture / PagedAttention / continuous batchingvLLML2
TensorRT-LLMIn-Flight Batching / Plugin / FP8TensorRT-LLML2
QuantizationGPTQ / AWQ / SmoothQuant / INT4 / INT8 / FP8Model Quantization Fundamentals, Weight-Only Quantization and Mixed PrecisionL2
KV cacheSize calculation / PagedAttention / MQA / GQABatching and Request Scheduling, The GPU Memory Hierarchy and the Bandwidth WallL2
BatchingContinuous / In-Flight / Prefill-Decode separationBatching and Request SchedulingL2
ServingTriton / FastAPI / K8s / canaryTriton Inference Server, Model Serving and OrchestrationL2
CUDAKernel / SIMT / bank conflict / Tensor CoreGPU Architecture and OptimizationL3
Triton DSLWrite attention / quantization kernelsKernel Fusion and Custom KernelsL3
FlashAttentionFA1/2/3 / tiling / recomputeKernel Fusion and Custom KernelsL3
CUDA GraphCapture & replay / launch overheadComputation Graph OptimizationL3
Graph optimizationKernel fusion / constant folding / torch.compileComputation Graph OptimizationL3
Distributed inferenceTP / PP / EP / disaggregatedDistributed Inference (TP/PP)L3
Speculative decodingMedusa / EAGLE / acceptance rateSpeculative Decoding and Medusa/EAGLEL3
Edge inferencellama.cpp / MLC / Metal / WebGPUllama.cpp and GGUF, Mobile DeploymentL3
Performance metricsTTFT / TPOT / Roofline / HBM bandwidthLatency, Throughput, and Concurrency, The Roofline Model and Compute AnalysisL3

How to use this table

Look up the "Page on this site" column row by row for each JD keyword, then mark "can do / can't / half-can" next to each page name — that gives you the entry point of your catch-up list. Don't try to fill the whole table — only fill the parts covered by "hard requirements + bonus items where you're one or two levels short."

6. Suggested Learning Priorities ​

Priorities are ranked by "return on investment + role fit":

Priority 1  L1 Python / C++ / Linux / performance analysis
            The foundation and the hard bar for all inference roles
            Investment: ongoing — 5–10 hours/week to maintain

Priority 2  L2 Inference engines (vLLM / TensorRT-LLM) + KV cache + batching
            The main interview line for LLM inference roles
            Investment: 2–4 weeks, close source reading + measurement

Priority 3  L2 Quantization + serving
            Standard equipment for all deployment roles
            Investment: 1–2 weeks, run comparison experiments

Priority 4  L3 Kernel development (CUDA / Triton) + graph optimization
            Differentiators for engine / optimization / kernel roles
            Investment: 3–6 weeks ongoing, starting from reduction

Priority 5  L3 Distributed inference + edge inference + speculative decoding
            Role-specific differentiators; fill per target role
            Investment: 1–2 weeks, papers + demos

The hard constraint on time allocation

L1 is always priority 1 — every role's hard bar includes Python + Linux, and it can't be skipped. L2 is priority 2 because LLM inference roles account for 70%+ of inference & deployment openings in 2026. L3 is customized to the target role: engine roles → kernels + graph optimization; systems roles → serving + K8s; edge roles → llama.cpp + quantization. See the "role type → learning path" mapping in Learning Paths: Three Routes.

7. Self-Assessment: The Five-Level Scale (Printable) ​

The exam points from the previous sections are consolidated into a printable self-assessment table. Scoring criteria:

LevelMeaningCriterion (grade yourself in the mirror)
1 Never heard of itCompletely unfamiliarHaven't even seen the term
2 Heard of itVague impressionSeen or heard it, but can't give a definition
3 AwareKnow the conceptCan give a definition and one example, but not the principles or trade-offs
4 FamiliarCan use it hands-onCan implement / tune / troubleshoot independently; knows pros, cons, and applicable scenarios
5 Can teach itCan teach othersCan explain principles + derivation + trade-offs + counterexamples in 5 minutes, and survive follow-ups

How to print and check

Print this section, check one box per row for your current, honest level. Honesty is the premise — the goal is to expose every row below "can teach it," not to make the table look good. Check only one level per row; when in doubt, check the lower one.

Block A: L1 Fundamentals ​

Exam pointNever heardHeardAwareFamiliarCan teach
Python GIL / asyncio□□□□□
C++ RAII / smart pointers / templates□□□□□
Linux performance analysis (perf / nsys)□□□□□
Reading Nsight Compute reports□□□□□
PyTorch Profiler□□□□□
TCP vs gRPC vs HTTP□□□□□
RDMA / NCCL communication□□□□□
GPU SIMT / memory hierarchy□□□□□
Tensor Core / TMA□□□□□
H100 vs A100 key differences□□□□□

Block B: L2 Engines and Scheduling ​

Exam pointNever heardHeardAwareFamiliarCan teach
vLLM architecture / Scheduler□□□□□
PagedAttention block table□□□□□
Continuous batching□□□□□
In-Flight Batching (TRT-LLM)□□□□□
Chunked Prefill□□□□□
Prefix caching□□□□□
KV cache size calculation□□□□□
KV cache quantization□□□□□
MQA / GQA□□□□□
Triton Server model repository□□□□□
K8s deployment and autoscaling□□□□□
Canary rollout and rollback□□□□□
TTFT / TPOT monitoring□□□□□

Block C: L2 Quantization and Compression ​

Exam pointNever heardHeardAwareFamiliarCan teach
PTQ vs QAT□□□□□
GPTQ / AWQ / SmoothQuant□□□□□
Weight-only vs weight+activation□□□□□
INT4 / INT8 / FP8 selection□□□□□
Kernel fusion / graph optimization concepts□□□□□
Limits of pruning / distillation for LLMs□□□□□

Block D: L3 Kernels and Graphs ​

Exam pointNever heardHeardAwareFamiliarCan teach
FlashAttention-1/2/3 principles□□□□□
Hand-writing a CUDA reduction kernel□□□□□
Triton DSL attention□□□□□
Kernel fusion strategies□□□□□
Tensor Core / WGMMA / TMA□□□□□
Bank conflict / coalesced access□□□□□
CUDA Graph capture□□□□□
torch.compile / Inductor□□□□□
AOT vs JIT compilation□□□□□

Block E: L3 Distributed and Edge ​

Exam pointNever heardHeardAwareFamiliarCan teach
Tensor Parallel communication overhead□□□□□
Pipeline Parallel bubble□□□□□
Expert Parallel (MoE)□□□□□
Disaggregated Prefill/Decode□□□□□
Ring Attention□□□□□
llama.cpp / ggml□□□□□
MLC-LLM / TVM□□□□□
ARM NEON / Apple Metal□□□□□
WebGPU / Vulkan□□□□□
Speculative decoding (Medusa / EAGLE)□□□□□
Roofline model□□□□□
HBM bandwidth and memory-bound□□□□□

8. How to Turn the Self-Assessment into a Personal Catch-Up List ​

After finishing the self-assessment, turn it into an executable weekly plan in four steps.

Step 1: Group by level ​

Classify each exam point into three buckets:

  • Class A (levels 1–2, completely unfamiliar): highest priority — first solve "what is it." Goal: reach level 3 within a week.
  • Class B (level 3, aware): what's missing is principles and trade-offs. Goal: reach level 4 within two weeks — the standard is "can implement independently + state applicable scenarios."
  • Class C (level 4, familiar): what's missing is "teach-it" ability. Goal: turn each point into a 5-minute spoken talk. This maps directly to interview performance.

Level-5 points need no more time unless you're targeting a teaching- or theory-heavy role.

Step 2: Rank by block and total the debt ​

Aggregate the three classes by block to get three "debt tables." Suggested catch-up order:

Priority 1  Block A L1 fundamentals     — the foundation under all other blocks
Priority 2  Block B L2 engines & scheduling — the main interview line for LLM
                                          inference roles; highest ROI
Priority 3  Block C L2 quantization      — required by nearly all deployment roles
Priority 4  Block D L3 kernels & graphs  — if the target role is engine /
                                          optimization / kernel
Priority 5  Block E L3 distributed & edge — fill per target role

Step 3: Apply the "catch-up list" template ​

WeekFocus blockClass A tasks (to level 3)Class B tasks (to level 4)Class C tasks (write talks)Deliverable
Week 1L1 fundamentals3 unfamiliar points2 points: add principles1 talk (e.g., GIL)Notes + talk
Week 2L2 engines + KV cache2 unfamiliar points3 points: add principles2 talks (e.g., PagedAttention, continuous batching)Source notes + talks
Week 3L2 quantization + serving1 unfamiliar point2 points: add principles2 talks (e.g., GPTQ vs AWQ, TTFT/TPOT)Comparison experiment + talks
Week 4L3 kernels / distributed1 unfamiliar point2 points: add principles2 talks (e.g., FlashAttention, TP)Hand-written kernel + talks
Week 5+Loop until A/B are clearedAdjust dynamicallyAdjust dynamicallyPush toward "combination" questionsMock-interview recordings

Step 4: Three execution principles ​

  • Principle 1: it doesn't count until you can say it. After reviewing each point, record a 5-minute self-talk on your phone and listen back — wherever you stumble is what you didn't actually understand.
  • Principle 2: source + measurement beats documentation. For engine-type points, you must run vLLM yourself once, change one parameter once, and read one Nsight report — "read the docs" and "ran the measurement" are two different things in an interview.
  • Principle 3: retest weekly. On Sunday night, rescore with the self-assessment table, cross out the points that leveled up, and flag the ones that didn't move — then ask why. It's usually not lack of effort but the wrong review method (e.g., only watching, never practicing).

One more thing

The goal of a catch-up list is not "coloring the whole table in" — it's getting every threshold item in your target role's JD to level 4 or above. Blocks your role doesn't need (e.g., edge inference for an engine role) can sit at level 4; put your time on the blade's edge.

9. Further Reading ​

On this site:

References (real resources):