Appearance
Interview Question Bank
Start by memorizing one sentence for this page: an interview doesn't test knowledge — it tests whether you can use knowledge to solve problems you've never seen. Rote question-drilling alone cannot get you through an inference & deployment interview, but walking in without ever looking at questions wastes the opportunity — this page gives you two things: the logic behind the questions, and the framework your answers should follow. The point of drilling is to turn high-frequency exam points into reflexes, not to hit the exact wording.
This page is the final step — "prove yourself" — of the career module's four-step pipeline. Finish the first three steps first (read the market, Deconstructing JD Knowledge Points, resume polishing), then come here. When time is tight, execute the checklist in Section 10, Interview Prep Plan.
1. The Interview Landscape: What Each Round Is Really Testing
An inference & deployment interview typically runs five to seven rounds; most companies complete 2–4 rounds in a single day. Every round has its own pass bar — don't substitute "my last round went well" for "how should I play this round."
Typical interview flow for inference & deployment roles (big and mid-size companies)
┌─────────────┬─────────────────────────────┬──────────────────────────────┐
│ Round │ Format │ Pass bar (interviewer's rubric)│
├─────────────┼─────────────────────────────┼──────────────────────────────┤
│ ① Self-intro│ 3–5 minutes verbal │ Clear logic, clear positioning,│
│ │ │ role fit │
│ ② Systems │ Python/C++/Linux Q&A │ Solid engineering foundation; │
│ basics │ │ can explain GIL/RAII │
│ ③ Engine & │ vLLM/TensorRT/PagedAttention│ Source-level understanding; │
│ scheduling│ │ can discuss trade-offs; can │
│ │ │ locate bottlenecks │
│ ④ Kernels & │ CUDA/Triton/INT8/FP8 Q&A │ Knows the mechanisms; can │
│ quantization│ │ discuss bank conflicts │
│ ⑤ Project │ Serial follow-ups on resume │ Really did it; has numbers; │
│ deep-dive│ projects │ survives follow-ups │
│ ⑥ System │ Design a 1000-QPS LLM │ Has a framework; makes │
│ design │ service │ trade-offs; knows boundaries │
│ ⑦ Live │ KV cache draft / reduction │ Correctness + complexity + │
│ coding │ kernel │ communication │
│ ⑧ HR │ Conversation │ Stability, motivation, │
│ │ │ soft skills │
└─────────────┴─────────────────────────────┴──────────────────────────────┘What each round is truly testing:
- ① Self-introduction: not a retelling of the resume. What the interviewer wants is a "who you are → what you've done → why you fit this role" three-sentence skeleton. 30 seconds of background, 2 minutes on 1–2 projects with performance numbers, 30 seconds of motivation. End with a hook ("I've been landing PRs on vLLM's chunked prefill lately — which maps exactly to the long-context optimization in your JD") to steer the rhythm toward territory you've prepared.
- ② Systems fundamentals: testing whether the engineering foundation is solid. Python GIL / asyncio, C++ RAII / smart pointers, and Linux performance analysis are the high-frequency trio. Fail this round and nothing later can save you.
- ③ Engine & scheduling: the "load-bearing round" of this role. How vLLM's Scheduler works, what PagedAttention's block table looks like, the difference between continuous and in-flight batching — these questions separate "read the docs" from "read the source."
- ④ Kernels & quantization: testing whether you can drop to the CUDA / Triton level. Why FlashAttention is fast, how to avoid bank conflicts, INT8 vs FP8 selection — this round is the watershed between engine roles and systems roles.
- ⑤ Project deep-dive: the interviewer uses serial follow-ups to verify authenticity — "actually did it" and "watched someone else's project" are instantly distinguishable after three levels of follow-up. Detail questions include: what did you set vLLM's
--gpu-memory-utilizationto and why; which metrics did you look at in the Nsight report; how did you pick batch and seq_len for the benchmark. See Section 6. - ⑥ System design: open-ended questions ("design an LLM service that sustains 1000 QPS," "design a multi-model inference platform"). The rubric is not "the standard answer" but whether you have a mental framework from requirements → engine selection → deployment architecture → monitoring & alerting → capacity planning, and whether you can make trade-offs under latency, cost, and availability constraints.
- ⑦ Live coding: generic algorithm problems (sorting / binary search / DP) plus inference-related handwriting (KV cache draft, reduction kernel). The bar is correct code, clearly stated complexity, and communication throughout. When stuck, explain the brute-force solution first and then optimize — vastly better than coding in silence.
- ⑧ HR round: stability (why are you leaving / why us), motivation, collaboration, stress tolerance. Core principle: positive, specific, never complain about your current employer. Salary negotiations most often collapse when you name an absolute number ("I want X") — quote a range and ask about their budget instead.
One principle that runs through every round: STAR
Project, behavioral, and design questions can all be answered with STAR: Situation (background) → Task (the problem) → Action (what you did — emphasize your decisions) → Result (quantified outcome). The interviewer retains zero from "we had a slow inference service back then," but remembers "tuned vLLM --max-num-batched-tokens from 2048 to 4096, TTFT from 1.8s to 0.6s, throughput +120%."
2. Performance & Metrics Questions (4)
1. What's the difference between TTFT and TPOT? What drives each?
- What's tested: the meta-concepts of this role — a guaranteed question.
- Reference answer:
- TTFT (Time To First Token): from request arrival to the first output token. Dominated by the prefill phase — compute grows quadratically with prompt length; affected by GPU compute, KV cache memory allocation, and in-batch scheduling.
- TPOT (Time Per Output Token): average time per token after the first one. Dominated by the decode phase — every generated token needs one forward pass containing a single token per request, which is memory-bound; driven most by HBM bandwidth, KV cache size, and batch size.
- The optimization directions differ: TTFT improves via chunked prefill, prefix caching, and faster attention kernels (FlashAttention); TPOT improves via continuous batching (more parallelism), KV cache quantization (less memory traffic), and speculative decoding (fewer forwards).
- Bonus points:
- "In production, the business side cares about P99, not the mean — report both TTFT and TPOT as P50/P99";
- "There's a trade-off between TTFT and TPOT: enabling chunked prefill raises TPOT slightly but cuts TTFT — tune by business weight."
- Map to this site: Latency, Throughput, and Concurrency, Batching and Request Scheduling, Inference Benchmarking in Practice.
2. How do you tell memory-bound from compute-bound?
- What's tested: intuitive grasp of the Roofline model.
- Reference answer:
- Use the Roofline model: the x-axis is arithmetic intensity (FLOPs / byte), the y-axis is throughput (FLOP/s). An operator's ceiling = min(peak compute, HBM bandwidth × arithmetic intensity).
- If the operator sits on the "bandwidth slope" → memory-bound, limited by HBM bandwidth;
- if on the "compute plateau" → compute-bound, limited by Tensor Core throughput.
- The signature of LLM inference: decode is almost always memory-bound (each generated token performs only a tiny batch×hidden matmul, but the entire weight set must stream from HBM — arithmetic intensity is extremely low); prefill trends compute-bound with long prompts.
- Measurement: check each kernel's
Achieved OccupancyandDRAM Throughputin Nsight Compute — low occupancy with saturated DRAM means memory-bound.
- Bonus points: "Llama-70B decode has arithmetic intensity around 0.5 FLOP/Byte; A100 HBM bandwidth 2TB/s × 0.5 = 1 TFLOP/s, far below the 312 TFLOP/s peak — hence memory-bound. That's why weight-only quantization speeds up decode roughly linearly: less bandwidth = more speed." See The GPU Memory Hierarchy and the Bandwidth Wall and The Roofline Model and Compute Analysis.
- Map to this site: The Roofline Model and Compute Analysis, The GPU Memory Hierarchy and the Bandwidth Wall, GPU Architecture and Optimization.
3. Where do LLM inference latency bottlenecks usually hide? How do you locate them?
- What's tested: hands-on performance-tuning experience.
- Reference answer: troubleshoot in four layers — "scheduling → prefill → decode → communication":
- ① Scheduling: vLLM's
Waiting/Runningqueue lengths, whethermax_num_seqsis the constraint; Triton'sdynamic_batchingqueue; - ② Prefill: long prompts with insufficient chunked-prefill budget, KV cache allocation failures;
- ③ Decode: batch too small to saturate HBM bandwidth, attention kernel falling off the FlashAttention path;
- ④ Communication: cross-node TP all-reduce consuming > 30% of forward time, NCCL going over PCIe instead of NVLink/RDMA.
- Toolchain: Nsight Systems to see the phase breakdown inside a forward; PyTorch Profiler for per-kernel costs; vLLM's
--profilemode to generate traces.
- ① Scheduling: vLLM's
- Bonus points: "With 70B across 2 nodes at TP=16, I once traced 38% of forward time to NCCL all-reduce; setting
NCCL_NET_GDR_LEVEL=5moved it to RDMA and cut it to 15%. That segment shows up asnccl-allreducein the nsys trace." - Map to this site: Latency, Throughput, and Concurrency, Distributed Inference (TP/PP), Tuning and Performance Optimization.
4. Why can't LLM inference batch size grow without limit?
- What's tested: understanding the scheduling-vs-memory trade-off.
- Reference answer: three constraints:
- ① Memory: KV cache grows linearly with batch × seq —
max_num_seqs × max_num_batched_tokens × 2 × num_layers × hidden_size × dtype_size ≤ available memory; - ② Latency: larger batches make each forward bigger, so single-forward time grows; in decode, TPOT rises with batch (arithmetic intensity rises but latency grows);
- ③ Scheduling: vLLM's continuous batching swaps finished requests for new ones at step boundaries; oversized batches lower swap frequency and can reduce throughput.
- The optimum usually sits between 32 and 256 and must be measured per hardware + model + SLO.
--max-num-seqsand--max-num-batched-tokensare vLLM's two core scheduling knobs.
- ① Memory: KV cache grows linearly with batch × seq —
- Bonus points: "I measured Llama-70B on A100 80G: throughput scaled linearly from batch 8 to 64, flattened from 64 to 128, and dropped beyond 128 as KV cache pressure triggered evictions."
- Map to this site: Batching and Request Scheduling, The GPU Memory Hierarchy and the Bandwidth Wall, Inference Benchmarking in Practice.
3. Quantization Questions (4)
1. GPTQ vs AWQ — differences? Which for which scenario?
- What's tested: quantization method selection — a must for optimization roles.
- Reference answer:
- GPTQ: weight quantization based on second-order Hessian information, solving for per-column weight updates that minimize output error. Needs calibration data (~128 samples); accuracy loss usually < 1%. Suits accuracy-sensitive scenarios (chat models, instruction following).
- AWQ: quantization based on "weight importance" — keeps high precision only for outlier channels (those with large activations) and quantizes the rest to INT4. No second-order information needed; fast and memory-friendly. Suits resource-constrained / rapid-quantization scenarios.
- Mechanistic difference: GPTQ minimizes output error with a backpropagation-style procedure (calibrated on pre-activations); AWQ directly identifies "important weights" by activation magnitude and protects them.
- Measured comparison: on Llama-70B, GPTQ raises perplexity 0.3%, AWQ 0.4%, but AWQ quantizes 3× faster.
- Bonus points: "SmoothQuant is a third idea — smooth activation outliers so both weights and activations can go INT8. LLM inference pushes weight-only (GPTQ / AWQ) because decode is memory-bound and weight quantization directly cuts bandwidth; activation quantization degrades LLM accuracy badly."
- Map to this site: Model Quantization Fundamentals, Weight-Only Quantization and Mixed Precision.
2. Why is weight quantization so effective for LLM inference?
- What's tested: the "why" behind quantization — depth of understanding.
- Reference answer: two reasons:
- ① Decode is memory-bound: arithmetic intensity around 0.5 FLOP/Byte; the bottleneck is HBM bandwidth. Weights dominate memory (Llama-70B FP16 = 140GB); INT4 weight-only compresses them to 35GB — a direct 4× cut in HBM traffic → roughly linear decode speedup.
- ② Weight outliers are structurally sparse: LLM weights have a few outlier channels, but the outliers are structural — the same channels are outliers across tokens — so group-wise quantization or AWQ's "protect important channels" mechanism handles them, keeping accuracy loss small.
- By contrast, activation quantization degrades LLMs badly — activation outliers are token-level (different tokens have outliers at different positions), so INT8 quantization gets dominated by outliers and accuracy collapses. This is the fundamental reason LLM inference pushes weight-only.
- Bonus points: "INT4 weight-only compresses Llama-70B from 140GB to 35GB, so a single H100 80G can run it — FP16 needs 2-GPU TP. Quantization here solves a 'won't fit' hardware problem, not just a speed problem."
- Map to this site: Model Quantization Fundamentals, Weight-Only Quantization and Mixed Precision, The GPU Memory Hierarchy and the Bandwidth Wall.
3. Where does INT4 quantization lose accuracy?
- What's tested: depth on quantization accuracy.
- Reference answer: four sources:
- ① Outlier-channel saturation: a few outlier channels have magnitudes far above the rest; with only 16 values in 4 bits, outliers get clamped to max → accuracy loss concentrates in outlier channels. AWQ mitigates by keeping outlier channels in FP16.
- ② Group-wise quantization error: larger group_size means larger intra-group spread between outliers and the rest, hence larger error; group_size=128 is the empirical sweet spot.
- ③ Zero-point drift: symmetric quantization (zero_point=0) suffers when the weight distribution is offset; asymmetric quantization needs a calibrated zero_point.
- ④ KV cache quantization stacking: enabling INT4 weights + INT8 KV cache simultaneously accumulates error — evaluate each separately.
- Typically, INT4 group-wise quantization raises perplexity 0.3–0.6% on Llama-family models, imperceptible in blind business-side testing.
- Bonus points: "I ran AWQ on Llama-3-8B: group_size=128 gave +0.4% perplexity, group_size=32 gave +0.2%, but group_size=8 actually increased memory (more scale and zero-point metadata) — 32 is the cost-effectiveness sweet spot."
- Map to this site: Model Quantization Fundamentals, Weight-Only Quantization and Mixed Precision.
4. FP8 vs INT8 — and why is FP8 fast on H100?
- What's tested: new hardware × precision formats.
- Reference answer:
- Format difference: INT8 is an 8-bit integer (-128 to 127) needing scale and zero-point; FP8 comes in E4M3 (4-bit exponent + 3-bit mantissa) and E5M2 (5-bit exponent + 2-bit mantissa) — floating point with a dynamic range, no scale needed.
- Accuracy difference: FP8 usually loses less accuracy on LLMs (outliers in weights and activations are safer in floating point); INT8 weight-only remains the LLM mainstream, while FP8 weight + activation is the new trend on H100.
- Why FP8 is fast on H100: H100 introduced the Transformer Engine + FP8 Tensor Core, doubling FP8 matmul throughput versus FP16 (1979 TFLOP/s vs 989 TFLOP/s). FP8 also skips the quantize/dequantize conversion overhead (compare INT8's dequantize kernel), saving kernel launches and memory traffic.
- When not to use FP8: A100 and earlier have no FP8 Tensor Core; accuracy-sensitive scenarios (e.g., RLHF training) still prefer BF16; INT8 is more universal on edge and CPU.
- Bonus points: "FP8 on H100 isn't just faster — NVIDIA deeply integrated the Transformer Engine with FP8: quantization, scale management, and loss scaling happen at the hardware layer, and the software stack (TransformerEngine, TensorRT-LLM, vLLM 0.8+) supports it natively."
- Map to this site: Model Quantization Fundamentals, GPU Architecture and Optimization, Hardware Primer.
4. KV Cache & Batching Questions (4)
1. How do you compute KV cache size?
What's tested: fundamentals — asked almost every time.
Reference answer: the formula:
KV cache (bytes) = 2 (K and V) × num_layers × batch_size × seq_len × head_dim × num_kv_heads × dtype_size- 2: two caches, K and V;
- num_layers: number of Transformer layers;
- batch_size × seq_len: every token stores one K, V pair;
- head_dim × num_kv_heads: per-head dimension and number of KV heads;
- dtype_size: FP16 = 2 bytes, INT8 = 1 byte.
Example: Llama-70B (80 layers, 64 heads, 8 KV heads, head_dim=128, FP16), batch=8, seq=2048:
2 × 80 × 8 × 2048 × 128 × 8 × 2 = 5.37GBMQA / GQA impact: num_kv_heads drops from num_heads to 1 (MQA) or num_heads/8 (GQA), cutting KV cache to 1/8 directly — a big part of why GQA matters so much for LLMs.
Bonus points: "Sizing the KV cache is the prerequisite for tuning vLLM's
--gpu-memory-utilization: available memory = total memory - weight memory - temporary buffers - KV cache. When I tuned 70B on a single A100 80G, I first computed FP16 weights at 140GB wouldn't fit → INT4 weights at 35GB → 45GB left for KV cache → derive the batch × seq ceiling."Map to this site: The GPU Memory Hierarchy and the Bandwidth Wall, Batching and Request Scheduling, Model Quantization Fundamentals.
2. How does continuous batching work? What problem does it solve?
- What's tested: the core mechanism of vLLM and TensorRT-LLM.
- Reference answer:
- The problem with static batching: requests in a batch have different lengths; short requests wait for the long ones while the GPU idles; new requests wait for the whole batch to finish, stacking latency.
- Continuous batching: at each forward boundary (one step), the scheduler checks which requests have finished, swaps them out, swaps new requests in from the queue, and starts the next step — transparent to the KV cache of unfinished requests.
- Key data structure: PagedAttention's block table — the KV cache isn't allocated contiguously but paged in blocks (e.g., 16 tokens per block); requests can share free blocks, eliminating memory fragmentation.
- Effect: 2–10× throughput over static batching, 10×+ over no batching. vLLM's continuous batching and TensorRT-LLM's in-flight batching are the same idea with different implementations.
- Bonus points: "Continuous batching's limit: each step's forward computes all tokens in the batch together, and when a short request finishes and a new request swaps in, the new request's prefill mixes with old requests' decode in the same batch — that's the problem chunked prefill solves, slicing long-prompt prefills so they don't hog the whole forward budget."
- Map to this site: Batching and Request Scheduling, vLLM, TensorRT-LLM.
3. What problem does PagedAttention solve?
- What's tested: vLLM's signature mechanism.
- Reference answer:
- The problem: the traditional KV cache allocates contiguous memory per request — a request reserves
max_seq_lenof contiguous memory up front, causing two issues: ① fragmentation: finished short requests leave small holes that long requests can't fit into; ② over-reservation: actual generation length is usually far belowmax_seq_len, so the reserved portion is wasted. - The PagedAttention fix: borrow OS virtual memory — the KV cache is stored in fixed-size blocks (e.g., 16 tokens per block), logically contiguous but physically discrete. The block table records each request's logical-block → physical-block mapping. Finished requests return blocks straight to the free pool; no fragmentation.
- Key mechanisms: ① block table: one per request, recording physical locations of the KV cache; ② copy-on-write: shared blocks for multi-branch scenarios like beam search, copied only on modification; ③ prefix sharing: requests with identical prefixes (e.g., the same system prompt) share KV blocks, saving recomputation.
- Effect: memory utilization rises from 30–50% to 90%+; throughput improves 2–24× over static batching (vLLM paper numbers).
- The problem: the traditional KV cache allocates contiguous memory per request — a request reserves
- Bonus points: "The side effect: the attention kernel must index KV through the block table instead of accessing it contiguously, costing some performance (cache misses). vLLM kept block_size=16 as the default even after 0.8 — it's a trade-off between performance and memory utilization; larger blocks approach contiguous access but waste more memory."
- Map to this site: Batching and Request Scheduling, vLLM, The GPU Memory Hierarchy and the Bandwidth Wall.
4. What is chunked prefill, and why did vLLM 0.8 introduce it?
- What's tested: understanding of vLLM's evolution — separates users from source readers.
- Reference answer:
- The problem: when continuous batching admits a new request, its prefill compute grows quadratically with prompt length; a long prompt can hog the entire forward budget, leaving old requests' decode waiting and TPOT jittery.
- The chunked-prefill fix: split a long prompt's prefill into chunks (N tokens each) and compute one chunk per step, mixed with old requests' decode in the same forward. Long prompts no longer monopolize the forward budget, and old decode latency stays stable.
- Key parameter:
--max-num-batched-tokenscaps the total token budget per step (e.g., 4096). Too small → long prompts sliced too thin, prefill drags; too large → old decode waits. Tune per business. - When to enable: mandatory when the business mixes long prompts (> 4k tokens) with short ones; pure short-prompt businesses can disable it to save scheduling overhead.
- Bonus points: "Chunked prefill vs TTFT / TPOT: enabling it raises TTFT (prefill slower overall once sliced) but lowers TPOT (old decode no longer blocked) — a classic latency-vs-latency trade-off."
- Map to this site: Batching and Request Scheduling, vLLM, Latency, Throughput, and Concurrency.
5. Kernel Questions (3)
1. Why is FlashAttention fast?
- What's tested: a kernel-optimization must — understanding the mechanism, not the name.
- Reference answer: three mechanisms:
- ① Tiling: Q, K, V are tiled into SRAM (shared memory), avoiding repeated HBM round trips. The naive attention's S = QKᵀ is an O(N²) intermediate that gets written back to HBM and read again for softmax; FlashAttention fuses the entire softmax inside the tile, so intermediates never hit HBM.
- ② Recompute: skip storing intermediate attention matrices for backward; the forward stores only Q, K, V and the softmax normalization statistics, and attention is recomputed in backward — trading compute for memory.
- ③ IO-aware: it doesn't reduce FLOPs (they actually increase slightly) — it reduces HBM accesses. Through the Roofline lens, attention is memory-bound; fewer HBM accesses = speed.
- Effect: FlashAttention-1 delivered 2.3×–4× attention speedups on A100 (small to long seq); FlashAttention-2 improved GPU thread scheduling for another ~2×; FlashAttention-3 uses TMA + FP8 on H100 for another 1.5–2×.
- Bonus points: "FlashAttention's core isn't a 'faster algorithm' — it's 'less IO.' In the Nsight Compute report, naive attention saturates
DRAMthroughput; with FlashAttention,DRAMutilization drops to ~30% and the bottleneck shifts toSM Compute." See Kernel Fusion and Custom Kernels and The GPU Memory Hierarchy and the Bandwidth Wall. - Map to this site: Kernel Fusion and Custom Kernels, The GPU Memory Hierarchy and the Bandwidth Wall, GPU Architecture and Optimization.
2. Where are the limits of kernel fusion? When should you not fuse?
- What's tested: judgment about kernel optimization.
- Reference answer: the benefits and the boundaries:
- Benefits: ① fewer HBM accesses (intermediates flow through SRAM); ② fewer kernel launches (one launch runs a group of operators); ③ less memory (no intermediate tensors allocated).
- Boundaries:
- ① Very different operator shapes: fusing Conv (N,C,H,W) with MatMul (M,K,N) needs layout conversion whose cost may exceed the fusion gain;
- ② Complex operator semantics: operators with if/else branches or dynamic shapes (e.g., padding masks) are hard to fuse — compilers often give up;
- ③ Cross-block reductions: reductions spanning blocks (e.g., full layer norm) need atomics or two-phase kernels — fusion can be slower;
- ④ Dynamic shapes: with dynamic shapes, fused tile sizes can't be predicted and performance suffers.
- When not to fuse: a single kernel already near peak (e.g., cuBLAS GEMM at 95% of peak on large matrices); or the fusion needs a custom kernel whose maintenance cost outweighs the gain (calling PyTorch eager is more stable).
- Bonus points: "Adjacent fusion like Linear+ReLU+Linear is obvious. But Transformer attention needs a standalone FlashAttention precisely because the fusion requires algorithm-level redesign (recompute + tiling), not simple kernel chaining. That's why the truly advanced form of fusion demands algorithm-hardware co-design."
- Map to this site: Kernel Fusion and Custom Kernels, Computation Graph Optimization, GPU Architecture and Optimization.
3. What is a bank conflict in CUDA, and how do you avoid it?
- What's tested: CUDA fundamentals — a must for kernel engineers.
- Reference answer:
- Bank conflict: shared memory is divided into 32 banks (4 bytes wide each, independently accessible). When 32 threads in a warp access shared memory simultaneously and two or more threads hit different addresses in the same bank → serialization → performance loss.
- Types: ① 2-way bank conflict: 2 threads in one bank → 2× latency; ② 32-way bank conflict: all 32 threads in one bank → 32× latency (worst case).
- Avoidance: ① padding: add a padding column to the data layout (33 columns instead of 32) to skew the banks; ② swizzling: remap threads to data (e.g., thread (i,j) reads position (j, i ⊕ j)) so consecutive threads hit different banks; ③
__ldmatrix(Hopper): hardware-level ldmatrix instructions avoid bank conflicts automatically. - Detection: the
L1/Histogram Shared Bankmetric in Nsight Compute — a 32-way conflict showing up there is a bug.
- Bonus points: "CUTLASS's GEMM swizzle is the classic case — the thread-to-tile mapping uses an XOR swizzle so a warp hits different banks across iterations. Reading CUTLASS's swizzle is harder than writing a working GEMM — it's the real watershed between 'can write CUDA' and 'understands CUTLASS.'"
- Map to this site: GPU Architecture and Optimization, Kernel Fusion and Custom Kernels.
6. Project Deep-Dive Questions (8)
The project round has the highest elimination rate for this role, because resume fraud and "credit-only projects" have nowhere to hide here. There is only one strategy: make sure the project is something you actually did, then organize it as "background → my decisions → performance data → retrospective." Below are answer frameworks for the high-frequency questions.
1. Tell me about the inference optimization project you're proudest of
- Answer framework: use STAR, and deliberately emphasize "decisions I made" over "things we did." Structure: one sentence of background (business problem + model + hardware) → three steps of action (engine selection, parameter tuning, kernel rewrite — and why each) → hard-metric results (TTFT, TPOT, QPS, GPU utilization before and after) → one retrospective sentence (what I'd do differently a second time). Keep it to 2–3 minutes and leave room for follow-ups.
- Common mistakes: narrating the team without a personal thread; only success, no trade-offs; no performance numbers.
- Handling follow-ups: prepare three "dig points" — ① source details (which file, which line you changed); ② measurement context (how the benchmark ran, warmed up or not, P50 or P99); ③ pitfalls (which parameter you set wrong, which upgrade broke what).
2. Why did you pick vLLM over TensorRT-LLM? Or anything else?
- Answer framework: the comparison-decision trio — model type (open-source LLM or not), business constraints (cost, maintainability, ecosystem), team stack (Python-first or C++-first). Example: "The business was an internal Llama-70B assistant with a Python-first stack; we chose vLLM 0.8 because ① the open-source community is active, so issues get community help; ② Python-led, so the team ramped fast; ③ chunked prefill and prefix caching arrived in 0.8, mapping to our long-prompt scenario. TensorRT-LLM is indeed 10–15% faster, but its build pipeline is complex, upgrades are painful, and the team lacked C++ capacity — vLLM was the pragmatic call."
- What's tested: not picking "right," but having an explicit decision logic — being able to say "why not X" is worth more than "why Y."
3. Your online inference P99 just spiked to 8 seconds — how do you investigate?
- Answer framework: troubleshoot in five layers — "traffic → scheduling → engine → kernel → hardware":
- ① Traffic: did QPS suddenly rise? Did the long-prompt share spike? Confirm on the Grafana dashboard.
- ② Scheduling: is vLLM's
Waitingqueue backing up? Ismax_num_seqsat its cap? Ismax_num_batched_tokenstoo small? - ③ Engine: did the vLLM version change recently? Did some prompt fall off the FlashAttention path? Did prefix caching silently fail?
- ④ Kernel: Nsight Systems trace — which kernel's time ballooned inside the forward? Is NCCL all-reduce consuming > 30%?
- ⑤ Hardware: did GPU utilization drop (suggesting an I/O stall)? Is memory nearly full (heavy KV cache eviction)? Did the cross-node NVLink drop?
- What's tested: engineering triage thinking. Answering with the ordering "suspect traffic first, then scheduling, then the engine, and only then hardware" earns strong points.
4. The service slowly degrades after launch — what do you do?
- Answer framework: this is a mix of distribution drift and KV cache fragmentation. The action list:
- ① build monitoring: five lines — TTFT, TPOT, queue length, GPU utilization, KV cache eviction rate — with alert thresholds set in advance;
- ② identify the drift type: has the user prompt-length distribution shifted (business peaks bring more long prompts)? Did the model version change?
- ③ countermeasures: a) restart instances regularly to clear KV cache fragmentation (vLLM 0.8+'s
PrefixCachingmitigates, but periodic restarts remain); b) re-tunemax_num_seqs: old values stop being optimal as the business changes; c) multi-replica canary: try new parameters on 10% of replicas first. - Also answer "how to prevent it": design the monitoring dashboards before launch, with automatic alerts when TTFT P99 exceeds thresholds.
- What's tested: whether you understand that "an inference service is a perishable good," plus MLOps instincts for the monitoring/scheduling loop. See Model Serving and Orchestration and Common Pitfalls and Anti-Patterns.
5. What's the biggest pit you hit in this project?
- Answer framework: honesty plus a complete loop: what the pit was → how I found it → how I fixed it → what it distilled into. Prefer technical pits over blaming others. Classic strong material:
- chunked prefill parameter pit:
max_num_batched_tokensset too large, stalling old decode; - NCCL over PCIe instead of NVLink: cross-node TP absurdly slow, fixed with
NCCL_NET_GDR_LEVEL; - vLLM upgrade regression: prefix caching behavior changed from 0.6 to 0.8, cache hit rate dropping 80% → 30%;
- quantization accuracy drop: INT4 group_size=128 blew up perplexity in one business scenario; traced to an abnormal outlier distribution, fixed with group_size=32.
- chunked prefill parameter pit:
- What's tested: authenticity and retrospection. Interviewers know projects never go smoothly — "never hit a pit" is itself the danger signal.
6. How did you run your benchmark? Is the methodology trustworthy?
- Answer framework: present benchmarking as an engineering project of its own:
- ① Warmup: 10 warmup runs to stabilize KV cache and CUDA context before timing;
- ② Sampling: 1000 iterations, report P50/P90/P99 — never the mean (long-tail distributions make means meaningless);
- ③ Fixed variables: model version, batch, seq_len, hardware, concurrency all pinned — change one variable at a time for attribution;
- ④ Control group: every run needs a comparator (old version, another engine, another parameter set) — absolute numbers alone mean nothing;
- ⑤ Reproducibility: publish the scripts and prompt sets to GitHub so the interviewer can clone and rerun.
- What's tested: whether you've actually benchmarked or just stitched a number together. See Inference Benchmarking in Practice.
7. How did you land a PR in vLLM? Walk me through the review
- Answer framework: tell three things:
- ① Motivation: what problem in production → what bug or bottleneck you found reading the source → what you changed;
- ② Design: why this change, what alternatives you compared, what benchmark showed how much improvement;
- ③ The review: what the maintainer pushed back on → how you revised → how it finally merged → whether later versions extended your approach.
- Bonus: cite concrete review feedback like "asked to add tests" or "asked to redesign the API" — proof you've actually lived open-source collaboration. One PR merged into mainline is worth five project entries.
8. If you had two months, how would you redo this project?
- Answer framework: show retrospection and prioritization:
- ① start from the original project's biggest weakness (unclear measurement context / no monitoring / no cross-node TP / no speculative decoding) and spend 30% of the time on it;
- ② monitoring and benchmarking first (if the measurement is wrong, everything after is wrong);
- ③ iterate baseline → increments, each with a benchmark and a canary;
- ④ reserve the last 1–2 weeks for experiments in speculative decoding / quantization / edge adaptation.
- The high-scoring signal in an answer: "I would build a complete benchmark and monitoring baseline before talking optimization."
- What's tested: project retrospection + engineering methodology + time planning.
7. System Design & Synthesis Questions (4)
1. Design an LLM inference service that sustains 1000 QPS
- What's tested: synthesis — a common system-design question.
- Reference framework:
- ① Clarify requirements: is 1000 QPS peak or average? Average prompt length? Generation length? SLOs on TTFT / TPOT? Budget?
- ② Capacity estimation: assume 1k-token prompts and 500-token generations — 1500 inference tokens per request. 1000 QPS × 1500 = 1.5M tokens/s. Llama-70B on a single H100 80G delivers ~2000 tokens/s (INT4 weights + continuous batching), needing 750 H100s — clearly unacceptable. So either a smaller model (8B at 10000+ tokens/s per card), more aggressive quantization (INT4 group-wise), or larger batches.
- ③ Architecture:
- Ingress: Nginx / Envoy + rate limiting (route by prompt length);
- Scheduling: a vLLM replica pool + dynamic routing (split by prompt length into long-context and short-context replica groups);
- Inference: per-replica vLLM + chunked prefill + prefix caching; long-context replicas enable chunked prefill, short-context replicas disable it to save scheduling overhead;
- Monitoring: Prometheus + Grafana with five metrics — TTFT / TPOT / queue / GPU utilization;
- Elasticity: K8s HPA + a GPU resource pool;
- Failure: replica health checks + auto-restart + traffic failover.
- ④ Trade-offs: multi-model scheduling (multi-model replicas consume GPU but enable on-demand routing)? speculative decoding (+30–50% throughput but more complex)? cross-node TP (mandatory when a model doesn't fit, otherwise unnecessary)?
- ⑤ Reflection: at 1000 QPS the bottleneck is usually KV cache memory, not compute — solvable with INT8 KV cache quantization + GQA models. That's the key insight interviewers want.
- Bonus points: "I'd do capacity planning before architecture: 1000 QPS is unrealistic on 70B but entirely feasible on 8B. Model choice is itself part of the architecture decision."
- Map to this site: Model Serving and Orchestration, Batching and Request Scheduling, Distributed Inference (TP/PP), Inference Benchmarking in Practice.
2. Design a multi-model inference platform supporting 200+ model versions
- What's tested: system design from the MLOps angle.
- Reference framework:
- ① Model registry: one metadata record per model version (path, version, SLO, hardware requirements, quantization scheme);
- ② Routing: route by model_id to the replica group, then second-order routing by prompt length / user priority;
- ③ Deployment strategies: a) resident replicas (high-frequency models); b) on-demand loading (low-frequency models — vLLM
--load-formatplus a weight cache); c) serverless (rarely called models — high startup cost, lowest cost overall); - ④ Version management: canary rollout (5% of traffic for 30 minutes, full rollout only without regression), A/B testing, rollback mechanisms;
- ⑤ Cost and quota: chargeback per team, per-token quotas, throttling beyond quota.
- This is the core question for ML Infra / MLOps engineers. See Model Serving and Orchestration.
- Map to this site: Model Serving and Orchestration, Triton Inference Server.
3. Why is speculative decoding lossless? Medusa vs EAGLE?
- What's tested: speculative-decoding principles.
- Reference answer:
- Why lossless: a small draft model first proposes k candidate tokens; the large model verifies all k in one forward (verification is cheaper than generation — one forward computes logits for k tokens). Accepted tokens are kept according to the large model's probability distribution; rejected ones are resampled from the rejection point. The key: accept/reject decisions follow the large model's distribution, so the final output distribution is exactly equivalent to direct sampling from the large model → lossless.
- Medusa: attaches multiple "Medusa heads" to the LLM's hidden states, each predicting a next token. Advantage: no separate draft model needed; drawback: lower acceptance rates.
- EAGLE: an autoregressive draft model taking the LLM's hidden state + previous token as input and predicting the next token. Advantage: high acceptance rates (0.6+) and 2–3× end-to-end speedups.
- When not to use: ① low acceptance (draft and target differ a lot); ② extremely long prompts (the draft still runs the full prompt — no prefill speedup); ③ strict-mode businesses (output must follow the target model exactly and draft deviation is unacceptable).
- Bonus points: "Real-world speedup depends on acceptance rate × k. EAGLE on Llama-70B achieves ~0.6 acceptance at k=4, for roughly 2.2× actual speedup. The cost is extra memory and compute (the draft model must also fit) — not worth it when resources are tight."
- Map to this site: Speculative Decoding and Medusa/EAGLE, vLLM, Latency, Throughput, and Concurrency.
4. Tensor Parallel vs Pipeline Parallel — communication costs?
- What's tested: distributed inference — a must.
- Reference answer:
- TP (Tensor Parallel): shards each layer's weights across GPUs; every forward needs an all-reduce. Communication = batch × seq × hidden × 2 bytes (FP16). Cheap over intra-node NVLink (NVSwitch 900GB/s), but 30%+ of step time across nodes (IB 200GB/s).
- PP (Pipeline Parallel): shards layers across GPUs; only activations cross stage boundaries. Communication = batch × seq × hidden × 2 bytes, once per boundary. But PP has the bubble problem — earlier stages idle while later stages finish the micro-batch; bubble ratio = (N-1) / (N+M-1), where N is the number of stages and M the number of micro-batches.
- When to choose: ① a single layer doesn't fit (70B per-layer weights > one card's memory) → TP is mandatory; ② the whole model doesn't fit but each layer does → PP; ③ in practice, TP+PP combined: TP=8 intra-node, PP across nodes.
- MoE inference: uses EP (Expert Parallel) — different experts on different GPUs, tokens routed to the owning expert. EP vs DP: DP holds a full model replica per GPU; EP holds only some experts per GPU.
- Bonus points: "Cross-node TP is slow not because of NVLink-vs-IB bandwidth (NVSwitch 900 vs IB 200), but because NCCL falls back to P2P instead of RDMA across nodes unless you explicitly set
NCCL_NET_GDR_LEVEL. That's a detail you only know after stepping in it." - Map to this site: Distributed Inference (TP/PP), Hardware Primer.
8. Live Coding Questions (5)
Live coding for this role splits into two families: generic algorithms (written tests / online assessments — see the prep plan) and inference-related handwriting (on-site Q&A plus core code). Below are the high-frequency list and reference implementations for the latter.
1. Implement a simple KV cache class
- What's tested: basic data structures for inference.
- Reference implementation:
python
import torch
class KVCache:
def __init__(self, num_layers, num_kv_heads, head_dim, max_batch, max_seq, dtype=torch.float16, device="cuda"):
self.num_layers = num_layers
self.num_kv_heads = num_kv_heads
self.head_dim = head_dim
self.max_batch = max_batch
self.max_seq = max_seq
# Pre-allocate (num_layers, max_batch, max_seq, num_kv_heads, head_dim)
self.k = torch.zeros(num_layers, max_batch, max_seq, num_kv_heads, head_dim, dtype=dtype, device=device)
self.v = torch.zeros_like(self.k)
# Current written length per batch slot
self.lengths = [0] * max_batch
def append(self, layer_idx, batch_idx, new_k, new_v):
# new_k, new_v: (new_seq, num_kv_heads, head_dim)
start = self.lengths[batch_idx]
end = start + new_k.shape[0]
self.k[layer_idx, batch_idx, start:end] = new_k
self.v[layer_idx, batch_idx, start:end] = new_v
self.lengths[batch_idx] = end
def get(self, layer_idx, batch_idx):
return self.k[layer_idx, batch_idx, :self.lengths[batch_idx]], \
self.v[layer_idx, batch_idx, :self.lengths[batch_idx]]- Follow-ups: "What's wrong with this implementation?" — ① wastes memory (pre-allocates max_seq); production uses PagedAttention-style paging; ② batch dimension can't dynamically grow or shrink (continuous batching needs to swap requests in and out); ③ no prefix sharing.
- Map to this site: Batching and Request Scheduling, vLLM.
2. PagedAttention draft: block table and block pool
- What's tested: understanding of the PagedAttention mechanism.
- Reference implementation:
python
import torch
class PagedKVCache:
def __init__(self, num_layers, num_kv_heads, head_dim, block_size, num_blocks, dtype=torch.float16, device="cuda"):
self.block_size = block_size # e.g., 16 tokens per block
self.num_blocks = num_blocks
# Physical storage: (num_layers, num_blocks, block_size, num_kv_heads, head_dim)
self.k_blocks = torch.zeros(num_layers, num_blocks, block_size, num_kv_heads, head_dim, dtype=dtype, device=device)
self.v_blocks = torch.zeros_like(self.k_blocks)
# Block pool: ids of free physical blocks
self.free_blocks = list(range(num_blocks))
# Per-request logical -> physical block mapping
self.block_tables = {} # request_id -> list of physical block ids
def allocate(self, request_id, num_tokens_needed):
# Allocate blocks on demand
n_blocks_needed = (num_tokens_needed + self.block_size - 1) // self.block_size
if n_blocks_needed > len(self.free_blocks):
raise RuntimeError("OOM: no free blocks")
blocks = [self.free_blocks.pop() for _ in range(n_blocks_needed)]
self.block_tables[request_id] = blocks
def free(self, request_id):
# Request finished: return blocks
for b in self.block_tables[request_id]:
self.free_blocks.append(b)
del self.block_tables[request_id]- Follow-ups: "How does this differ from real vLLM?" — ① the real implementation is C++ with a custom allocator; ② copy-on-write support for beam search; ③ prefix sharing (identical prefixes share blocks, copy-on-write on modification). See vLLM.
- Map to this site: Batching and Request Scheduling, vLLM, The GPU Memory Hierarchy and the Bandwidth Wall.
3. Hand-write a reduction kernel (CUDA draft)
- What's tested: CUDA fundamentals.
- Reference implementation:
cuda
__global__ void reduce_sum(const float* input, float* output, int n) {
__shared__ float sdata[256];
int tid = threadIdx.x;
int i = blockIdx.x * blockDim.x + tid;
// Each thread grabs one element (a more general version uses a strided loop)
sdata[tid] = (i < n) ? input[i] : 0.0f;
__syncthreads();
// Tree reduction
for (int s = blockDim.x / 2; s > 0; s >>= 1) {
if (tid < s) {
sdata[tid] += sdata[tid + s];
}
__syncthreads();
}
if (tid == 0) {
atomicAdd(output, sdata[0]);
}
}- Follow-ups: "Where are the bank conflicts in this version?" — the shared-memory accesses
sdata[tid]andsdata[tid + s]: at s=128 the stride spans 128 floats = 512 bytes = 128 banks, no conflict; at s=1 accesses are adjacent, no conflict. But certain strides do conflict; a better version uses warp shuffle (__shfl_down_sync). - Map to this site: GPU Architecture and Optimization, Kernel Fusion and Custom Kernels.
4. A simple fused matmul + bias + relu kernel in Triton
- What's tested: Triton DSL basics.
- Reference implementation:
python
import triton
import triton.language as tl
@triton.jit
def matmul_bias_relu_kernel(
a_ptr, b_ptr, bias_ptr, c_ptr,
M, N, K,
stride_am, stride_ak,
stride_bk, stride_bn,
stride_cm, stride_cn,
BLOCK_M: tl.constexpr, BLOCK_N: tl.constexpr, BLOCK_K: tl.constexpr,
):
pid_m = tl.program_id(0)
pid_n = tl.program_id(1)
offs_m = pid_m * BLOCK_M + tl.arange(0, BLOCK_M)
offs_n = pid_n * BLOCK_N + tl.arange(0, BLOCK_N)
offs_k = tl.arange(0, BLOCK_K)
acc = tl.zeros((BLOCK_M, BLOCK_N), dtype=tl.float32)
for k in range(0, tl.cdiv(K, BLOCK_K)):
a = tl.load(a_ptr + offs_m[:, None] * stride_am + offs_k[None, :] * stride_ak,
mask=(offs_m[:, None] < M) & (offs_k[None, :] < K + k * BLOCK_K), other=0.0)
b = tl.load(b_ptr + offs_k[:, None] * stride_bk + offs_n[None, :] * stride_bn,
mask=(offs_k[:, None] < K) & (offs_n[None, :] < N), other=0.0)
acc += tl.dot(a, b)
bias = tl.load(bias_ptr + offs_n, mask=offs_n < N, other=0.0)
acc += bias[None, :]
acc = tl.maximum(acc, 0.0) # ReLU
tl.store(c_ptr + offs_m[:, None] * stride_cm + offs_n[None, :] * stride_cn,
acc, mask=(offs_m[:, None] < M) & (offs_n[None, :] < N))- Follow-ups: "How do you pick the tile sizes?" — BLOCK_M / BLOCK_N / BLOCK_K depend on SRAM capacity; 128×128×32 is common on A100; too large overflows SRAM, too small wastes parallelism.
- Map to this site: Kernel Fusion and Custom Kernels, GPU Architecture and Optimization.
5. Practice scope for generic-algorithm live coding
- What's tested: whether you can "still crank through problems" — the high-frequency scope matters more than full coverage.
- Reference list:
- Binary search and variants: first element ≥ target in a sorted array; peak finding;
- Sorting: quicksort, mergesort, heapsort;
- DP: longest increasing subsequence, knapsack, edit distance;
- Graphs: BFS / DFS, topological sort, Dijkstra;
- Linked lists: reverse, cycle detection, merging sorted lists;
- Stacks / queues: monotonic stack, sliding-window maximum.
- Practice order: a working brute-force solution first, then optimize to the best complexity the interviewer expects; state time / space complexity for every problem — the most important hidden scoring item. Common pitfalls and coding hygiene live in Common Pitfalls and Anti-Patterns.
9. How to Answer "I Don't Know"
Getting stuck is not scary — answering ugly is. The interviewer is also watching your composure under pressure; that is itself a soft-skills question. Principles and scripts:
- Principle 1: never bluff. Interviewers ask deeper than you can fake; fabrication gets punctured by follow-ups, the score collapses, and you lose the "honest" impression too. The right posture: state the boundary clearly + show your thinking path.
- Principle 2: respond in layers. Not every "don't know" deserves instant surrender — judge the layer: never heard of it (outside your field) → heard of it but can't derive it (missing principles) → know part of it (fuzzy memory).
- Principle 3: turn "I don't know" into evidence that you can learn.
Three script templates:
Template 1 (completely unfamiliar — the honest route):
"I haven't systematically studied this concept. That said, from my understanding it may be
related to X (adjacent knowledge). If it's solving problem Y (reasoning backward from the
problem), my initial approach would be... I'll go study this direction after the interview."
Template 2 (heard of it, can't derive it):
"I've seen this in papers/blogs — I can say roughly what problem it solves (a sentence or
two), but I can't reconstruct the derivation in detail right now. If possible, could you
hint at its assumptions? Or let me describe my understanding with an example — I believe
it's related to A and B..."
Template 3 (fuzzy memory — guided):
"I remember its relationship to PagedAttention, but I'm not sure about the block table
indexing details. Let me first state what I'm sure about: ... For the uncertain part I'd
rather say so than mislead you. If you give me a concrete scenario, I can reason through
its mechanism on the spot."Three accompanying bonus moves:
- Show the learning path proactively: "I've been reading the FlashAttention-3 source lately — I didn't answer the TMA details well today, but I can explain that direction clearly next month" — turns "can't" into "in progress."
- Fall back on analogy: when you don't know a term, "my understanding is it does something like ..., similar to ..." lets the interviewer see your analogical reasoning.
- Counter with a question: in an unfamiliar area, ask one specific, non-offensive question ("Is this mechanism implemented with vLLM or TensorRT-LLM in your inference service?") — buys time and shows engagement. Don't ask freebie questions like "can you give me a hint?"
One answer you must avoid
"I haven't learned this — can we skip it?" — burns the question's score and exposes weak learning intent. The right posture is always: acknowledge the boundary + give a partial understanding + commit to filling the gap. Interviewers don't expect you to know everything; they expect you to "explain what you know thoroughly, and not lose points on what you don't."
10. Interview Prep Plan (1-Week and 1-Month Checklists)
Designed on the start-with-the-end-in-mind principle from Learning Paths: Three Routes: first figure out what your target role's JD tests, then check yourself against the list. Both versions assume you've completed the self-assessment in Deconstructing JD Knowledge Points.
1-Month Version (standard pace, 2–3 hours/day)
| Week | Main line | Fixed 30-minute daily action |
|---|---|---|
| Week 1 | L1 fundamentals + metrics: Python GIL / C++ RAII / Nsight / TTFT-TPOT — with Latency, Throughput, and Concurrency, GPU Architecture and Optimization | One generic-algorithm coding problem |
| Week 2 | L2 engines + quantization: vLLM source deep reading / PagedAttention / continuous batching / GPTQ vs AWQ — with vLLM, Model Quantization Fundamentals | One concept question answered aloud (recorded) |
| Week 3 | L3 kernels + distributed: FlashAttention / CUDA reduction / TP-PP-EP — with Kernel Fusion and Custom Kernels, Distributed Inference (TP/PP) | Daily self-question: "why did I make this decision in my project?" |
| Week 4 | Full simulations: rehearse the "I don't know" scripts from Section 9; book 2 mock interviews; build an error log | Review the previous three weeks' errors and recordings |
1-Week Version (emergency sprint, 4–6 hours/day)
Days 1–2 Rapid sweep of fundamentals: answer the 4 metrics + 4 quantization questions
(Section 2) yourself before reading answers; fix gaps same day; pair with the
latency-throughput and quantization pages
Day 3 Engines & scheduling: read vLLM Scheduler / PagedAttention source closely,
derive the KV cache size formula by hand, write the PagedAttention draft
once (Section 8)
Day 4 Kernels + distributed: FlashAttention principles / hand-written CUDA reduction /
TP-PP communication costs
Day 5 Project deep-dive: STAR-rewrite resume projects; rehearse questions 1, 3, 5, 6
from Section 6; memorize the "I don't know" scripts
Day 6 Two mock interviews (friends or recorded video), focusing on self-intro and
project narration
Day 7 Error-log review + light retrospective; learn nothing newThree suggestions that run through the whole plan
- Record and re-listen to your spoken answers — 90% of "felt great at the time" answers reveal broken logic on playback; this is the cheapest improvement tool available.
- Build a "one-sentence definition" list: every high-frequency concept expressible in one sentence + one number (e.g., "TTFT is the prefill-phase latency; Llama-70B on H100 with a 2k prompt is roughly 600ms") — far more useful than reciting paragraphs.
- Keep terminology consistent: use canonical terms in answers (PagedAttention, continuous batching, TP, EP, TTFT, TPOT), and calibrate any time against the Glossary — terminology is the interviewer's hidden signal for "trained professional."
11. Further Reading
- Concept catch-up: Latency, Throughput, and Concurrency | The GPU Memory Hierarchy and the Bandwidth Wall | The Roofline Model and Compute Analysis | GPU Architecture and Optimization | Hardware Primer
- Engines and kernels: vLLM | TensorRT-LLM | Triton Inference Server | Kernel Fusion and Custom Kernels | Computation Graph Optimization
- Quantization and compression: Model Quantization Fundamentals | Weight-Only Quantization and Mixed Precision | Pruning and Sparsification | Knowledge Distillation
- Scheduling and distributed: Batching and Request Scheduling | Model Serving and Orchestration | Distributed Inference (TP/PP) | Speculative Decoding and Medusa/EAGLE
- Edge inference: llama.cpp and GGUF | Mobile Deployment
- Practice: Inference Benchmarking in Practice | Tuning and Performance Optimization | Inference Engine Comparison | Common Pitfalls and Anti-Patterns
- Career companions: career module guide | JD List | Deconstructing JD Knowledge Points | Skills Benchmarking | Learning Paths: Three Routes
- Practice platforms: LeetCode (generic algorithms); vLLM / SGLang / llama.cpp source (reference implementations for inference-related handwriting); Hugging Face courses (LLM track)
References
- Kwon et al. Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP 2023) — the original vLLM/PagedAttention paper
- Dao. FlashAttention-2: Fast Attention with Better Parallelism and Work Partitioning (2023) — the FlashAttention-2 paper
- Shah et al. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision (2024) — FlashAttention-3 / H100 TMA / FP8
- Frantar et al. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (ICLR 2023) — GPTQ quantization
- Lin et al. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (MLSys 2024) — AWQ quantization
- Xiao et al. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models (ICML 2023) — SmoothQuant activation quantization
- Cai et al. Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads (ICML 2024) — Medusa speculative decoding
- Li et al. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty (ICML 2024) — EAGLE speculative decoding
- Shoeybi et al. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism (2019) — the classic TP/PP paper
- NVIDIA. H100 Transformer Engine Technical Brief (2022) — official H100 FP8 / TMA overview
- NVIDIA. Nsight Systems Documentation — official docs for the performance-analysis toolchain
- NVIDIA. CUDA C++ Programming Guide — the official textbook for bank conflicts / reduction kernels
- OpenAI Triton Language Tutorials — getting started with the Triton DSL
- vLLM Project Source Code — source-reading material for engines
- TensorRT-LLM Documentation — in-flight batching / FP8 implementation reference
- LeetCode Problem Set — the standard practice platform for generic-algorithm live coding