Appearance
Model Deployment Interview Question Bank (80+ Questions)
Organized into eight modules, each question comes with reference answer points (3-5 bullets) and follow-up directions. The goal is not to memorize answers but to build an answer skeleton of "concept → scenario → trade-off": open with a one-sentence definition, explain the mechanism or flow, and close with the engineering trade-offs.
How to Use This Question Bank
First, self-assess against the JD knowledge map to locate your weak spots. Second, work through each module answering out loud; when you can't answer, go back to the relevant concept page to fill the gap. Third, run mock interviews: spend 2-3 minutes per question, and you've only passed once you can handle the follow-ups. Use this bank alongside the job-hunt sprint track in the learning paths.
Module 1: Inference Fundamentals (11 Questions)
1. What is the difference between training and inference?
- Training runs forward + backward passes; inference only runs the forward pass. Training updates weights; inference uses fixed weights
- Training prioritizes throughput (large batches); inference prioritizes latency (small batches). Training uses FP32/AMP to preserve gradients; inference can use FP16/INT8 quantization
- Training is compute-bound; inference is bandwidth-bound (very little computation per weight read)
- Inference demands real engineering: model loading, serving, monitoring, capacity planning
Follow-up: Why is inference latency-sensitive? What batch size is typical for inference?
2. What computations happen in a single forward pass?
- CNN: convolution → activation → pooling → fully connected, a layer-by-layer tensor transformation
- Each layer has three tensors—input, weights, output—and intermediate activations occupy memory
- Operators can be fused (conv+bn+relu) to reduce memory traffic and kernel launches
- Transformer: embedding → self-attention → FFN → LayerNorm, plus the KV cache
Follow-up: Why do intermediate activations often take more memory than the weights?
3. What makes up inference latency?
- Model loading and graph compilation (one-time cost)
- Request queuing (batch aggregation, rate limiting)
- Preprocessing (decoding, tokenization, resizing) and postprocessing (softmax, NMS)
- Operator compute time (the core part, usually broken down via profiling)
- Network transfer (serialization, RTT)
Follow-up: How do you tell whether the bottleneck is compute or bandwidth? How would you break it down with Nsight or torch.profiler?
4. How do you estimate model size? Why is GPU memory the critical resource?
- Weight size = parameter count × bytes per parameter: 7B at FP16 ≈ 14 GB, FP32 ≈ 28 GB
- Inference memory must also hold intermediate activations, the runtime, and GPU context
- LLMs add the KV cache: 2 × num_layers × hidden × batch × seq_len × bytes
- If it doesn't fit in memory, switch to quantization (quantization) or go distributed
Follow-up: How many A100-80G cards does a 70B FP16 model need? (Four—weights alone are 140 GB.)
5. CPU vs. GPU inference: what are the trade-offs?
- CPU: low per-request latency, no VRAM limit, simple deployment—good for small models or low-concurrency online serving
- GPU: massive parallelism and large batches—good for large models and high throughput
- The bottlenecks differ: CPUs are limited by memory bandwidth; GPUs by VRAM bandwidth and kernel launch overhead
- Edge/mobile often goes CPU + INT8 (see the TensorRT edge case study)
Follow-up: When can a GPU actually be slower than a CPU? (batch=1 on a small model, kernel-launch overhead dominating)
6. What do latency, throughput, and QPS mean? How are they related?
- Latency: time to complete a single request (P50/P99). Throughput: requests completed per unit of time
- QPS (RPS) = throughput; throughput = concurrency / average latency (an approximation of Little's Law)
- The usual levers for raising throughput (batching) sacrifice latency—a genuine trade-off
- Capacity planning should use "max QPS at the target latency"
Follow-up: Why can P99 be far higher than P50? (Long tail, GC pauses, resource contention)
7. What is the memory bandwidth bottleneck? Why is inference bandwidth-bound?
- Arithmetic intensity = total compute / total memory traffic (FLOPs per byte)
- Inference does only a few FLOPs per weight read, so arithmetic intensity is low → bandwidth-bound
- In bandwidth-bound scenarios, reducing bytes (quantization, operator fusion) pays off the most
- GPU VRAM bandwidth >> CPU memory bandwidth—one source of the GPU's inference advantage
Follow-up: Why does quantization help most in bandwidth-bound scenarios?
8. What is a batch? Why is dynamic batching useful?
- A batch processes multiple samples in one forward pass, amortizing kernel launches and weight reads
- GPU utilization rises with batch size, but so does latency
- Dynamic batching aggregates queued requests (time window / max batch size); see serving
- The batch ceiling is set by memory (activations + KV cache) and compute
Follow-up: What parameters does dynamic batching expose (window, max_batch)? How do you set the trade-off?
9. Why do inference services need warm-up?
- The first inference triggers graph compilation, operator selection, memory allocation, and cache misses—noticeably slow
- Only after warm-up does performance enter its steady state; otherwise load-test and production numbers are skewed
- In Kubernetes, use a readiness probe to hold back traffic until warm-up completes
- CUDA Graph capture and TensorRT engine building are part of warm-up too
Follow-up: What happens if you ship without warming up? (Timeout alerts, misjudged capacity)
10. How do you validate model accuracy at deployment time?
- Use an evaluation set drawn from the same distribution as training and compare the baseline (training-side metrics) with deployment-side metrics
- Check preprocessing consistency (scaling, normalization, tokenizer)—the most common source of accuracy drops
- Run an accuracy regression after quantization on the full evaluation set, not a sample
- After launch, keep validating with online metrics (CTR/accuracy); see the rollout process
Follow-up: Production results are 5% worse than offline—what do you check first?
11. What is a computation graph? Dynamic vs. static graphs?
- A computation graph is a dataflow graph: nodes are operators, edges are tensors
- Dynamic graphs (eager) execute operator by operator—flexible, easy to debug. Static graphs build first and execute later, enabling global optimization
- Static graphs enable operator fusion, constant folding, and memory reuse; ONNX and TensorRT both take the static-graph route
- The trend:
torch.compile/ JIT to get the best of both
Follow-up: Why is ONNX a static graph? Why are dynamic shapes hard to handle?
Module 2: Inference Engines (11 Questions)
12. What is ONNX? Why convert models to it?
- ONNX is an open neural network exchange format—an intermediate representation that works across frameworks (PyTorch/TF)
- It standardizes the operator set and graph representation so backends (ORT/TensorRT) only have to optimize for one format
- The ecosystem is mature: conversion tools, quantization, and visualization are all there (see model formats)
Follow-up: Common pitfalls when converting PyTorch checkpoints to ONNX? (Dynamic shapes, unsupported operators, control flow)
13. What is ONNX Runtime's optimization approach?
- Graph optimization: operator fusion, constant folding, redundant-node elimination
- Execution Providers (CPU/GPU/DirectML/OpenVINO) map the graph onto different hardware
- Supports quantized models (QDQ format) and INT8 execution
- Memory planning: static analysis to reuse activation buffers
Follow-up: Why is ORT faster than PyTorch eager mode? (Graph optimization plus kernel selection)
14. What is TensorRT's full optimization pipeline?
- Pipeline: parse the model → build the engine (operator selection/fusion/precision) → serialize → run inference
- INT8 requires a calibration set to determine quantization scales
- Engines are tightly bound to the GPU model, driver, and TensorRT version
- Supports CUDA Graph capture and multi-stream concurrency
Follow-up: Can a serialized engine be used directly on a different GPU? (No—it must be rebuilt.)
15. What is operator fusion?
- Combine adjacent operators into a single kernel to cut kernel launches and intermediate tensor reads/writes
- Typical examples: conv+bn+relu, GELU folded into attention, residual + normalization fusion
- Fusion pays off most in bandwidth-bound scenarios
- TensorRT, ORT, and vLLM all fuse at different granularities
Follow-up: Can fusion hurt accuracy? What numerical issues should you watch for?
16. How does TensorRT INT8 calibration work?
- Run a representative calibration set (a few hundred to a few thousand samples) to collect per-layer activation distributions
- Choose a quantization scheme: min-max / entropy calibration / percentile
- Produce per-layer scales (possibly per-channel) and build the INT8 engine
- If the calibration set doesn't match the real inference distribution, accuracy collapses—verify manually
Follow-up: How do you pick a calibration set? What happens with too few samples?
17. What are Triton Inference Server's core features?
- Serves models from multiple frameworks through one interface (ONNX/PyTorch/TensorRT/TF)
- Dynamic batching + concurrent model instances (multi-stream parallel execution)
- Model ensembles / business logic (script-based) orchestration
- gRPC/HTTP endpoints plus performance analysis tooling (Model Analyzer)
Follow-up: How do Triton and vLLM differ in positioning? (General-purpose model serving vs. LLM-specific)
18. What are vLLM's core mechanisms?
- PagedAttention: page-based KV cache management that eliminates memory fragmentation and waste
- Continuous batching: token-level scheduling; a finished sequence frees its slot immediately
- Native support for loading quantized weights (GPTQ/AWQ/FP8) and tensor parallelism
- Block-table-based management lets memory be reused for long contexts
Follow-up: What does PagedAttention save over a traditional KV cache? (See LLM inference)
19. How do you choose an inference engine?
- Start from the framework: PyTorch models go ONNX Runtime / TorchServe first; for peak performance pick TensorRT
- Then hardware: in the NVIDIA ecosystem use TensorRT/Triton; on Intel use OpenVINO; for general cases ORT
- Then scenario: online LLM serving picks vLLM/SGLang; mixed multi-framework fleets pick Triton; edge picks TensorRT-lite/TFLite
- Score candidates on the framework comparison dimensions: performance, ecosystem, ease of use, team experience
Follow-up: You're handed a CV model that must serve 1,000 QPS in production. What do you pick, and why?
20. Eager mode vs. graph mode? What is a CUDA Graph?
- Eager: one kernel launch per operator—flexible but slow. Graph: capture the whole sequence, then submit it once
- CUDA Graph captures a chain of kernels as a graph; replaying it avoids launch overhead, ideal for fixed shapes
- With dynamic shapes, CUDA Graphs need tricks like graph pools and segmented capture
torch.compileperforms graph optimization + kernel fusion
Follow-up: Why do CUDA Graphs cut small-batch inference latency so dramatically?
21. What is GGUF? Why is llama.cpp popular?
- GGUF is llama.cpp's model format, with multiple built-in quantization levels (Q4_K_M and friends)
- Weights are stored pre-quantized: load and run, no runtime calibration
- The universal key to running LLMs on CPU and Apple Silicon—common for edge and local deployment
- Versus safetensors: GGUF also carries quantization parameters and tokenizer info
Follow-up: In GGUF's Q4_K_M, what does K_M stand for? (Higher precision for the key part)
22. How do you optimize slow model loading?
- Preload and keep models resident (memory/VRAM pool) so requests never trigger a load
- Put model files on local NVMe to avoid network/shared-storage reads
- Spread loading across replicas; auto-retry failed loads and roll back to the previous version
- In Kubernetes, use readiness to control when traffic arrives
Follow-up: During a hot model update, how do you guarantee zero dropped requests?
Module 3: Quantization and Compression (10 Questions)
23. What is the basic principle of quantization?
- Map floating-point values into an integer range using scale + zero_point
- Symmetric quantization (zero_point=0, typical for weights) vs. asymmetric quantization (typical for activations)
- FP32 → INT8 cuts size by 75% and lowers bandwidth/compute demand
- Quantization works for inference, not training: gradients need high precision. See quantization
Follow-up: Per-tensor vs. per-channel quantization—what's the difference?
24. PTQ vs. QAT: how do you choose?
- PTQ: post-training quantization with a calibration set—fast, and good enough for most INT8 cases
- QAT: quantization simulated during training (fake quant)—higher accuracy, much higher cost
- Prefer QAT for accuracy-sensitive tasks (detection/segmentation) or low bit-widths (INT4)
- Hybrid strategy: keep sensitive layers in high precision and quantize the rest
Follow-up: Accuracy collapsed after PTQ—what do you do? (Calibration set / mixed precision / switch to QAT)
25. How do you evaluate and mitigate INT8 accuracy issues?
- Compare INT8 vs. FP16 metrics on the full evaluation set instead of eyeballing a few samples
- Mitigations: better calibration data, per-channel scales, skipping sensitive layers, mixed precision
- Find sensitive layers with per-layer error analysis (sensitive layer analysis)
- Activation outliers are the main accuracy killer; clipping helps
Follow-up: How do you find the most sensitive layer?
26. What are the differences between FP16, BF16, and FP8?
- FP16: 1+5+10 bits—narrow range, overflows easily. BF16: 1+8+7 bits—FP32's range with less precision
- BF16 is training-friendly (wide dynamic range); FP16 is more mature for inference
- FP8 (E4M3/E5M2): the new-generation inference format—INT8-sized with better accuracy
- Choice depends on hardware support (BF16 on Ampere and later; FP8 on Hopper/Blackwell)
Follow-up: Why does LLM training standardize on BF16 rather than FP16?
27. GPTQ vs. AWQ: what's the difference?
- Both are 4-bit weight quantization methods aimed at LLMs
- GPTQ: minimizes quantization error layer by layer using second-order (Hessian) information, in one global calibration pass
- AWQ: activation-aware—protects important weight channels with large activations via per-channel scaling
- In practice the two are close in accuracy; AWQ is faster and simpler to implement
Follow-up: Why is 4-bit weights + 16-bit activations (W4A16) the mainstream combination?
28. Weight quantization vs. activation quantization?
- Weight-only quantization (W8A16): quantize weights, keep activations in FP16—simple to implement, broadly compatible
- Activation quantization (W8A8): requires tracking activation dynamic ranges; wide activation distributions make accuracy hard to hold
- Only with activation quantization do you get the full INT8 compute speedup (int8 GEMM)
- Common LLM choices: W8A8 (e.g., SmoothQuant) or W4A16
Follow-up: Why is activation quantization hard? (The dynamic range differs token by token)
29. Distillation, pruning, sparsification, quantization—how do they differ?
| Technique | How it works | Benefit | Cost |
|---|---|---|---|
| Quantization | Low-bit numeric representation | Size / bandwidth | Low (PTQ) |
| Pruning | Remove unimportant weights/channels | Sparsified size | Medium |
| Distillation | A large model teaches a small one | Smaller architecture | High (requires retraining) |
| Sparsification | Zero out + sparse storage | Size / compute (when hardware supports it) | Medium |
Suggested order: prune/distill first, then quantize; see model compression.
Follow-up: Why does sparsification underdeliver on GPUs? (Limited hardware support for sparse acceleration)
30. Why does quantization speed things up? Does it still help when compute-bound?
- Bandwidth-bound: halving weight bytes halves the transfer time
- Compute-bound: INT8 can use tensor cores (2-4x FLOPS), so there's still a win
- Smaller size also means better cache hit rates and room for bigger batches
- For very small models, quantization overhead can cancel the gains—measure it
Follow-up: How do you tell whether an inference workload is compute-bound or bandwidth-bound?
31. How do you choose a quantization bit-width?
- Work backward from the memory budget: the lowest bit-width that fits with headroom wins—don't chase the minimum blindly
- By accuracy needs: stay conservative with INT8 or mixed precision for sensitive tasks (detection/segmentation/medical)
- By hardware: check format support (FP8 needs Hopper or newer)
- Balance the accuracy-size-speed triangle and let measurements decide; see quantization
Follow-up: How would you configure a 7B model to fit on a single 16 GB card?
32. What is SmoothQuant, and what problem does it solve?
- It fixes the accuracy problem of activation quantization by "smoothing" the difficulty from activations onto weights
- Weights are scaled per channel so the activation distribution becomes more uniform
- It's the classic representative of W8A8 schemes and makes models quantization-friendly
- It needs a small calibration set to gather activation statistics
Follow-up: What idea does SmoothQuant share with AWQ? (Both use channel scaling)
Module 4: Serving (10 Questions)
33. REST vs. gRPC for inference services: what are the trade-offs?
- REST: HTTP/JSON—easy to debug, broad ecosystem, good for external APIs
- gRPC: protobuf + HTTP/2—multiplexing, low latency, strong typing, good for internal services
- Internal hops between services usually use gRPC; public endpoints usually expose REST
- Triton offers both; prefer gRPC for performance-sensitive paths
Follow-up: What HTTP/1.1 problem does HTTP/2 multiplexing solve?
34. How does dynamic batching work?
- Requests queue up and wait for either the time window to elapse or max_batch to fill
- The aggregated batch runs one forward pass, amortizing kernel launches; results are split back into individual responses afterward
- Parameters: max_batch_size, delay (wait duration), queue policy
- Trade-off: a larger window means higher throughput and higher latency
Follow-up: How do you find the optimal window time? (Load-test sweep + a latency budget)
35. How do you design model loading and hot swaps for an inference service?
- Store model versions in a model repository and load them into a VRAM pool
- Hot swap: new version ready → switch routing → release the old version (no downtime)
- Auto-rollback on load failure; no traffic until readiness passes
- Rolling updates across replicas with a "one version at a time" policy
Follow-up: The hot load fails because VRAM ran out—what's your fallback plan?
36. How do you implement rate limiting?
- Token bucket / leaky bucket algorithms, keyed on QPS or concurrency
- Concurrency caps (semaphores) + queues + drop on timeout
- Return 429 with Retry-After so clients back off and retry
- Rate limiting and backpressure: propagate load upstream when the service is overwhelmed
Follow-up: How do you set the limit? (From the saturation point found in load testing; see load testing)
37. What threading and process model should an inference service use?
- Python is GIL-bound: CPU-heavy inference belongs in separate processes or the native layer
- In FastAPI, run inference in a thread pool or a separate process so the event loop never blocks
- Keep model replicas (processes) resident and shard requests by concurrency
- Triton uses model instances + streams for parallelism
Follow-up: Why can't a Python inference service scale its QPS? (GIL, interpreter overhead)
38. How do you scale a model service horizontally?
- Keep the service stateless (state lives outside) and scale replicas with traffic
- GPU services are bound by VRAM and card count: each card holds N model replicas
- Autoscale with the Kubernetes HPA on custom metrics (QPS, GPU utilization)
- Load balancing (Ingress/Service) distributes traffic across replicas
Follow-up: What HPA pitfalls affect GPU inference services? (Covered in Module 5 below)
39. What is a model instance?
- A Triton concept: multiple instances of the same model, each with its own execution thread/stream
- When instances exceed what the CPU/GPU can execute concurrently, they queue
- It pairs with dynamic batching: batching within an instance, parallelism across instances
- Concurrency is capped by memory and compute resources
Follow-up: What happens if you set too many instances? (Queuing, VRAM exhaustion)
40. How do you implement graceful shutdown?
- Stop accepting new requests → drain in-flight requests → release resources
- Kubernetes preStop hook + terminationGracePeriodSeconds
- Close connection pools, flush state to disk, log the final batch
- A hard kill turns in-flight requests into 5xx errors
Follow-up: How do readiness probes and graceful shutdown work together?
41. How do you package a model into a deployable artifact?
- Docker image: base image (matching CUDA version) + inference engine + model files
- Mount model files externally (images get bloated and slow to update); put code in the image
- Pre-build ORT/TensorRT engine artifacts to avoid runtime compilation
- Use multi-stage builds to shrink the image and pin dependency versions
Follow-up: Model baked into the image vs. mounted volume—how do you decide?
42. How do you handle API versioning and compatibility?
- Version via path or header (/v1/...) and avoid breaking changes
- Keep request/response schemas compatible: new fields are optional
- Run versions side by side for gradual rollout so clients upgrade smoothly
- Track callers and communicate before changes; see the rollout process
Follow-up: How do you make a model upgrade (changing input format) seamless?
Module 5: Kubernetes and Containers (11 Questions)
43. What is the Pod lifecycle?
- Pending → Running → Succeeded/Failed; on trouble, ContainerCreating/CrashLoopBackOff
- initContainers run before the main container (environment setup)
- On deletion: preStop runs first, then SIGTERM, then SIGKILL after the grace period
- Controllers (Deployment/StatefulSet) handle replica repair
Follow-up: What are the common causes of CrashLoopBackOff, and how do you debug it?
44. What's the difference between liveness, readiness, and startup probes?
- liveness: is the container alive? Failure triggers a restart
- readiness: can it take traffic? Failure removes it from rotation (no restart)
- startup: designed for slow-starting containers, protecting liveness from false kills
- For inference services: startup covers warm-up, readiness gates traffic, liveness probes the process
Follow-up: Why not use liveness to probe "business health" in an inference service? (Risk of false restarts)
45. How does the HPA work?
- The control plane periodically reads metrics (metrics-server) and computes the desired replica count
- Desired replicas = ceil(current usage / target usage)
- Supports custom and external metrics (QPS, GPU utilization)
- Cooldown windows (scaling down slower than up) prevent flapping
Follow-up: HPA pitfalls for GPU inference services? (VRAM never shrinks back, new Pods need warm-up, cold-start storms)
46. How are GPU resources scheduled in Kubernetes?
- The NVIDIA device plugin exposes GPUs as an extended resource (nvidia.com/gpu)
- Memory granularity: one card is one unit—if it doesn't fit on a single card, it won't get scheduled
- Request GPU count rather than VRAM (VRAM-level requests need MIG or a custom scheduler)
- Node labels + taints/tolerations pin GPU workloads to GPU nodes
Follow-up: On an 8-GPU machine, how do you stop eight small jobs from hogging every card so big jobs can't schedule?
47. How do you set requests and limits?
- requests drive scheduling (a reservation); limits are the hard ceiling (CPU throttling, memory OOM)
- GPUs only have limit semantics (no oversubscription)
- Memory limits too low → OOMKilled restarts; too high → waste
- Set them from load-test data: watch actual usage and leave 1.2-1.5x headroom
Follow-up: What happens after a container OOMs? How do you find who consumed the memory?
48. Deployment vs. StatefulSet?
- Deployment: stateless, interchangeable, rolling updates. StatefulSet: stable identity, ordered operations, persistent volumes
- Model inference is usually stateless → Deployment
- Use StatefulSet when you need a fixed network identity or storage (e.g., inference caches, distributed engines)
Follow-up: Give an example where an online inference service would need a StatefulSet.
49. How do you use ConfigMaps and Secrets?
- ConfigMaps hold non-sensitive configuration (timeouts, batching parameters)
- Secrets hold sensitive data (API keys, model encryption keys)
- Changes require a restart or rolling update to take effect (no automatic hot-reload)
- Never bake secrets into images or commit them to repositories
Follow-up: Where should model paths and timeout parameters live?
50. How do you build and optimize container images?
- Multi-stage builds: separate the build environment from the runtime for smaller images
- Match the base image to your CUDA/driver versions to avoid missing runtime libraries
- Exploit layer caching by putting stable layers first; use .dockerignore to exclude large files
- Inference engines are huge: pre-build engine artifacts instead of compiling at runtime
Follow-up: The image is too large and pulls too slowly—how do you optimize? (Layering, external models, private registry)
51. What's the traffic path through Service and Ingress?
- Service: ClusterIP for in-cluster load balancing; NodePort exposes a node port
- Ingress: layer-7 routing (host/path) to Services, with TLS and rate limiting
- Request path: Client → Ingress → Service → Pod
- Session affinity and traffic splitting (canary) can be handled at the Ingress layer
Follow-up: Why introduce a Service instead of connecting to Pod IPs directly?
52. Node affinity and taint tolerations?
- nodeSelector/nodeAffinity: schedule Pods onto specific nodes (GPU nodes)
- taints/tolerations: mark nodes as "dedicated" so only Pods with the matching toleration land there
- Combine them: taint GPU nodes and add a toleration to inference Pods to keep non-GPU jobs from grabbing the cards
- Topology spread constraints (podTopologySpread) spread replicas across availability zones for resilience
Follow-up: How do you keep regular jobs from crowding out GPU nodes?
53. What is a Kubernetes Operator, and why do ML platforms favor them?
- An Operator = CRDs that extend the API + a controller that continuously reconciles toward the desired state
- It turns complex applications like "deploy a model service" into declarative custom resources
- KServe, BentoML, and Kubeflow all follow the Operator pattern
- The payoff: loading, scaling, rollouts, and rollbacks handled automatically
Follow-up: To write an Operator for an inference service, what custom resources would you need to define?
Module 6: Monitoring and Reliability (10 Questions)
54. What do SLO, SLI, and SLA mean?
- SLI: the measured indicator (e.g., P99 latency). SLO: the target (e.g., P99 < 100 ms, 99% of the time)
- SLA: the external contract, with penalties for breach
- Error budget = 100% - SLO%; how much you've burned decides whether to freeze releases
- Process: business goals → key indicators → targets → monitoring and alerting
Follow-up: For a recommendation model service, which SLIs would you pick for the SLO?
55. What are the four golden signals (RED/USE)?
- Latency, Traffic, Errors, Saturation
- For inference, saturation means: GPU utilization, VRAM, queue depth, batch backlog
- "GPU utilization + queue depth" reflects inference health far better than CPU utilization
- Every signal needs a clear target and an alert
Follow-up: How do you pick saturation metrics? (GPU util, VRAM, queue length)
56. What are Prometheus's metric types, and how is P99 computed?
- counter (cumulative), gauge (instantaneous), histogram (bucketed cumulative)
- Histogram buckets feed histogram_quantile for P99; the error depends on bucket widths
- You need the full latency distribution—averages alone aren't enough
- Metric collection costs performance: control the sampling rate and label cardinality
Follow-up: How do you divide up histogram buckets? (Bucket by the characteristics of your latency distribution)
57. How do you detect data drift?
- Feature drift: statistical distribution changes (PSI, KS test). Concept drift: the input-output relationship changes
- Monitor input/output/prediction distributions and compare against a baseline on a schedule
- Lightweight approach: sample key features, compute PSI, alert past a threshold
- Response to drift: alert → manual review → rollback or retrain; see monitoring
Follow-up: Drift alerts fire too often—how do you handle false positives?
58. What are the principles of good alert design?
- Severity levels: P0 (service down) / P1 (SLO at risk) / P2 (latent risk)
- Multi-window rules (sustained for N minutes) cut flappy false positives; avoid "alert storms"
- Every alert carries: symptom, impact, and remediation steps (a runbook)
- Prefer error budgets over raw thresholds: alert only when the budget is nearly spent
Follow-up: How do you tame an alert storm? (Severity tiers, grouping, silencing, automated handling)
59. How do you do distributed tracing and locate slow requests?
- Instrument with OpenTelemetry: traces + spans, correlated across services
- The chain covers: gateway → preprocessing → inference engine → postprocessing → storage
- To find slowness: sort spans by duration and hunt the hot span
- Weigh trace sampling rate against cost (head sampling + tail sampling)
Follow-up: P99 regressed on an inference request—what's your investigation order? (See observability in practice)
60. What should structured logs record?
- Per request: request_id, model version, latency, batch size, result summary
- System level: OOM, load failures, timeouts, degradation events
- JSON structure makes logs searchable; keep sensitive data (user content) out of the log store
- Tie logs, metrics, and traces together with request_id
Follow-up: How do logs and metrics divide the work? (Metrics aggregate, logs drill down)
61. Fault drills and chaos engineering?
- Inject failures on purpose: kill Pods, cut the network, throttle, simulate model-load failures
- Verify self-healing: HPA scale-out, Pod recreation, traffic failover
- Turn drill results into runbooks and feed improvements back into SLOs
- Start with low-risk components and keep the blast radius small
Follow-up: What fault-injection experiments would you run against an inference service?
62. How do you plan capacity and warn on saturation?
- Build a QPS-to-resource model from load testing; see load testing methodology
- Watch trends and warn early (e.g., GPU utilization climbing past 80% and still rising)
- Keep headroom for scale-out (say 30%) and reserve time for cold starts
- Run dedicated capacity reviews before big sales or events
Follow-up: Doubling QPS—how many more GPUs do you need? Show your estimation process.
63. How do you combine metrics, logs, and traces to debug production issues?
- Metrics flag the anomaly (P99 rising) → logs show what requests in that window looked like → traces drill into specific spans
- request_id is the primary key joining all three
- Common patterns: queue backlog (metrics), missing warm-up (logs), engine compilation jitter (traces)
- Build a knowledge base of "known issue → symptom → fix"
Follow-up: Walk through a real troubleshooting scenario end to end, on the spot.
Module 7: MLOps (10 Questions)
64. How does CI/CD for ML differ from software CI/CD?
- Software CI/CD only manages code; ML adds two volatile inputs: data and models
- The ML pipeline: data validation → training → evaluation → packaging → deployment → monitoring
- Model evaluation needs offline benchmarks plus online validation—unit tests alone don't cut it
- Reproducibility: pin data versions, code versions, and environments (see the MLOps pipeline)
Follow-up: What's fundamentally different between "model tests passed" and "software tests passed"?
65. What problem does a model registry solve?
- Centralized versioning and metadata management (metrics / data versions / hyperparameters)
- Stage management (staging/production) with an audit trail of who registered and who shipped
- It turns "which version to deploy" into a configuration change, not a code change
- MLflow or an in-house registry is a core ML platform component
Follow-up: What metadata belongs in a model registry?
66. How do you design a model rollout process?
- Flow: offline evaluation passes → staging verification → canary (5% → 25% → 50%) → full rollout
- Each step has explicit promotion criteria (online metric comparisons)
- Automatic rollback on failure; see the rollout process
- Automate the launch: configuration-driven, not manual
Follow-up: How does a canary decide the new model is safe? (Metric confidence intervals, error budget)
67. What rollback strategies are there?
- Model-level rollback: switch back to the previous model version (seconds—a config flip)
- Code-level rollback: revert the inference service image
- Data-level rollback: rare, but you need a plan for it
- After rollback, monitor until you confirm you're back at baseline—don't just flip and walk away
Follow-up: An hour after full rollout, the new model's results drop. What's your rollback procedure?
68. A/B testing and shadow deployment?
- A/B: split real traffic between two versions by ratio and compare on business metrics
- Shadow deployment: mirror traffic to the new version without returning its responses—validation only
- Splitting tools: Istio/Ingress weighted routing or in-house splitters
- Keys: consistent metric definitions, enough samples, no sample contamination
Follow-up: How large does an A/B sample need to be? (Estimate from effect size and confidence level)
69. What happens when training and serving preprocessing disagree?
- If training-time normalization/scaling differs from production, the input distribution shifts and quality collapses
- Common traps: image resize method, tokenizer version, missing-value imputation
- Fix: share preprocessing code (same repo/same package) and run consistency tests before launch
- This is one of the core motivations for feature stores
Follow-up: Production underperforms offline—how do you tell whether preprocessing is to blame?
70. How does experiment tracking work, and why is it needed?
- Record: data version, code commit, hyperparameters, environment, evaluation metrics
- Goals: reproducible, comparable, replayable
- Tools: MLflow Tracking, W&B, and friends
- For comparisons: a shared baseline and controlled variables
Follow-up: Two experiments look identical but produce different results—how do you debug? (Environment or data version differences)
71. Why do you need data version control (DVC, etc.)?
- Data changes; model reproducibility depends on data versions
- Data lives in object storage (large files); DVC tracks file → version → repository
- Tie it to code versions so "this commit + this data = this model" always holds
- Data compliance: versioning makes audits and rollbacks possible
Follow-up: Data got overwritten—how do you recover the old training result?
72. What is the pre-launch model acceptance checklist?
- Offline: full evaluation-set metrics, per-scenario/per-segment metrics, comparison against baseline
- Robustness: handling of malformed inputs, missing fields, over-length inputs
- Performance: load-test targets met (latency/QPS), warm-up, capacity
- Engineering: logging/monitoring/alerting in place, rollback plan, SLO alignment
Follow-up: For a model you own, what are the last three items on the acceptance checklist before launch?
73. Build or buy an MLOps platform?
- Build: maximum control but high cost (headcount + maintenance)—fits large, mature teams
- Assemble open source: Kubeflow/MLflow/Argo plus cloud services—fast to start
- Buy SaaS: fastest start, but limited customization and data-egress risk
- Decide based on: team size, data compliance, business complexity
Follow-up: For a 10-person team adopting MLOps, what's the lowest-cost starting point?
Module 8: LLM Inference (10 Questions)
74. What is the autoregressive generation process?
- Generate one token at a time, feeding it back as input for the next step
- Serial dependency: token t+1 depends on the previous t tokens → full parallelism is impossible
- Two phases: prefill (processes the input in parallel, computing the KV cache) and decode (generates token by token)
- Key metrics: TTFT (time to first token), TPOT (time per output token)
Follow-up: Why is GPU utilization low during decode? (Small batches, bandwidth-bound)
75. What is the KV cache? How do you estimate its memory?
- Cache the K/V each attention layer computes so decode doesn't recompute them
- Memory ≈ 2 × layers × hidden × seq_len × batch × bytes
- Example: LLaMA-70B, 80 layers, hidden 8192, seq 2048, batch 32, FP16: 2×80×8192×2048×32×2 B ≈ 171 GB
- The KV cache is the dominant memory cost for long contexts and large batches
Follow-up: Given the formula, estimate KV cache memory on the spot.
76. How does continuous batching work?
- Traditional static batching: the whole batch waits for the slowest request before any slot frees up
- Continuous batching: at the end of each decode step, finished sequences exit immediately and new requests join anytime
- Step-level scheduling → a dramatic boost in GPU utilization
- vLLM made it practical at scale (see the vLLM case study)
Follow-up: What does continuous batching demand from the scheduler? (Check finished sequences every step, manage memory)
77. What problem does PagedAttention solve?
- Traditional KV caches allocate contiguously: fragmentation + waste from pre-allocation
- Paging + block tables: allocate 16-token blocks on demand—physically non-contiguous but logically contiguous
- Eliminates internal/external fragmentation and enables longer contexts and bigger batches
- Block-table mapping enables copy-on-write, which supports parallel sampling
Follow-up: How do you choose the page size (block size)? How does it affect efficiency?
78. How does speculative decoding work?
- A draft model quickly proposes k candidate tokens; the large model verifies them in parallel
- If all are correct, k tokens are accepted in one step → an effective speedup with the output distribution unchanged
- The speedup ceiling depends on the draft model's accuracy (acceptance rate)
- In practice, N-gram drafts can replace the small model
Follow-up: Does speculative decoding lose accuracy? (No—rejection sampling guarantees the distribution stays identical)
79. How does tensor parallelism work?
- Split a single weight matrix by rows/columns across multiple cards; each card computes its shard, then all-reduce
- Fits models with heavy per-layer compute (LLMs); the workhorse for serving 70B/175B models
- Communication overhead grows with card count → interconnect (NVLink/RDMA) is critical
- Compare with pipeline parallelism (split by layer) and data parallelism (replicate the model)
Follow-up: 8-way tensor parallel vs. 8-way data parallel—what's the trade-off? (Per-request latency vs. throughput)
80. How do you estimate GPU memory for LLM inference?
- Total memory = weights + KV cache + activations + framework overhead
- Weights: parameter count × bytes (FP16: 2 B); the KV cache uses the formula from the previous question
- Activations matter during prefill and shrink during decode
- Example: 7B FP16 weights are 14 GB; add the KV cache and CUDA overhead, and you want a 32 GB card to run long contexts comfortably
Follow-up: To fit 70B onto 2×A100-80G, what do you need? (Quantization + tensor parallelism)
81. How does optimization differ between prefill and decode?
- Prefill: compute-bound and parallelizable; the focus is memory (long-sequence activations)
- Decode: bandwidth-bound and serial; batch size is the key throughput lever
- Different levers: prefill uses FlashAttention/long-sequence optimizations; decode relies on continuous batching
- Latency budget: TTFT (first token) is set by prefill; TPOT is set by decode
Follow-up: An LLM service has a normal P50 but long TTFT—what could be causing it?
82. How does quantization affect LLM inference?
- W4A16 (GPTQ/AWQ): 75% smaller size and bandwidth, a clear decode speedup
- FP8 (W8A8): good accuracy with Hopper+ hardware acceleration—the emerging trend
- The KV cache can be quantized too (INT8/FP8) for further memory savings
- Quantization frees up memory → larger batches/longer contexts → higher throughput
Follow-up: What scenarios suit an INT4 weights + FP8 KV cache configuration?
83. What are the challenges of long-context inference, and how do you handle them?
- The KV cache grows linearly with seq_len, exploding memory
- Responses: quantize the KV cache, reuse memory with PagedAttention, cache eviction/compression
- Long contexts stretch prefill time-to-first-token; chunked processing helps
- Related optimizations: FlashAttention, sparse attention, position encoding extrapolation
Follow-up: Serving 128K contexts—name three memory and latency challenges.
Module 9: System Design (5 Questions)
84. Design a high-concurrency inference service (CV classification/detection)
Design points:
- Clarify requirements: QPS, P99, model size, batching preferences
- Architecture: gateway (rate limiting/auth) → preprocessing → inference engine (Triton/ORT) → postprocessing → storage
- Dynamic batching + multiple replicas + horizontal scaling; warm-up and readiness
- Monitoring: latency distribution, GPU utilization, error rate; SLOs and alerting
- Capacity: load-test modeling → GPU count estimation → scale-out headroom
Follow-up: QPS goes from 1k to 10k—which part of the architecture breaks first? How does it evolve?
85. Serve a 70B model
Design points:
- Memory estimation: weights (FP16 140 GB / INT4 35 GB) + KV cache
- Parallelism: tensor parallelism (8 cards) as the primary strategy for latency; pipeline parallelism to extend throughput
- Engine choice: vLLM/TensorRT-LLM (continuous batching + PagedAttention)
- Quantization: FP8/INT4 to balance accuracy and capacity
- Elasticity: separate deployments for online latency vs. offline throughput (see deployment patterns)
Follow-up: How do online (low-latency) and offline (high-throughput) deployments differ? How do you adjust when the card allocation changes?
86. Capacity-plan a recommendation service
Design points:
- Set targets: peak QPS, P99, availability
- Load-test to measure per-machine/per-card capacity (QPS@P99) and build a linear model
- Formula: replicas = peak QPS / per-replica capacity × redundancy factor (1.3-1.5)
- Account for failure domains (AZs) and scale-out time (cold start/image pull)
- Run dedicated load tests and pre-scale before big sales events; watch trends the rest of the time
Follow-up: Traffic suddenly triples—can you scale out in time? (Warm-up time vs. scale-out time)
87. Design a unified multi-model inference platform
Design points:
- Goal: one unified entry point for multiple frameworks (ONNX/TensorRT/PyTorch)
- Foundation: Kubernetes + custom resources (KServe-style) + a model repository
- Capabilities: automatic loading, dynamic batching, gradual rollout, A/B, monitoring, quotas (multi-tenant)
- Layers: access (routing/auth/rate limiting) → scheduling (GPU allocation) → execution (inference engines)
- Extensions: model warm-up, traffic-based scaling, cost accounting
Follow-up: Many small models vs. one large model—how does the scheduling strategy differ?
88. Design a model deployment plan for edge devices
Design points:
- Hardware constraints: compute/memory/bandwidth/power; see hardware fundamentals
- Model side: distilled small model + INT8 quantization + pruning (compression)
- Engine choice: TensorRT/TFLite/OpenVINO/llama.cpp, depending on the hardware ecosystem
- Edge-cloud split: low latency at the edge with the cloud as fallback; OTA model updates
- Evaluation: measure on real hardware (latency/power/accuracy)—paper FLOPS don't count
Follow-up: An edge model update fails—how do you roll back? How do you monitor devices running offline?
Module 10: Live Coding Prompts (6 Questions)
Interviewers often test short problems that reveal how you think about concurrency and batching. Each item below lists what's being tested and a reference approach; practice full implementations in the build-your-own projects until they're second nature.
89. Write a simple batch aggregation function
What's tested: sliding-window aggregation, timeout and max capacity, thread safety.
Reference approach: build an aggregator with max_batch and max_wait; requests enter a queue, and a timer or a full queue triggers flush(); on flush, take a snapshot of the queue, run inference in the callback, and dispatch the results.
Follow-up: How would you implement it with asyncio? What about multiple threads?
90. Implement a fixed-capacity FIFO queue
What's tested: ring buffers, thread safety, blocking vs. non-blocking.
Reference approach: array + head/tail pointers with a size counter; put waits or returns false when full; get waits when empty; use locks or atomic counters for concurrency.
Follow-up: How would you optimize for a read-heavy workload?
91. Implement sliding-window P99 statistics
What's tested: streaming data, approximate statistics.
Reference approach: a bucketed histogram (fixed buckets) to approximate P99; or maintain an ordered structure (SortedList/tree); or sample. State the accuracy vs. complexity trade-off.
Follow-up: Production requires that P99 stats never block the inference thread—what do you do?
92. Implement a request aggregator with a timeout (time-boxed batching)
What's tested: concurrency primitives (timeout, Condition).
Reference approach: check elapsed wait time on every put using a condition variable + timestamps; the timeout triggers flush; watch for races (flush running concurrently with put).
Follow-up: How do you prevent new requests arriving during flush from scrambling the batch?
93. Implement a token-bucket rate limiter
What's tested: algorithm + concurrency + timestamps.
Reference approach: track last_refill_time and the token count; refill tokens at the configured rate; return 429 when tokens run out; use single-flight/atomic operations for concurrency safety.
Follow-up: Token bucket vs. leaky bucket—what's the difference, and how does each behave under bursty traffic?
94. Given an LLM service, write a generate wrapper with concurrency control (pseudocode)
What's tested: semaphore + queue + timeout, plus engineering judgment.
Reference approach: Semaphore(max_concurrency); enqueue and wait; wait_for with a timeout; once you get a slot, call the inference endpoint, return the result, and release in a finally block.
Follow-up: Why cap concurrency instead of allowing unlimited concurrent requests? (Memory, runaway queuing)
Appendix: Tips for the Interview Room
| Tip | What it means |
|---|---|
| Define first, then expand | For every question, say what it is in one sentence, then explain the mechanism, then give the engineering trade-offs |
| Volunteer the trade-offs | "I'd pick Triton over X because Y" reads far better than listing tool names |
| Reason out loud when stuck | "I'm not sure of the exact implementation, but I'd guess it relates to X, so I'd start by checking X" |
| Steer questions toward your resume projects | Back up abstract answers with real numbers from your own work (pair with resume writing) |
| Ask questions at the end | Ask about the team's deployment stack, SLOs, and GPU fleet—it's your chance to screen the company right back |
Further Reading
- JD knowledge map — trace each question's weak spots back to its page
- JD list: open roles at companies in China and abroad — see what your target roles emphasize
- Learning paths: three routes — the companion plan for a 2-week job-hunt sprint
- Inference fundamentals and anatomy of an inference system — the deep foundation for Modules 1 and 2
- Quantization and LLM inference — the high-frequency deep-dive areas for Modules 3 and 8
- Common pitfalls — every trap interviewers love to test lives here
- Build your first inference service — the practice ground for live coding and portfolio projects