Appearance
Breaking Down JD Requirements
The most intimidating part of a JD isn't the "3 years of experience" — it's the string of nouns: K8s, TensorRT, vLLM, INT8 quantization, dynamic batching, capacity planning... Behind each term sits a whole body of knowledge, and the depth required differs completely — some only demand "general awareness," others "derive the formulas on the spot."
This page maps high-frequency JD keywords to the corresponding pages on this site and adds recommended mastery depth and common interview follow-ups, so you know what to learn and how deep to go — avoiding the mismatch of grinding on K8s for three months only to end up at "aware" level.
Depth definitions
Aware = can say in one sentence what it is and what problem it solves. Proficient = can set it up and use it hands-on, and explain the key parameters. Expert = can explain the underlying principles, tune it, and answer counterexamples and edge cases.
1. The Master Mapping Table
Organized as "JD keyword → page on this site → recommended depth → interview follow-up direction." Get the full picture first, then read group by group.
| JD Keyword | Page on This Site | Recommended Depth | Common Interview Follow-ups |
|---|---|---|---|
| Kubernetes | deployment patterns | Proficient | Pod lifecycle, how HPA works |
| Docker | anatomy of an inference system | Proficient | Image layering, resource limits |
| ONNX / ONNX Runtime | model formats | Proficient | PTH→ONNX pitfalls, what an execution provider is |
| TensorRT | model formats | Proficient | The optimization pipeline, INT8 calibration |
| Triton | model serving, Triton case study | Proficient | Dynamic batching, concurrent model instances |
| vLLM | LLM inference, vLLM case study | Proficient | PagedAttention, continuous batching |
| MLflow | MLOps pipelines | Aware | Model registry, experiment tracking |
| Quantization | quantization | Expert (high differentiation) | PTQ/QAT, INT8 accuracy, GPTQ/AWQ |
| Model compression | model compression | Aware | Pruning vs. distillation vs. sparsification |
| Inference optimization | performance optimization, inference fundamentals | Expert | Bottleneck localization methods, bandwidth-bound scenarios |
| Dynamic batching | model serving | Proficient | The latency-throughput trade-off |
| KV Cache | LLM inference | Proficient | VRAM estimation, long contexts |
| Performance tuning | performance optimization, load testing | Proficient | How to diagnose a P99 regression |
| Capacity planning | load testing | Proficient | Peak QPS to GPU-count estimation |
| Monitoring & alerting | monitoring, observability in practice | Proficient | SLOs, alert design, drift detection |
| CI/CD | MLOps pipelines | Proficient | ML CI/CD vs. software CI/CD |
| Gradual rollout | rollout process | Proficient | Judging rollout quality, rollback |
| GPU / VRAM | hardware fundamentals | Proficient | Memory bandwidth, why inference is bandwidth-bound |
| Security | security | Aware | Model theft, unauthorized access, auditing |
| Glossary / tool overview | Glossary, tool inventory | Look up as needed | — |
2. Group-by-Group Deep Dive
1. Frameworks and Tools
| Keyword | Core Points | Recommended Depth | Page on This Site |
|---|---|---|---|
| K8s | Declarative orchestration, Pod/Deployment/Service/HPA, GPU scheduling | Proficient: can draw the scheduling flow end to end | deployment patterns |
| Docker | Images, container isolation, GPU passthrough (--gpus) | Proficient | anatomy of an inference system |
| ONNX Runtime | Graph optimization, operator fusion, the execution-provider mechanism, quantization support | Proficient | model formats |
| TensorRT | parse→build→engine, INT8 calibration, hardware binding | Proficient (Expert for engine roles) | model formats, TensorRT edge case study |
| Triton | Multi-framework support, dynamic batching, concurrent instances, ensembles | Proficient | model serving, Triton case study |
| vLLM | PagedAttention, continuous batching, quantized model loading | Proficient (Expert for LLM roles) | LLM inference, vLLM case study |
| MLflow | Experiment tracking, model registry, stage management | Aware | MLOps pipelines |
Too many tools to remember?
Start with the tool inventory to build a mental map of "where each tool sits in the pipeline," then go deep on the 2-3 core tools for your target role. Interviewers don't test whether you memorized flags — they test "why this tool, here."
2. Domain Knowledge
| Keyword | Core Points | Recommended Depth | Page on This Site |
|---|---|---|---|
| Quantization (PTQ/QAT) | scale/zero_point, calibration, accuracy acceptance | Expert | quantization |
| Model compression | Pruning, distillation, sparsification, knowledge transfer | Aware | model compression |
| Inference optimization | Operator fusion, kernel selection, CUDA Graphs | Expert | inference fundamentals, performance optimization |
| Dynamic batching | Request aggregation, timeout triggers, latency cost | Proficient | model serving |
| KV Cache | Autoregressive caching, VRAM footprint, paging | Proficient | LLM inference |
3. Engineering Skills
| Keyword | Core Points | Recommended Depth | Page on This Site |
|---|---|---|---|
| Performance tuning | Profiling workflow, P99 analysis, bottleneck taxonomy | Proficient | performance optimization |
| Capacity planning | Load-test modeling, QPS→GPU conversion, headroom reserves | Proficient | load testing |
| Monitoring & alerting | Golden signals, SLOs, drift detection | Proficient | monitoring, observability in practice |
| CI/CD | Train/evaluate/deploy pipelines, model registry | Proficient | MLOps pipelines |
| Gradual rollout | Canary, blue-green, automatic rollback | Proficient | rollout process |
4. Soft Skills
| Keyword | Core Points | Recommended Depth | Page on This Site |
|---|---|---|---|
| Cross-team collaboration | Interfaces with algorithm/backend/SRE teams, translating requirements | Aware | anatomy of an inference system |
| Documentation | Deployment guides, runbooks, incident postmortems | Aware | observability in practice |
Don't wing the soft skills
When asked "how do you collaborate with the algorithm team?", the winning answer is a concrete workflow: who provides the model → who owns serving → who signs off on quality → how production issues get escalated. The responsibility boundaries in the rollout process make a great answer template.
3. Tie It into the Site's Learning Paths
Studying isolated pages is inefficient. Chain your learning with the learning paths: three routes:
- Job-search sprint (2 weeks): sweep through the items marked "Proficient" above following the route map, prioritizing "can answer, can implement."
- Systematic deepening (8 weeks): go deep on the "Expert" items — quantization and inference optimization — and produce tangible results with the hands-on build.
- Desk reference (ongoing): index the "Aware"-level items as keywords, and consult the Glossary and tool inventory on demand.
4. Self-Assessment: 10 Questions to Gauge Your Level
You pass with ✅ if you can answer "what + why + how"; ⚠️ if you can only recite definitions and stall at follow-ups; ❌ if the question means nothing to you. After the self-test, map your ❌ and ⚠️ items back to the corresponding pages in the tables above and patch the gaps.
| # | Self-Test Question | Related Page | Result |
|---|---|---|---|
| 1 | Why is inference a better fit for quantization than training? | inference fundamentals | □ |
| 2 | What problem does each step solve when converting a PTH model to ONNX and then to TensorRT? | model formats | □ |
| 3 | After INT8 quantization, accuracy drops 3 points — in what order do you debug? | quantization | □ |
| 4 | A service's P99 jumped from 60ms to 200ms — how do you diagnose it? | performance optimization | □ |
| 5 | Dynamic batching doubled QPS — so why did latency get worse? | model serving | □ |
| 6 | How much VRAM does a 70B model need at FP16? How do you compute the KV Cache? | LLM inference | □ |
| 7 | Why isn't HPA enough for GPU inference services? | deployment patterns | □ |
| 8 | How would you design SLOs and alerts for an inference service? | monitoring | □ |
| 9 | When rolling out a model gradually, how do you judge whether the new model "works"? | rollout process | □ |
| 10 | Training-time and serving-time preprocessing differ on the same data — what happens? | MLOps pipelines | □ |
Once done, take your ❌/⚠️ questions to Model Deployment Interview Questions for the corresponding reference answers, then use Resume Benchmarking to write what you learned into your resume.
Further Reading
- Job Description Checklist — where these keywords come from
- Career: Overview and Job Landscape — decide which direction to focus on
- Model Deployment Interview Questions — the final exam for your self-test
- Learning Paths: Three Routes — turn this page's list into a weekly plan
- What Is Model Deployment — the entry point and foundation for all concept pages
- Tool Inventory — a quick-reference manual at the tool level