Skip to content

Breaking Down JD Requirements

At a glance Which bodies of knowledge hide behind JD keywords like K8s, ONNX, TensorRT, vLLM, and performance tuning? This page maps high-frequency JD keywords to the relevant pages on this site, with recommended mastery depth, so you know what to learn and how deep to go.

Breaking Down JD Requirements ​

The most intimidating part of a JD isn't the "3 years of experience" — it's the string of nouns: K8s, TensorRT, vLLM, INT8 quantization, dynamic batching, capacity planning... Behind each term sits a whole body of knowledge, and the depth required differs completely — some only demand "general awareness," others "derive the formulas on the spot."

This page maps high-frequency JD keywords to the corresponding pages on this site and adds recommended mastery depth and common interview follow-ups, so you know what to learn and how deep to go — avoiding the mismatch of grinding on K8s for three months only to end up at "aware" level.

Depth definitions

Aware = can say in one sentence what it is and what problem it solves. Proficient = can set it up and use it hands-on, and explain the key parameters. Expert = can explain the underlying principles, tune it, and answer counterexamples and edge cases.

1. The Master Mapping Table ​

Organized as "JD keyword → page on this site → recommended depth → interview follow-up direction." Get the full picture first, then read group by group.

JD KeywordPage on This SiteRecommended DepthCommon Interview Follow-ups
Kubernetesdeployment patternsProficientPod lifecycle, how HPA works
Dockeranatomy of an inference systemProficientImage layering, resource limits
ONNX / ONNX Runtimemodel formatsProficientPTH→ONNX pitfalls, what an execution provider is
TensorRTmodel formatsProficientThe optimization pipeline, INT8 calibration
Tritonmodel serving, Triton case studyProficientDynamic batching, concurrent model instances
vLLMLLM inference, vLLM case studyProficientPagedAttention, continuous batching
MLflowMLOps pipelinesAwareModel registry, experiment tracking
QuantizationquantizationExpert (high differentiation)PTQ/QAT, INT8 accuracy, GPTQ/AWQ
Model compressionmodel compressionAwarePruning vs. distillation vs. sparsification
Inference optimizationperformance optimization, inference fundamentalsExpertBottleneck localization methods, bandwidth-bound scenarios
Dynamic batchingmodel servingProficientThe latency-throughput trade-off
KV CacheLLM inferenceProficientVRAM estimation, long contexts
Performance tuningperformance optimization, load testingProficientHow to diagnose a P99 regression
Capacity planningload testingProficientPeak QPS to GPU-count estimation
Monitoring & alertingmonitoring, observability in practiceProficientSLOs, alert design, drift detection
CI/CDMLOps pipelinesProficientML CI/CD vs. software CI/CD
Gradual rolloutrollout processProficientJudging rollout quality, rollback
GPU / VRAMhardware fundamentalsProficientMemory bandwidth, why inference is bandwidth-bound
SecuritysecurityAwareModel theft, unauthorized access, auditing
Glossary / tool overviewGlossary, tool inventoryLook up as needed—

2. Group-by-Group Deep Dive ​

1. Frameworks and Tools ​

KeywordCore PointsRecommended DepthPage on This Site
K8sDeclarative orchestration, Pod/Deployment/Service/HPA, GPU schedulingProficient: can draw the scheduling flow end to enddeployment patterns
DockerImages, container isolation, GPU passthrough (--gpus)Proficientanatomy of an inference system
ONNX RuntimeGraph optimization, operator fusion, the execution-provider mechanism, quantization supportProficientmodel formats
TensorRTparse→build→engine, INT8 calibration, hardware bindingProficient (Expert for engine roles)model formats, TensorRT edge case study
TritonMulti-framework support, dynamic batching, concurrent instances, ensemblesProficientmodel serving, Triton case study
vLLMPagedAttention, continuous batching, quantized model loadingProficient (Expert for LLM roles)LLM inference, vLLM case study
MLflowExperiment tracking, model registry, stage managementAwareMLOps pipelines

Too many tools to remember?

Start with the tool inventory to build a mental map of "where each tool sits in the pipeline," then go deep on the 2-3 core tools for your target role. Interviewers don't test whether you memorized flags — they test "why this tool, here."

2. Domain Knowledge ​

KeywordCore PointsRecommended DepthPage on This Site
Quantization (PTQ/QAT)scale/zero_point, calibration, accuracy acceptanceExpertquantization
Model compressionPruning, distillation, sparsification, knowledge transferAwaremodel compression
Inference optimizationOperator fusion, kernel selection, CUDA GraphsExpertinference fundamentals, performance optimization
Dynamic batchingRequest aggregation, timeout triggers, latency costProficientmodel serving
KV CacheAutoregressive caching, VRAM footprint, pagingProficientLLM inference

3. Engineering Skills ​

KeywordCore PointsRecommended DepthPage on This Site
Performance tuningProfiling workflow, P99 analysis, bottleneck taxonomyProficientperformance optimization
Capacity planningLoad-test modeling, QPS→GPU conversion, headroom reservesProficientload testing
Monitoring & alertingGolden signals, SLOs, drift detectionProficientmonitoring, observability in practice
CI/CDTrain/evaluate/deploy pipelines, model registryProficientMLOps pipelines
Gradual rolloutCanary, blue-green, automatic rollbackProficientrollout process

4. Soft Skills ​

KeywordCore PointsRecommended DepthPage on This Site
Cross-team collaborationInterfaces with algorithm/backend/SRE teams, translating requirementsAwareanatomy of an inference system
DocumentationDeployment guides, runbooks, incident postmortemsAwareobservability in practice

Don't wing the soft skills

When asked "how do you collaborate with the algorithm team?", the winning answer is a concrete workflow: who provides the model → who owns serving → who signs off on quality → how production issues get escalated. The responsibility boundaries in the rollout process make a great answer template.

3. Tie It into the Site's Learning Paths ​

Studying isolated pages is inefficient. Chain your learning with the learning paths: three routes:

  1. Job-search sprint (2 weeks): sweep through the items marked "Proficient" above following the route map, prioritizing "can answer, can implement."
  2. Systematic deepening (8 weeks): go deep on the "Expert" items — quantization and inference optimization — and produce tangible results with the hands-on build.
  3. Desk reference (ongoing): index the "Aware"-level items as keywords, and consult the Glossary and tool inventory on demand.

4. Self-Assessment: 10 Questions to Gauge Your Level ​

You pass with ✅ if you can answer "what + why + how"; ⚠️ if you can only recite definitions and stall at follow-ups; ❌ if the question means nothing to you. After the self-test, map your ❌ and ⚠️ items back to the corresponding pages in the tables above and patch the gaps.

#Self-Test QuestionRelated PageResult
1Why is inference a better fit for quantization than training?inference fundamentals□
2What problem does each step solve when converting a PTH model to ONNX and then to TensorRT?model formats□
3After INT8 quantization, accuracy drops 3 points — in what order do you debug?quantization□
4A service's P99 jumped from 60ms to 200ms — how do you diagnose it?performance optimization□
5Dynamic batching doubled QPS — so why did latency get worse?model serving□
6How much VRAM does a 70B model need at FP16? How do you compute the KV Cache?LLM inference□
7Why isn't HPA enough for GPU inference services?deployment patterns□
8How would you design SLOs and alerts for an inference service?monitoring□
9When rolling out a model gradually, how do you judge whether the new model "works"?rollout process□
10Training-time and serving-time preprocessing differ on the same data — what happens?MLOps pipelines□

Once done, take your ❌/⚠️ questions to Model Deployment Interview Questions for the corresponding reference answers, then use Resume Benchmarking to write what you learned into your resume.

Further Reading ​