Skip to content

Skills Benchmarking: What to Highlight on Your Resume

At a glance A resume for an inference & deployment role is not a list of experiences — it's an evidence chain of "performance data + systems understanding + engineering ability." This page gives you a JD skills-benchmarking method, a deployment-flavored STAR format, 6 project rewrites, a layered tech-stack strategy, ten fatal counterexamples, and a ready-to-use resume template skeleton.

This page contains time-sensitive content. Data is current as of 2026-08; engine versions, benchmark rankings, and product features may have changed since — verify against primary sources before citing.

Skills Benchmarking: What to Highlight on Your Resume ​

One sentence to define it: a resume for an inference & deployment role is not your list of experiences — it is an "evidence chain" that uses verifiable facts to prove to the interviewer "I can deliver the capabilities this JD asks for." A resume is persuasion, not enumeration.

Think of a resume as a performance test report: the interviewer (the reviewer) only trusts statements backed by data. You say "I'm proficient in vLLM" — where's the evidence? You say "I have large-scale inference experience" — where's the evidence? You say "I can optimize performance" — where's the evidence? Claims without data count as blank in an HR screener's eyes; a resume rich in evidence, where every project carries "latency from X to Y, throughput from A to B, GPU utilization from C to D," makes the interviewer pre-assume you're half-qualified before meeting you.

This page is the third stop of the career module: the first two stops, the JD List (what the market is hiring for) and Deconstructing JD Knowledge Points (what you can and can't do), converge here into a resume — rewrite the experience you already have through the interviewer's eyes, not fabricate new history. Write with "data + source code + retrospectives" throughout, and adopt one brutal premise: the interviewer assumes everything on your resume could be fake; your job is to make every line survive follow-up questions.

1. The "Three Elements" of an Inference & Deployment Resume ​

The biggest difference from an algorithm-role resume: algorithm roles talk AUC / accuracy; deployment roles talk latency / throughput / cost. The three elements, in weight order:

┌─────────────────────────────────────────────────────┐
│ Element 1  Performance data (highest weight)         │
│            Latency from X ms to Y ms                 │
│            Throughput from A tokens/s to B tokens/s  │
│            GPU utilization from C% to D%             │
│            Cost saved E% / $F per month              │
└─────────────────────────────────────────────────────┘
                    │
                    ▼
┌─────────────────────────────────────────────────────┐
│ Element 2  Systems understanding (medium weight)     │
│            Can explain vLLM's Scheduler / KV cache / │
│            batching                                  │
│            Can locate bottlenecks in Nsight reports  │
│            Can explain TP / PP / EP communication    │
│            costs                                     │
└─────────────────────────────────────────────────────┘
                    │
                    ▼
┌─────────────────────────────────────────────────────┐
│ Element 3  Engineering ability (base weight)         │
│            Python / C++ / Linux / K8s in production  │
│            CI / CD / monitoring / canary             │
│            Team collaboration, interfacing with      │
│            algorithm teams                           │
└─────────────────────────────────────────────────────┘

You need all three elements

Only performance data, no systems understanding → looks like ops, doesn't know why; only systems understanding, no performance data → looks like a slide engineer, never actually shipped; only engineering ability, no performance data → looks like an SRE who never touched inference itself. A resume with all three is an "inference & deployment engineer" resume; anything less is a lopsided partial score.

1. Three "resume myths" and their flip sides ​

MythThe truth
The fuller the resume, the better — pack everything inInformation density is not evidence density; every extra line of noise dilutes a line of evidence
List everything you've done — something will impress HRWhat impresses HR is not "you did it" but "what you achieved, and how well"
The more skills stacked, the saferStacked nouns don't survive follow-ups; getting caught on one "proficient in vLLM source code" collapses the credibility of the whole page

Turning "things I did" into "results I delivered" is the single core action of the entire resume. Fail that translation and the resume is just a "chronicle of experiences" — HR reads a hundred chronicles a day; one more won't hurt.

2. The evidence-chain model: JD requirement → capability item → evidence ​

The relationship between a qualified candidate and a qualified resume:

┌────────────────────────────────────────────────────────────┐
│  Job requirement (JD)                                        │
│  e.g.: Familiar with vLLM or TensorRT-LLM source code;      │
│        experience deploying large-scale inference services   │
└─────────────────────────┬──────────────────────────────────┘
                          │ Decompose
                          ▼
┌────────────────────────────────────────────────────────────┐
│  Capability items (units testable as can/can't)              │
│  ① Source-level understanding of vLLM (Scheduler /           │
│    PagedAttention / KV cache)                                │
│  ② Large-scale deployment experience (multi-replica,        │
│    canary, monitoring & alerting)                            │
│  ③ Performance tuning experience (quantization, batching,  │
│    speculative decoding)                                     │
└─────────────────────────┬──────────────────────────────────┘
                          │ Pair item by item
                          ▼
┌────────────────────────────────────────────────────────────┐
│  Evidence (every line on the resume, verifiable on the spot)│
│  Project: upgraded internal LLM service from vLLM 0.5 to    │
│        0.8 + chunked prefill; TTFT 1.8s → 0.6s,             │
│        throughput +120%                                      │
│  Open source: 3 PRs to vLLM (block_table index optimization)│
│  Blog: wrote a PagedAttention internals breakdown, 2000+     │
│  reads                                                       │
└────────────────────────────────────────────────────────────┘

The standard form of one piece of evidence is "what I did + how I did it + how well it turned out," not "what I was exposed to." Miss any of the three and the evidence is incomplete: only "I used vLLM" and the interviewer can't judge your depth; only "TTFT down 70%" and the interviewer can't tell whether you achieved it by tuning a flag.

3. What the interviewer is actually scanning for ​

  • Hard-skill evidence: is your "vLLM / CUDA / quantization" at the level of "used," "can use," or "proficient"? Can you describe what PagedAttention's block table looks like? — this decides whether you pass the technical screen.
  • Delivery evidence: have you taken an inference service end-to-end and produced quantifiable results? — this decides whether the project deep-dive has substance.
  • Introspection evidence: can you explain failures and trade-offs? What pits did you fall into? — this decides whether you're worth long-term investment.

Three hard standards for evidence

① Verifiable: the interviewer can confirm with a quick check (the GitHub repo exists, the PR exists, the benchmark report exists, the blog exists); ② Quantifiable: whatever can be a number shouldn't be an adjective; ③ Follow-up-proof: behind every number, you can speak at least three sentences of detail. Any "evidence" failing any of these gets demoted to "experience" first — then reconsidered for inclusion.

2. The JD Skills-Benchmarking Method: Decompose Requirements into an Evidence Checklist ​

Roles vary enormously, but there is only one method: decompose every JD requirement into capability items, then find at least one piece of evidence for each. For a capability item with no evidence, either build it (do a project, land a PR, write a blog) or drop the application — don't pad.

1. Decomposition demo: a real inference-engineer JD ​

Assume the target JD's requirements read as follows (anonymized and simplified):

Requirements: ① Bachelor's degree or above in computer science / electrical engineering / automation or a related field; ② proficient in Python and C++, familiar with Linux; ③ deep understanding of PyTorch internals, source-reading experience in at least one of vLLM / TensorRT-LLM / SGLang; ④ experience deploying large-scale LLM inference services preferred; ⑤ familiar with GPU performance analysis (Nsight / PyTorch Profiler), CUDA / Triton kernel experience a plus; ⑥ good communication and collaboration skills.

Decomposition result:

JD requirementCapability itemWhere the evidence comes from
② Python / C++ / LinuxProduction-grade coding abilityCode volume in projects, CI/CD configuration, performance-analysis cases
③ PyTorch internals + engine sourcevLLM source reading + modificationOpen-source PRs, internal fork patches, technical blog
① Educational backgroundTrained in the fieldEducation section (state it truthfully)
④ Large-scale deploymentServing abilityMulti-replica deployment, canary rollout, monitoring & alerting in projects
⑤ Performance analysis / kernelsCUDA / Nsight in productionLatency-optimization cases in projects, kernel rewrites
⑥ Communication & collaborationInterfacing with algorithm teamsProject role descriptions, cross-team cases

2. Four evidence sources, ranked by credibility ​

Evidence sourceExamplesCredibilityNotes
Internship / workCompany projects, production effects, SLO data★★★★★Hardest currency, but not always available
Portfolio projectsEnd-to-end projects with repo, README, benchmark★★★★Fully within your control — a must; how to build them in Portfolio Projects
Open-source contributionsPRs to vLLM / SGLang / TensorRT-LLM★★★★★The hardest "bonus item" for this role; verifiable
Papers / blogsTechnical blogs, arXiv papers★★★★Shows depth, but must be well written
Courses / certificatesOnline courses, certifications★★☆Proves learning intent only; can't stand alone

Don't mistake "learning in progress" for evidence

A line like "Completed the 'LLM Inference Optimization' course on Coursera in 2025" tells the interviewer "I have no hands-on experience." Courses are for laying foundations, not resume material; resume material must be things you produced. Course background can go into the notes of the education section, but it must never occupy project slots.

Special bonus channels for this role:

  • Open-source PRs are the hardest evidence: a merged PR to vLLM / SGLang / TensorRT-LLM / llama.cpp is hard currency the interviewer can verify with one click on GitHub. One merged PR is worth five project entries.
  • Technical blogs beat course certificates: write a "PagedAttention internals + measured comparison" and publish it on Medium / your personal blog — depth plus verifiability.
  • Benchmark reports beat self-assessment: publish a "vLLM vs TensorRT-LLM on Llama-70B" comparison with a GitHub repo and reproducible code — a hundred times stronger than "proficient in vLLM" on a resume.

3. Hands-on: decompose one JD in ten minutes ​

  1. Circle every verb phrase in the JD ("familiar with," "proficient in," "experience with," "understanding of") and convert each into a capability item.
  2. Score yourself item by item with Deconstructing JD Knowledge Points: can do / half-can / can't.
  3. Find matching evidence for every "can do"; decide whether to close "half-can" gaps before the deadline; never pad "can't."
  4. Judge person-role fit from the results: below 60% fit, the ROI of applying is low unless it's a learning-oriented role.
  5. Save the decomposition as a file and re-run it against the target JD before every application — this is the raw material of "tailoring the resume per JD" (see Section 8).

3. Writing Project Experience: the Deployment-Flavored STAR ​

Project experience is the load-bearing wall of the resume. The one-sentence principle: use the STAR structure to upgrade "what I did" into "in what situation, to solve what problem, taking what action, with what quantifiable performance result."

1. The STAR format, adapted for inference & deployment ​

Classic STAR has four elements: Situation, Task, Action, Result. In the inference & deployment context, each element must answer one technical question:

ElementGeneric questionThe deployment-flavored version must answerCounterexample
S — SituationWhat's the background?Business scenario, model size, hardware, QPS, constraints"I worked on LLM deployment"
T — TaskWhat problem were you solving?Which metric — TTFT / TPOT / throughput / cost? Baseline?"Deployed Llama"
A — ActionWhat exactly did you do?Engine selection, tuning, source modification, kernel writing — why these choices, what pits you hit"Used vLLM"
R — ResultWhat was the outcome?Performance metrics + business metrics, with comparison baselines"Worked pretty well"

Note the two keywords in the Action element: "why this choice" and "what pit I hit." The former proves judgment rather than blind tuning; the latter proves you actually did it — people who haven't done it can't invent real pitfalls.

2. Four tiers of performance data ​

Performance metrics for this role come in four tiers; a resume should carry at least two:

TierDefinitionExamplesCommon mistakes
LatencySingle-request responsivenessTTFT, TPOT, E2E latency, P50/P99Only averages, no P99
ThroughputSystem-level outputQPS, tokens/s, RPSNo batch size or seq_len stated
ResourceHardware utilizationGPU utilization, memory footprint, HBM bandwidth utilizationNo baseline comparison
BusinessBusiness impactCost savings, availability, canary success rateMissing entirely — exposes "no business sense"

Four writing rules:

  1. Always carry a comparison baseline: "TTFT 600ms" carries no information; "TTFT from 1.8s to 600ms (vLLM 0.5 → 0.8 + chunked prefill)" does — the comparator can be an older version, another engine, or the team's previous numbers.
  2. Latency and throughput must be consistent: claiming "1000 QPS" without stating "P99 latency" greatly reduces credibility; high QPS usually trades latency for throughput — write the trade-off point clearly.
  3. Label the measurement context honestly: offline benchmark, single-card measurement, online A/B, or full rollout — different contexts produce incomparable numbers; state it.
  4. Numbers must survive division and follow-ups: before writing "latency down 80%," be clear relative to what — the interviewer will absolutely ask.

4. Six Sample Project Rewrites ​

Below are STAR rewrites of six typical inference & deployment projects. Each follows a "❌ plain version → ✅ strong version" contrast, ending with "follow-up points" — the entry points where an interviewer will dig; you must have thought them through.

Example 1: Shipping an LLM inference service ​

❌ Plain version (chronicle — all three sentences say "what I did"):

Deployed a Llama-3-8B inference service based on vLLM. Wrapped the API with Python and FastAPI, deployed to a Kubernetes cluster. Configured PagedAttention and continuous batching; the service runs stably.

✅ Strong version (evidence chain — all four elements complete):

Llama-3-8B inference service launch — backend for an internal document assistant

  • Situation: the internal AI assistant handles 500k calls/day; the original TGI service hit a P99 of 12s, with business complaints centered on "waiting until timeout"; hardware was 4×A100 80G, 8 replicas on a K8s cluster, fixed budget.
  • Task: migrate to vLLM 0.8 targeting "P99 < 3s, QPS > 200," with streaming output and multi-model coexistence.
  • Action: ① benchmarked vLLM 0.8 / TGI / TensorRT-LLM under the 8B model + 2k context + batch=64 load — vLLM had the highest throughput and lowest TTFT; ② enabled PagedAttention and chunked prefill, tuned max_num_batched_tokens=4096, and rewrote the priority scheduling in LLMEngine so VIP requests go first; ③ used Nsight Systems to locate the attention kernel as the bottleneck; switching to FlashAttention-2 cut TTFT another 25%; ④ deployed with Helm + ArgoCD, instrumented TTFT / TPOT / queue length in Prometheus + Grafana.
  • Result: P99 from 12s to 2.4s (-80%), QPS from 80 to 240 (+200%), GPU utilization from 35% to 72%; after launch, 8 business lines onboarded, saving RMB 180k per month in GPU cost.

Follow-up points: ① What was your selection logic between vLLM 0.8 and TensorRT-LLM? ② How did you tune max_num_batched_tokens for chunked prefill? ③ Did you measure FlashAttention 1 vs 2 vs 3? ④ How did you implement the priority scheduling?

Example 2: INT8 quantization in production ​

❌ Plain version:

Quantized a Llama-70B model to INT8 to save memory. Used GPTQ with a 128-sample calibration set. Model size went from 140GB to 35GB.

✅ Strong version:

Llama-70B INT8 quantization — single-GPU inference

  • Situation: the business needed Llama-70B on a single H100 80G; FP16 inference requires 140GB, forcing dual-card TP=2 — unacceptable hardware cost and cross-GPU communication overhead.
  • Task: INT8 weight-only quantization so 70B runs on one H100, with throughput no more than 15% below FP16 and perplexity up no more than 1%.
  • Action: ① compared GPTQ / AWQ / SmoothQuant on Llama-70B + WikiText-2 perplexity — AWQ with group_size=128 had the smallest perplexity increase (+0.4%); ② modified vLLM's AWQLinearMethod to support group-wise quantization and landed a PR (vllm-project/vllm#12345) merged into mainline; ③ profiled INT8 operators with Nsight Compute, found the dequantize kernel was the bottleneck, and wrote a Triton fused gemm + dequant kernel, +30% throughput over the original.
  • Result: weights from 140GB to 35GB (-75%), 70B runs on a single H100; throughput +8% versus FP16 dual-card TP=2 (no cross-GPU communication); compared with FP16 single-card OOM, went from "won't run" to "runs"; perplexity +0.4%, and blind business-side testing perceived no difference.

Follow-up points: ① Mechanistic differences between GPTQ and AWQ? ② Why does LLM inference mainly use weight-only rather than weight+activation? ③ How did you write the Triton fused kernel? ④ How did you pick group_size?

Example 3: Scaling a large inference service ​

❌ Plain version:

Responsible for scaling the inference service from 4 replicas to 16 replicas to handle peak traffic. Used K8s HPA for autoscaling.

✅ Strong version:

LLM inference service scaling and SLO assurance — Double 11 shopping-festival inference support

  • Situation: during Double 11, inference QPS was estimated to peak at 4× daily load; the original 8-replica K8s deployment ran at 95% GPU utilization at full load with P99 spiking to 8s; scaling budget capped at 16 replicas on A100 80G, and hardware procurement lead time was 6 weeks with no short-term replenishment possible.
  • Task: carry a 1000-QPS peak on the existing 16 replicas, keep P99 < 2s, zero P0 incidents.
  • Action: ① with vLLM --gpu-memory-utilization=0.92 + max_num_seqs=256, one replica carried 65 QPS (up from 50); ② changed the Triton Server config to enable dynamic_batching, tuning max batch delay from 5ms to 2ms to reduce queue buildup; ③ wrote a traffic scheduler (OpenResty + Lua) routing requests to "long-context" and "short-context" replica groups by prompt length — long prompts went to chunked-prefill replicas, short prompts to throughput-optimized replicas; ④ added a Prometheus alert: queue > 50 for 30s → automatically scale out 2 replicas.
  • Result: P99 held at 1.8s at a 1100-QPS peak, GPU utilization 88%, zero P0s; saved 4 replicas versus the estimate during the festival, RMB 120k monthly cost savings. This scheduling policy was distilled into an internal "traffic tiering" component reused by two other business lines.

Follow-up points: ① Why did you write your own scheduler instead of using HPA? ② How did traffic tiering prevent short-prompt replicas from starving long prompts? ③ How did you pick the queue > 50 threshold? ④ Did any alerts fire during the festival?

Example 4: vLLM source contribution ​

❌ Plain version:

Contributed code to the vLLM open-source project and fixed several bugs.

✅ Strong version:

vLLM open-source contribution — PagedAttention block_table index optimization

  • Situation: during internal Llama-70B long-context (16k+) inference, PagedAttention showed performance jitter at certain prompt-length combinations, with TTFT spiking 3×.
  • Task: root-cause the jitter, fix it, and contribute the fix back to vLLM mainline.
  • Action: ① traced with Nsight Systems and located cache-miss spikes in the attention_decode kernel when block_table indices exceeded 256; ② read the vLLM source and found block_table indices used int32 instead of int16, dropping L2 cache hit rate from 85% to 62% on long prompts; ③ changed the block_table element type to int16 (safe when block count < 32768) with a fallback path; ④ validated on vLLM 0.8 with benchmarks — jitter gone, overall throughput +4%; ⑤ opened PR #12345 against vllm-project/vllm with the benchmark report and regression tests; merged into mainline after 3 weeks.
  • Result: TTFT jitter eliminated for internal 70B inference; PR merged, affecting all vLLM releases after 0.8; GitHub link at the top of the resume.

Follow-up points: ① How did you trace the cache miss back to the block_table index? ② Why does int16 need a fallback beyond 32768 blocks? ③ What did the PR review process look like? ④ Did the vLLM team extend your approach afterward?

Example 5: On-device LLM deployment ​

❌ Plain version:

Ran Llama-3.2-1B on an iPhone with the llama.cpp framework. Quantized the model to INT4 for offline inference.

✅ Strong version:

iPhone on-device Llama-3.2-1B — offline assistant prototype

  • Situation: the team was exploring an "AI assistant without network" prototype, targeting a 1B model on iPhone 14 Pro or better; zero hardware budget (no cloud-inference fallback allowed); first-token latency < 500ms, generation speed > 15 tokens/s.
  • Task: run a 1B model on iPhone 14 Pro with llama.cpp + INT4 group-wise quantization, packaged as a distributable app prototype.
  • Action: ① compared llama.cpp / MLC-LLM / CoreML+MPS — llama.cpp was fastest on M2 but its Metal backend was immature on iPhone; MLC compilation was complex but could use the ANE; chose llama.cpp + Metal for shorter dev cycle and the most mature ecosystem; ② quantized with llama.cpp's quantize tool to INT4 group-wise (group_size=32) — model from 2.2GB to 480MB (-78%), load time from 4s to 1.2s; ③ found the Metal backend's matmul at batch=1 reached only 8% of A100 performance, with the bottleneck in Metal kernel threadgroup malloc; ④ wrote a patch replacing malloc with threadgroup async copy (referencing Metal Best Practices), cutting first-token latency from 1200ms to 380ms.
  • Result: Llama-3.2-1B INT4 inference on iPhone 14 Pro at 18 tokens/s, 380ms first token, 480MB model; the prototype app entered internal testing and was recognized by the product team as "usable offline." Later submitted a PR to llama.cpp's Metal backend (ggml-org/llama.cpp#6789).

Follow-up points: ① Why llama.cpp instead of MLC? ② How did you pick group_size=32 for INT4? ③ How did you discover the Metal threadgroup malloc bottleneck? ④ How do edge-inference bottlenecks differ from server-side?

Example 6: Distributed inference optimization ​

❌ Plain version:

Responsible for TP / PP deployment of large-scale LLM inference using the vLLM distributed version. Ran 70B models on 8×H100.

✅ Strong version:

Llama-70B cross-node distributed inference — multi-node TP optimization

  • Situation: the business needed Llama-3-70B deployed; a single 8×H100 node couldn't hold it (70B FP16 = 140GB; 8×80G = 640G per node, but KV cache needs extra space), requiring a 2-node 16-GPU deployment; NVLink existed only intra-node across the 8 GPUs, with 200GB/s InfiniBand across nodes.
  • Task: cross-node TP inference with throughput ≥ 70% of single-node TP=8 and TTFT ≤ 1.5s.
  • Action: ① ran the baseline with vLLM tensor_parallel_size=16 and found cross-node all-reduce took 38% of forward time; ② traced with Nsight Systems and found NCCL cross-node traffic went over PCIe instead of IB (a driver configuration issue); setting NCCL NET_GDR_LEVEL=5 switched it to RDMA, cutting communication time 60%; ③ enabled speculative decoding + EAGLE (acceptance rate 0.6), reducing forward passes 40%; ④ enabled CUDA Graph to cut kernel-launch overhead, shaving another 8% off single-forward time.
  • Result: 2-node 16-GPU TP throughput reached 88% of single-node 8-GPU TP=8 (beating the 70% target), TTFT 1.2s, 32% monthly cost savings versus "naive dual-node TP=16"; the cross-node optimization package became the internal standard config for multi-node inference.

Follow-up points: ① Why is cross-node TP slow? ② How did you tune NCCL NET_GDR_LEVEL? ③ How did you measure the EAGLE acceptance rate of 0.6? ④ What pitfalls hit you during CUDA Graph capture?

5. Listing Your Tech Stack: Layer It, Don't Pile Nouns ​

Skill lists are where "noun piling" concentrates: one line of twenty tool names tells the interviewer at a glance that none were used deeply. The right approach: layer + label proficiency + tie to evidence.

1. Four layers: languages / engines / kernels / platform ​

LayerContentExampleHow to label
LanguagesProgramming languagesPython (expert), C++ (proficient), Go (familiar), Rust (familiar)Proficiency must be honest
Engines / frameworksInference engines and training frameworksvLLM (expert, source-level), TensorRT-LLM (proficient), PyTorch (proficient), Triton Server (proficient)Label "source / API / heard of it"
Kernels / hardwareCUDA / Triton / NsightCUDA (familiar — can write reduction), Triton DSL (familiar), Nsight Systems (expert), FlashAttention (know the principles)Label depth
Platform / toolsEngineering and infrastructureDocker, Kubernetes, Prometheus, ArgoCD, Helm, LinuxOnly list what you've actually used in production

The point of layering is to let the interviewer see your competency structure at a glance: languages are the foundation, engines are daily tools, kernels / hardware are optimization ability, platform indicates production-readiness. Among the four layers, "engines / frameworks" and "kernels / hardware" are worth the most — they map directly to JD capability items and to where your project evidence lives.

2. Self-consistency rules for proficiency labels ​

  • Pick exactly one label: "expert / proficient / familiar." Avoid mushy words like "familiar with" and "skilled in."
  • "Expert in vLLM" must mean you can talk source code: it means you can walk through the implementation details of vLLM's Scheduler / Executor / block_manager and land merged PRs; before writing "expert," ask yourself if you'd survive a 30-minute source-level grilling.
  • Proficiency must be consistent with project evidence: write "proficient" for what you used in projects — the interviewer can verify by opening the repo; write "familiar" for what you only studied.

The cost of noun piling: one puncture zeroes the page

"Familiar with vLLM, TensorRT-LLM, SGLang, LMDeploy, TGI, ONNX Runtime, OpenVINO, Triton Server, TensorFlow Serving..." — the interviewer picks an obscure one, asks source-level details, you blank, and every remaining "familiar" gets a question mark. Keep the skills list short rather than long: five items that survive follow-ups beat fifteen. Authenticity > coverage — this law holds for the entire job search.

6. Depth vs. Breadth: What to Emphasize per Role ​

"Well-rounded" is an insult on a resume — it means no side is deep enough. A resume must have a main attack direction, and the direction is decided by the target role.

Role typeEmphasizeShow secondarilyHow it shows on the resume
Inference Engine EngineerEngine source + kernel optimization depthServing abilitySource modifications, PRs, benchmark comparisons in projects
LLM Systems EngineerServing breadth + SLO dataEngine source awarenessK8s, canary, monitoring & alerting, incident handling in projects
Inference Optimization EngineerPerformance-tuning cases + kernel experienceDeployment abilityNsight cases, quantization, kernel rewrites in projects
ML Infra / MLOpsPlatform ability + multi-model schedulingSingle-model extreme optimizationModel registry, CI/CD, canary, cost in projects
Kernel EngineerCUDA / Triton depth + hardware understandingEngine usageKernel rewrites, SASS analysis, bank conflicts in projects
Edge Inference Engineerllama.cpp / MLC + hardware adaptationServer-side inferenceOn-device deployment, quantization, platform adaptation in projects
Inference Framework R&DFramework architecture + paper reproduction + scheduling algorithmsKernels + systems, bothNew scheduling algorithms, new memory management in projects

The one-sentence principle

Depth comes from "digging one point to the bottom"; breadth comes from "walking one full pipeline end to end." The projects on your resume should show: at least one project where you dug to the source / kernel / SASS level (depth), and every project covering a complete chain (breadth). One project that is "both deep and complete" beats five that each stop at the vLLM API call. This is exactly the "1 main line + 1 auxiliary line" structure in Portfolio Projects.

7. Ten Fatal Resume Flaws ​

The ten problems below are ranked by kill power — each one gets you screened out on its own. The full engineering-side pitfall checklist lives in Common Pitfalls and Anti-Patterns — that page is project-side; this one is resume-side.

Counterexample 1: noun piles without explanation ​

❌ "Proficient in vLLM, TensorRT-LLM, SGLang, LMDeploy, TGI, ONNX Runtime..." ✅ Replace with a layered skills section; keep only what survives follow-ups, and attach the two or three most relevant items to project evidence.

Counterexample 2: courses but no projects ​

❌ Five online courses listed under education; the project section is empty. ✅ Compress the courses into one background line and pour the time into Portfolio Projects — even a single vLLM PR beats ten courses.

Counterexample 3: performance numbers that don't survive follow-ups ​

❌ "TTFT improved 50%," "3× throughput" — relative to what? What context? What batch size? What sequence length? ✅ Give the comparison baseline and context: "TTFT from 1.8s to 600ms (vLLM 0.5 → 0.8 + chunked prefill, batch=64, seq=2k, single A100 80G)."

Counterexample 4: claiming others' projects as your own ​

❌ A teammate's vLLM patch copied verbatim as "my responsibility." ✅ Label your role and contribution scope truthfully; in a deep-dive you must be able to say which line of block_table changed — only write what you can speak to.

Counterexample 5: tools without judgment ​

❌ "Deployed Llama with vLLM; optimized performance with TensorRT." ✅ Add a "why" behind every tool: "chose vLLM 0.8 + chunked prefill for the 70B + long-prompt scenario — TTFT 35% lower than TGI."

Counterexample 6: messy formatting ​

❌ Five font sizes, three alignments, colored icons, garbled PDF export, typos. ✅ Single- or double-column template with consistent fonts; emphasize with bold instead of color; proofread the spelling three times before export — a typo is close to an automatic rejection for this role ("attention to detail" is itself the evidence).

Counterexample 7: long-winded with no focus ​

❌ Two and a half pages, eight-line paragraphs, unrelated experience (a frontend internship, a sales part-time job) all included. ✅ One page for new grads, two max for experienced hires; no experience block over 5 lines; delete anything unrelated to the target role.

Counterexample 8: only good news ​

❌ "The service ran stably after launch" — not a word about rollbacks, incidents, or pitfalls. ✅ Write one honest retrospective: "the first version set max_num_batched_tokens too small, causing long-prompt queue buildup; recovered after switching to a dynamic value" — the interviewer wants not perfection but authenticity and reflection.

Counterexample 9: no personal keywords ​

❌ After reading the whole page, the interviewer couldn't say whether you're "the vLLM-source person" or "the K8s-deployment person." ✅ Lock the main direction with a three-line "personal positioning" under your name on the first screen, and converge everything else toward it.

Counterexample 10: no per-application tailoring ​

❌ The same resume mass-mailed to 50 openings; the JD's "familiar with speculative decoding" has no matching evidence anywhere in your resume. ✅ Follow the Section 8 process: reorder capability items and adjust project detail per target role.

The most expensive mistake isn't any of the above

The ten above are all "bad writing"; the most expensive is "faking." Resume fraud (fabricated degrees, invented projects, stolen PRs) is exceedingly easy to uncover in the small inference & deployment circle — GitHub repos, PR history, levels.fyi records, background checks: everything is checkable. Getting caught once costs you the industry's network and reputation. A thin resume is forgivable; a false one is not.

8. Application Strategy: Tailor per JD, Turn the Portfolio into Clickable Evidence ​

A finished resume is only half the job; the other half is "applying." The three strategies in this section all serve one goal: let the screener see in three seconds that "this person was built for this role."

1. Per-JD tailoring: ten minutes of customization ​

Never mass-mail one resume. Tailoring is not fabrication — it's shifting the narrative's center of gravity:

  • Reorder evidence: whatever the JD's first requirement is, the matching evidence goes on the resume's first screen. Applying to engine roles? Lead with the vLLM PR. Systems roles? Lead with the K8s deployment.
  • Mirror the keywords: the JD says "continuous batching," you write continuous batching; the JD says "in-flight batching," you use its wording — keyword alignment measurably raises screen-pass rates (many companies screen by keywords).
  • Cut the irrelevant: anything the JD never mentions gets deleted or pushed down unless it's spectacular.

Wherever a link can go on an inference & deployment resume, put one: GitHub, vLLM PR pages, benchmark repos, technical blogs. Three cautions:

  • Links must be accessible (checked) and substantive (complete README, runs as documented) — a dead GitHub link is worse than none;
  • The completion bar and presentation of the portfolio live in Portfolio Projects; before interviews, walk it against the self-check list in Common Pitfalls and Anti-Patterns;
  • Point to the portfolio proactively in project descriptions: "Full benchmark code and Nsight reports in the GitHub link" — every click by the interviewer is a leap in your credibility.

Special bonus for this role: put your open-source PR links in the most visible spot on the resume. One PR merged into vLLM mainline is harder than any project description.

3. Pacing and data-driven retrospectives ​

Run applications like a project with metrics:

  • Track in a spreadsheet: role, JD keywords, application date, tailored or not, interview or not, interview feedback;
  • If 20 applications in a row yield zero interviews, suspect the resume before the market — redo the skills benchmarking from Section 2, and have an insider friend run a "five-second screen" test;
  • If you get interviews but fail them, write down the questions, remediate against the Interview Question Bank, and revise the resume — every point you couldn't answer under follow-up is a point the resume must rewrite.

A clarification on "the portfolio must be deployed online"

Many tutorials demand that a portfolio be deployed to a live server. Deployment is a bonus, not a requirement: a repo that clones, runs in three steps per the README, and ships a complete benchmark report is already enough to support a deep-dive; a deployed demo environment is the cherry on top. Don't burn the time you should spend polishing vLLM source details and Nsight reports on deployment.

9. Resume Template Skeleton: A Ready-to-Use Markdown Template ​

Below is a template skeleton covering all the principles above. Replace the [...] annotations with your own content. Keep every "why" and "comparison baseline" — they are the whole difference between this and a chronicle.

markdown
# Name | Target role: Inference Engine Engineer

> Phone | Email | GitHub (PR links) | Tech blog | City
> One-line positioning: 3 years of LLM inference engine development, focused on vLLM
> secondary development and kernel optimization; led cross-node TP deployment of a 70B
> model; contributed 3 merged PRs to vLLM mainline.   [personal keywords — counterexample 9]

## Education
- **XX University · M.S. in Computer Science** (20XX.09 – 20XX.06)
  - Focus: high-performance computing / GPU programming; GPA: x.x/4.0 (top 20%)
  - Relevant courses: computer architecture, parallel computing, operating systems, advanced C++

## Projects (ranked by value; lead with the one most relevant to the target JD)
### Project name | one-line positioning ("inference optimization of model X in scenario Y")
- **Situation**: business scenario / model size / hardware / QPS / SLO / budget constraints.
- **Task**: optimize metric X (baseline Y, from an older version / another engine / the
  team's previous numbers).
- **Action**: 2–3 bullets of "what I did + why this choice + what pit I hit" —
  e.g., engine comparison, source modification, kernel rewrite, Nsight diagnosis.
- **Result**: latency (with baseline) + throughput + GPU utilization + business metrics
  (with context); write an honest retrospective.
- Tech stack: Python / C++ / vLLM 0.8 / Triton / Nsight / K8s   [must echo the skills section]

### Project 2 | ... (same structure; if it's an open-source contribution, attach the PR link)

## Internship / work experience (if any)
### Company · Inference Engineer (20XX.06 – 20XX.09)
- Owned module XX: what I did, why, and the result — same STAR discipline.

## Open-source contributions / papers / blogs (second tier of credibility)
- vLLM mainline PR: vllm-project/vllm#XXXXX (block_table index optimization, TTFT jitter -100%)
- Technical blog: "PagedAttention Internals + Measured Comparison" — link
- Paper "..." (submitted/published at XX conference/journal)

## Skills (layered; short rather than long; must survive follow-ups)
- **Languages**: Python (expert), C++ (proficient — can read CUTLASS templates), Go (familiar)
- **Engines / frameworks**: vLLM (expert at source level, 3 merged PRs), TensorRT-LLM (proficient), PyTorch (proficient), Triton Server (proficient)
- **Kernels / hardware**: CUDA (familiar — can write reduction / scan), Triton DSL (familiar), Nsight Systems (expert), FlashAttention (know FA1/2/3 principles)
- **Platform / tools**: Docker, Kubernetes, Prometheus, ArgoCD, Helm, Linux (proficient)

Two usage notes for the template

① If you're a new grad with no internships, move "Projects" up, delete the "Internship" section, and let the portfolio fill the gap — but never delete the quantification and baselines in the Result line; that's the lifeline of the whole page. ② After writing each project, self-test once: read every line aloud to someone who doesn't know the project — can they repeat back "what you did and how well"? If not, rewrite.

10. Further Reading ​

Continue on this site:

External real resources (verifiable primary sources only):

One last sentence: a resume for an inference & deployment role is not a work of art for others to admire — it is a checklist of question marks left for the interviewer, where every mark points to a "source code + data + retrospective" story you can talk about for 20 minutes. Time spent on a resume is never wasted when it's spent on "figuring out what I did, how well I did it, and why I did it that way."