Skip to content

Module Guide and Job Landscape

At a glance A panoramic view of inference acceleration roles — from Silicon Valley to China. Responsibilities, competency models, and salary bands for seven role families (engine engineer, systems engineer, optimization engineer, infra/MLOps engineer, kernel engineer, edge engineer, framework R&D), plus three persona profiles and a map of the four career-module pages.

This page contains time-sensitive content. Data is current as of 2026-08; engine versions, benchmark rankings, and product features may have changed since — verify against primary sources before citing.

Module Guide and Job Landscape ​

Start by memorizing one sentence for this page: the career module solves two things — first, see how many roles the inference & deployment market actually has and what each one demands; second, prepare yourself against the gaps. This page is the navigation map for both jobs: the job landscape gives you market intelligence, the three persona profiles help you locate yourself, and the content map tells you which four pages to dig into.

The most dangerous sequence in a job search is "grind first, position later": you bury yourself in CUDA programming for three months, only to find the role you were aiming at actually wants vLLM secondary development. The career module pulls you back on track — do the information work before the learning work. Inference & deployment roles have a wider skill radius than traditional ML (they span systems, kernels, serving, and hardware), which makes "learn first, aim later" even less forgivable here.

1. Module Positioning: Read the Market, Then Prepare Yourself ​

The module's four pages chain into a job-search pipeline, where each stage's output is the next stage's input:

Your situation: job hunting / switching into inference acceleration or model deployment
         │
         ▼
┌──────────────────────────────────────────────────────┐
│ ① Read the market — JD List                           │
│    What roles are open? ByteDance, Alibaba, OpenAI,   │
│    Groq... What the vLLM / TensorRT / CUDA in JDs     │
│    actually mean; salary bands, level bars, weights   │
│    of bonus items                                     │
└──────────────────────────────────────────────────────┘
         │
         ▼
┌──────────────────────────────────────────────────────┐
│ ② Read yourself — Deconstructing JD Knowledge Points  │
│    Map JD skill words to "can do / can't / half-can", │
│    and generate an L1→L2→L3 catch-up route instead of │
│    reading everything cover to cover                  │
└──────────────────────────────────────────────────────┘
         │
         ▼
┌──────────────────────────────────────────────────────┐
│ ③ Package yourself — Resume Analysis                  │
│    Rewrite project experience from the interviewer's  │
│    seat: latency, throughput, batch size, speedup,    │
│    and the pits you fell into                         │
└──────────────────────────────────────────────────────┘
         │
         ▼
┌──────────────────────────────────────────────────────┐
│ ④ Prove yourself — Interview Question Bank            │
│    Metrics / quantization / KV cache / speculative    │
│    decoding / system design / whiteboard KV cache,    │
│    for self-testing                                   │
└──────────────────────────────────────────────────────┘

This order wasn't invented by this site — it is the backbone of the Learning Paths: Three Routes sprint: start with the end in mind and back-solve your learning from JDs and interview questions, aiming not for full coverage but for complete coverage of high-frequency exam points. With only two or three weeks, follow this main line; with more time, cross it with the site's regular path of "guide → core concepts → case studies → practice".

How this module relates to the rest of the site

The career module does not re-teach knowledge; it does only two things: market intelligence + job-search method. For concepts, go to What Is Inference Acceleration? to nail down the boundary, and to the Glossary for definitions. The most common mistake at the job-hunt stage is spending time on "one more chapter of CUDA" instead of "organizing the PagedAttention you already know" — the career module is built for the latter.

2. Job Landscape: A Panorama of Seven Inference & Deployment Roles ​

Skim the master table first, then we will unpack each role. Note: the taxonomy here follows "job character," not "job title" — the same role can carry completely different titles at different companies.

RoleCommon English titleOne-line positioningMain deliverables
Inference Engine EngineerInference Engine EngineerSecondary development on vLLM / TensorRT-LLM / SGLangEngine patches, custom kernels, scheduling policies
LLM Systems EngineerLLM Systems EngineerTurn open-source LLMs into stable servicesLive models, autoscaling, monitoring & alerting
Inference Optimization EngineerInference Optimization EngineerPerformance tuning, quantization, kernel selectionCases that halve latency or double throughput
ML Infra / MLOps EngineerML Infra / MLOps EngineerModel serving-ization, CI/CD, platformizationModel platforms, inference gateways, canary systems
Kernel EngineerKernel EngineerHand-written CUDA / Triton kernelsCustom attention / quantization kernels
Edge Inference EngineerEdge Inference EngineerMobile / embedded / browser inferencellama.cpp / MLC / WebGPU deployments
Inference Framework R&DInference Framework R&DArchitecture and evolution of the engine itselfNew scheduling algorithms, new memory management, new execution modes

How to read the salary numbers (please read this first)

Salaries in this section are coarse, composite ranges compiled from public job postings and third-party statistics (levels.fyi and similar) as of the time of writing (dataAsOf: 2026-08). They are order-of-magnitude references for choosing a direction, not precise offers. "Tier-1 Chinese cities" means full-time annual salary in RMB in Beijing, Shanghai, Shenzhen, Hangzhou, and similar markets (covering common levels); "overseas" means US tech companies (USD). Inference & deployment pay is more polarized than traditional ML: engine engineers who can modify vLLM source code are extremely scarce and priced accordingly, while engineers who merely know how to deploy and call models hit a salary ceiling fast. Pay for the same role can vary 2–3× with education, experience, company, and interview performance; treat anything older than one quarter as a trend reference, and defer to actual offers and the latest statistics.

1. Inference Engine Engineer — the Engine "Secondary Developer" ​

One of the scarcest roles at Chinese tech giants and frontier labs. The job is to modify open-source inference engines (vLLM / TensorRT-LLM / SGLang / LMDeploy) at the source level: patching in custom attention variants, extending scheduling policies, wiring up new model architectures, and fixing performance blowups at specific shapes. It demands that you have actually read vLLM's scheduler source and know how PagedAttention's block table is maintained — not that you merely know how to run python -m vllm.entrypoints.api_server.

  • Core skills: production-grade Python + systems-level C++; PyTorch internals (autograd / dispatch); source-reading experience in at least one inference engine; CUDA fundamentals; Linux performance analysis (perf / nsys / Nsight Compute).
  • Typical pay: roughly RMB 500k–1.2M per year in tier-1 Chinese cities (core teams at LLM startups pay a clear premium); roughly $200k–400k per year overseas (L5+ at teams like OpenAI / Anthropic / Meta infra).
  • Map to this site: source-level understanding in vLLM, TensorRT-LLM, Speculative Decoding and Medusa/EAGLE; scheduling and batching in Batching and Request Scheduling.

2. LLM Systems Engineer — Turn Models into Services ​

Often confused with the engine engineer, but the division of labor is different: the engine engineer modifies source code; the systems engineer deploys models into stable services using off-the-shelf engines. Responsibilities include model loading and weight management, inference service wrapping (FastAPI / Triton), autoscaling and traffic routing, monitoring and alerting (TTFT / TPOT / queue length), release and canary management, and cost accounting plus GPU utilization optimization. It is a cross between SRE and ML engineer.

  • Core skills: Python + Go (some teams); Kubernetes and container orchestration; productionizing with Triton / vLLM / TGI; Linux + networking; observability (Prometheus / Grafana / OpenTelemetry); GPU resource management.
  • Typical pay: roughly RMB 400k–800k per year in tier-1 Chinese cities; roughly $150k–280k per year overseas.
  • Map to this site: serving and orchestration in Model Serving and Orchestration, Triton Inference Server; distributed inference in Distributed Inference (TP/PP).

3. Inference Optimization Engineer — Find Bottlenecks, Fix Kernels, Write Quantization ​

The role laser-focused on the word "performance." Responsibilities: locate inference bottlenecks, then choose and implement optimizations — kernel fusion, KV cache optimization, quantization (INT8 / FP8 / INT4 weight-only), speculative decoding, CUDA Graph capture, batching policy tuning. It demands both the ability to find bottlenecks with Nsight and the ability to drop to the CUDA level to write or modify kernels — a hybrid of "systems engineering + kernel skills."

4. ML Infra / MLOps Engineer — Platforms and Pipelines ​

Provides the "model-to-production" infrastructure for algorithm teams: model registries, training / inference pipelines, feature platforms, experiment management, canary and rollback, cost and quota management. The inference & deployment slice of this role is mainly multi-model scheduling, model version management, and inference gateways — extreme optimization of a single model is left to the inference optimization engineer.

5. Kernel Engineer — Squeeze Nanoseconds at the CUDA Level ​

Roles dedicated to writing CUDA / Triton / CUTLASS kernels, found mostly in the infrastructure groups of LLM teams, SDK teams at chip companies, and inference groups in autonomous-driving / multimodal teams. The work: hand-write or optimize FlashAttention variants, quantization kernels, GEMM, reduction, scan, and more; read Nsight Compute reports; understand Tensor Core / WGMMA / TMA instructions; and reason backward from PTX assembly to bottlenecks.

  • Core skills: CUDA / C++ / CUTLASS; Triton (OpenAI / Triton-Lang); PTX / SASS reading; Tensor Core programming (mma / wgmma); TMA (Hopper+); fusion strategies; memory alignment and bank conflicts.
  • Typical pay: roughly RMB 600k–1.5M per year in tier-1 Chinese cities (H100-class kernel engineers command extreme premiums at top companies); roughly $220k–450k per year overseas (NVIDIA CUDA team, Groq, hardware startups).
  • Map to this site: GPU Architecture and Optimization, Kernel Fusion and Custom Kernels, Computation Graph Optimization, Hardware Primer.

6. Edge Inference Engineer — Put an LLM in a Phone ​

A role created after 2024 as open-source small models matured (Phi-3, Qwen2.5-0.5B/1.5B, Llama-3.2-1B/3B). The job: run LLMs on phones / embedded devices / browsers — llama.cpp porting, MLC-LLM compilation, WebGPU / Vulkan backends, Metal on Apple Silicon, NEON optimization on ARM CPUs. It demands a double competency of "kernel-level understanding + edge engineering."

  • Core skills: C / C++ (embedded-grade); llama.cpp / ggml source; MLC / Apache TVM; Metal / Vulkan / OpenCL / WebGPU; ARM NEON / SVE; model quantization (INT4 / INT8 / INT4 group-wise).
  • Typical pay: roughly RMB 350k–700k per year in tier-1 Chinese cities (phone makers / algorithm startups); roughly $140k–250k per year overseas (Apple ANE team, Google Pixel, Meta Llama On-Device).
  • Map to this site: llama.cpp and GGUF, Mobile Deployment, Model Quantization Fundamentals.

7. Inference Framework R&D — Write the Next vLLM ​

Rare but top-tier: own the architecture and evolution of the inference framework itself, mostly in core teams of vLLM / SGLang / LMDeploy / TensorRT-LLM, OpenAI's Triton inference team, and platform teams at Anyscale / Modal / Together AI. Responsibilities include designing new scheduling algorithms, new memory management, new execution modes (e.g., Chunked Prefill, Disaggregated Prefill / Decode), and cross-hardware abstraction layers. This is a trinity of "research + engineering + systems."

  • Core skills: distributed-systems fundamentals; full-stack C++ / Rust / Python; CUDA / ROCm / TPU abstraction layers; attention frontiers (FlashAttention-3, Mamba, Hybrid); scheduling algorithms; paper reading and reproduction.
  • Typical pay: roughly RMB 800k–2M per year in tier-1 Chinese cities (startup core teams + equity); roughly $300k–600k per year + equity overseas (OpenAI / Anthropic / Anyscale and similar).
  • Map to this site: Distributed Inference (TP/PP), Batching and Request Scheduling, Classic Papers in Depth, Frontier Advances.

3. Role Skill Radar: Systems · Kernels · Hardware · Engineering ​

Plot the seven roles on four axes and the differences become obvious. Axis definitions:

            Systems (scheduling / memory / serving / distributed)
                      ▲
                     /|\
                     │
     Kernels ──────────┼────────── Hardware
  (CUDA / Triton /    │        (GPU architecture / Tensor Core /
   FlashAttention /   │         TPU / HBM / NVLink)
   quantization       │
   kernels)           │
                     \|/
            Engineering (CI/CD / K8s / observability / collaboration)

Weights are 1–5 per axis (1 = barely needed, 5 = core of the role):

RoleSystemsKernelsHardwareEngineeringRole character
Inference Engine Engineer5433Engine-internals type
LLM Systems Engineer4225Deployment-ops type
Inference Optimization Engineer4443Performance-tuning type
ML Infra / MLOps3225Platform-infra type
Kernel Engineer3552Kernel-deep-dive type
Edge Inference Engineer3443Edge all-rounder type
Inference Framework R&D5443Framework-architecture type

How to use this table

Don't just stare at the biggest number in a row — ask which column you can stomach working in for years: systems-strong people tolerate the grind of scheduling-algorithm details; kernel-strong people enjoy shaving nanoseconds against PTX; engineering-strong people need the reassurance of watching a service run stably; hardware-strong people love reasoning about bandwidth from architecture diagrams. Rank "which kind of work I would rather do long-term" above "which role is hotter." Kernel engineers earn more but walk a narrower path; MLOps roles are plentiful but have a low ceiling — face that trade-off honestly.

4. Three Typical Personas ​

Reading JDs is easy; imagining what doing the job actually feels like is hard. Here are three personas that make the differences concrete down to the daily level.

Profile A: Engine Engineer · Lin ​

Lin is an inference engine engineer at an LLM startup, 3 years in, previously a PyTorch backend engineer. She gets to the office at 10 a.m. and first checks last night's vLLM performance benchmark report: Qwen2.5-72B on 4×H100 shows TTFT 8% slower than last week. She opens the Nsight Systems trace, finds that a newly added attention path doesn't go through FlashAttention-3, and by the afternoon has a PR up for vLLM swapping flash_attn_varlen_func for flash_attn_3_varlen_func, with a set of comparison benchmarks attached. At 9 p.m. CI finishes and it merges into the internal fork. Her weekly report always contains numbers like "latency dropped from X to Y."

She has read far more code than she has written — vLLM's scheduler, block manager, and model runner add up to tens of thousands of lines, and she can explain "what happens on the GPU during one forward pass," from the H2D copy to kernel launch to attention to all-reduce. Her resume highlight: "Contributed 6 PRs to vLLM, including a PagedAttention block_table indexing optimization that cut 72B-model TTFT by 18%."

Profile B: Systems Engineer · Wei ​

Wei is a systems engineer on the LLM platform team of a major Chinese internet company, 5 years in, previously a traditional web backend engineer. He doesn't touch vLLM source code; his job is deploying 40+ open-source and in-house models as callable services. His daily driver is a Grafana dashboard: QPS per model, P99 latency, GPU utilization, queue backlog. One day the P99 of a Llama-3.1-70B service suddenly spikes to 8 seconds; within half an hour he isolates a prompt 28k tokens long that triggers GPU-memory thrash in the prefill phase, scales out two replicas, caps max_num_batched_tokens, and recovers to 1.2 seconds.

His most fluent tools are Helm, ArgoCD, Triton Server's model repository configuration, and vLLM's --gpu-memory-utilization tuning. His resume highlight: "Built a multi-model inference platform sustaining 1000+ QPS; the canary-release mechanism runs 5% of traffic for 30 minutes with no regression before full rollout; GPU utilization raised from 38% to 71%."

Profile C: Infra Engineer · Sam ​

Sam does ML Infra at a Silicon Valley tech giant, 6 years in, an SRE turned MLOps. He doesn't manage models directly; he manages "whether algorithm engineers can push models to production." His day is writing Argo Workflows templates, maintaining the model-registry API, collecting requirements from algorithm teams for new model launches, and running the monthly cost review to tell his boss: "GPT-4o calls cost us $800k this month; we replaced 30% of that traffic with Qwen2.5-72B and saved $240k." Most of his code is Go, with some Python glue.

What he knows well isn't the vLLM source but Triton Inference Server's model.json, KServe's InferenceService CRD, and Kubeflow's pipeline DSL. His resume highlight: "Designed and shipped a multi-model inference gateway supporting canary across 200+ model versions; built a GPU quota and billing system covering 4 business lines, cutting monthly GPU cost by 22%."

What the three personas share

The common keyword in all three resumes is numbers: latency cut by X, throughput up by Y, cost saved by Z, GPU utilization from A to B. A resume for an inference & deployment role without numbers is as good as blank — that is its biggest difference from an algorithm-role resume. Algorithm candidates can still argue "AUC up 3 points"; deployment candidates who don't speak in numbers have nothing left but "used vLLM," which to an interviewer reads as "no experience." The method lives in Skills Benchmarking: What to Highlight on Your Resume.

5. Which Direction Should You Pick: A Decision Checklist ​

Work through the three checklists below and your answers will converge on one or two roles.

1. By Background: Where Are You Standing Now? ​

  • [ ] Python / PyTorch engineering background, have written training pipelines → prioritize Inference Engine Engineer / Inference Optimization Engineer
  • [ ] C / C++ / systems programming background, have written network or storage services → prioritize LLM Systems Engineer / Kernel Engineer
  • [ ] SRE / DevOps background, fluent in Kubernetes → prioritize ML Infra / MLOps Engineer
  • [ ] Embedded / mobile development background, have written ARM NEON → prioritize Edge Inference Engineer
  • [ ] CUDA / HPC background, have written GEMM or reduction → prioritize Kernel Engineer / Inference Framework R&D
  • [ ] Algorithm engineer background, want to move into engineering → Inference Optimization Engineer is the most natural bridge

2. By Interest: Which Kind of "Done" Do You Enjoy? ​

  • [ ] "I turned a vLLM bottleneck into an X× speedup" → engine / optimization direction
  • [ ] "Our service ran at 99.99% availability with zero P0 incidents all year" → systems engineer / MLOps
  • [ ] "I spotted a bank conflict in the SASS disassembly, made it coalesced, and gained 20%" → kernel engineer
  • [ ] "I squeezed a 1.5B model into a phone running 15 tokens/s offline" → edge inference
  • [ ] "My scheduling algorithm lifted cluster throughput by 40%" → inference framework R&D

3. By Trend: How Much Certainty Are You Willing to Pay For? ​

The market structure of inference & deployment roles is shifting fast. Facts worth pricing into your decision (as of dataAsOf: 2026-08):

  • Demand for LLM inference optimization roles is exploding: after vLLM / TensorRT-LLM matured in 2023, every LLM team needs at least one engineer who can modify an engine; estimated new openings nationwide run into the thousands per year;
  • Kernel engineer salaries are leading the market: as H100 / H200 / B200 deployments scale, kernel engineers who understand Tensor Core / TMA / FP8 are seeing 15–25% annual raises;
  • MLOps roles are standardizing and pay is stabilizing: demand is large but the bar is relatively low, and the salary ceiling is becoming clear;
  • Edge inference roles haven't scaled yet: phone-side LLMs are still early, and headcount is smaller than server-side;
  • Domestic-chip inference roles are rising: Huawei Ascend, Moore Threads, Biren, Hygon and other vendors began hiring inference & deployment engineers at scale after 2025 — fast-growing, but with ecosystems less mature than NVIDIA's.

Read the three checklists together: if all three point the same way, commit; when they conflict, prioritize by "background gets you in the door, interest keeps you going, trends only decide how crowded the door is" — background determines whether you can get in, interest determines how far you will go, and trends only affect how crowded the entrance is.

Don't mistake a job title for an identity

The same title can mean completely different jobs at different companies: the "Inference Optimization Engineer" at company A may modify CUDA kernels all day, while the same title at company B may tune vLLM flags all day. Reading the JD always beats reading the title — which is why the first page of this module is the JD List, not a "role encyclopedia."

6. Content Map of This Module: What Each of the Four Pages Does ​

The career module has four pages, mapping to four actions in the job-search flow:

PageOne-line purposeWhen to use
JD ListAggregates real JDs for open inference & deployment roles at major companies, dissecting high-frequency requirements and barsDay one of the job search: first learn what the market wants
Deconstructing JD Knowledge PointsMaps JD skill words to site pages and generates your personal L1→L3 catch-up listAfter reading the market: locate your own gaps
Skills Benchmarking: What to Highlight on Your ResumeAn interviewer's-eye review standard for resumes, with 5–8 rewritten project examplesWhen preparing to apply: turn experience into evidence
Interview Question BankHigh-frequency Q&A with answer frameworks, for self-testingOne week before interviews: final self-test

The order is the main line, not an option

Learning Paths: Three Routes already arranges the four pages into a two-week main line: JD List → Deconstructing JD Knowledge Points → Skills Benchmarking → Interview Question Bank, where each page's output feeds the next. Reading them out of order forfeits most of the value.

7. Two Reminders ​

1. Check the dataAsOf Before Trusting Any Number ​

Everything in this module — roles, salaries, JD requirements, technology keywords — is time-sensitive: the half-life of the inference & deployment field is short; a role's JD may be rewritten twice within a year, and the Tensor Core programming model changes with every GPU generation. Therefore every time-sensitive page in this module carries a dataAsOf (data cutoff month) in the frontmatter, displayed at the top of the page. Before citing any specific number from it, check that date:

  • Within one quarter of today → usable as a current reference;
  • Older than one quarter → treat as a "trend reference" and verify against official sources;
  • Conceptual content (KV cache math, the Roofline model) decays very slowly and is not bound by this rule.

2. The JD vs. the Actual Job: Two Versions of the World ​

A JD describes "the person we hope you are"; day-to-day work is "what we actually need you to be," and there is a systematic gap between the two:

  • The JD says "expert in vLLM source code," but the reality may be mostly operations and firefighting — reading source is a small slice of the work;
  • The JD says "optimize latency," but what gets evaluated is cost and availability — cutting latency 30% while doubling cost may not make the business side happy;
  • The JD says "own it independently," but the reality is one cog in a large inference platform — in an interview, clearly explaining "the part I owned" is far more credible than overselling "I owned the whole system";
  • The JD says "H100 cluster experience preferred," but the team itself has only had H100s for 3 months — discussing your understanding of that uncertainty honestly is itself a plus.

There is exactly one countermove: do reconnaissance on the target team before the interview (interview write-ups, conversations with current employees, engineering blogs). Treat the JD as a hypothesis and verify it with research. The Interview Question Bank includes a category of "questions to ask the interviewer" built precisely for this.

8. Further Reading ​

Continue on this site

References (real sources for verifying salaries and trends)

  • levels.fyi — compensation database for US tech companies, searchable by company / level / city; a reliable source for overseas salaries
  • NVIDIA Careers — official NVIDIA job board, source of TensorRT / CUDA team openings
  • OpenAI Careers — official source for Triton inference and similar roles
  • vLLM Project Contributors — the "open-source resume" channel for engine-engineer hiring
  • [PaperMC / Moonshot / Zhipu Careers] — career pages of Chinese LLM startups; reference for role demand and salary bands