Appearance
JD List: Open Inference & Deployment Roles at Major Companies
This list answers a very specific question: open a job board right now — what are the inference acceleration roles on the market actually asking for?
This page does not reproduce any real JD verbatim. Instead, it generalizes the common patterns in inference & deployment JDs at major Chinese and international companies between 2024 and 2026, organized by "role profile + shared skill set." JDs are fluid and titles drift — a role open today may be frozen next month. What you should take away is a reusable analysis framework, not a list of openings that will expire.
On data timeliness
This page is a generalization of job descriptions publicly visible around the time of writing (dataAsOf: 2026-08) on company career sites and job platforms (Maimai, BOSS Zhipin, Liepin, LinkedIn, levels.fyi). It is not a guarantee of what any company is hiring for right now. For specific openings, locations, headcount, and requirements, always defer to the official career pages in real time. Every "hard requirement / bonus item" below is a cross-JD summary and does not point to any specific company or position.
1. How to Read This List
1. JDs change; the skill combination is the constant
Look at a real case of title drift: in 2021, everyone was hiring "Machine Learning Platform Engineers"; after vLLM went viral in 2023, the same work became "LLM Inference Engineer"; by 2025 it had morphed into "LLM Inference Optimization Engineer" and "Infra Engineer."
How one team's job title evolved over five years (illustrative):
2021 Machine Learning Platform Engineer (deployment track)
2022 Algorithm Engineer — Inference Services
2023 LLM Inference Engineer
2024 LLM Inference Optimization Engineer
2025 Inference Infra Engineer
2026 Senior LLM Serving Engineer
The skill core barely changed: Python + Linux + inference engine + GPU basics.
Only the titles kept moving.Job titles are a thermometer of market sentiment; the skill combination is the long-term skeleton. The first principle of reading JDs: translate the title into a "skill vector" instead of memorizing titles. LLM Inference Engineer, Inference Optimization Engineer, and Serving Engineer may all come from the same group, with 70% overlapping requirements; while two identically titled "Inference Engineers" at a cloud vendor and a startup may work on platformization and kernel optimization respectively.
2. Distinguish "Hard Requirements" from "Bonus Items"
Nearly every JD has two paragraphs: "Requirements" and "Bonus / Preferred." The two paragraphs mean completely different things:
| Hard requirements (requirements) | Bonus items (preferred) | |
|---|---|---|
| What it actually means | Cross this line to get the interview | Nice to have; usually not a blocker |
| Typical content | Degree, years of experience, must-have languages and frameworks | Papers, open-source contributions, specific model experience |
| How it filters | Filter: fail it and you're screened out | Ranker: separates passers |
| Candidate strategy | Check line by line; miss none | Treat as "differentiation material"; one or two good stories suffice |
A common misreading is to stress over bonus items as if they were hard requirements: the JD says "FlashAttention-3 experience preferred," and you decide you can't apply because you've never touched H100. In practice, most bonus items mean "better if present," not "must have." Conversely, if the hard requirements say "familiar with vLLM or TensorRT-LLM source code" and you've never read either, that is a real gap to worry about.
Three signals for telling hard requirements from bonus items
- Wording: "proficient in," "expert in," "X+ years of," "deep understanding of" are usually hard; "experience with XX preferred," "familiarity a plus," "bonus" are bonus items.
- Count: hard requirements rarely exceed 5 and are orthogonal to each other; a JD listing 12 "hard requirements" is really describing an "ideal candidate" — the real bar is lower than the paper.
- Job type: campus-recruiting JDs have few, broad requirements (testing fundamentals); experienced-hire JDs have many narrow ones (testing fit). For new grads, checking themselves line by line against experienced-hire JDs is the most efficient way to talk themselves out of applying.
3. Three Common Misreadings of JDs
- Reading only the title and company name: that's judging a book by its cover. Two "Inference Engineers" may do kernels and scheduling respectively — a bigger difference than "algorithm engineer" vs. "frontend engineer."
- Memorizing skill lists: the JD lists vLLM, so you put vLLM on your resume — then the interviewer asks "have you modified its scheduler?" and you're exposed. Every skill must expand into a project story.
- Looking only at the giants: big-company JDs use standardized wording with fine division of labor, good for calibrating the bar; but the offer you actually land may come from a mid-size company, a unicorn, or a vertical player. OpenAI / Anthropic have extremely high bars, but LLM startups like Moonshot / Zhipu / MiniMax have more inference openings and more room to negotiate.
4. How to Use This List
① Shortlist 2–3 target roles → Pick the 2–3 closest to your background from
Sections 2 and 3 of this page
② Check off against the map → Annotate line by line with
[Deconstructing JD Knowledge Points](/career/knowledge-map):
can do / can't / half-can
③ Build a prioritized catch-up → "Can't" items from hard requirements are P0;
list "half-can" items from bonus items are P1The full methodology of "back-solving learning from JDs" lives in Learning Paths: Three Routes.
2. Typical Roles at Major Chinese Companies
This section covers eight inference & deployment role families that appear steadily on the career pages of major Chinese companies and startups. Each entry gives "company + team + title + core responsibilities + must-have skills + bonus items + level range." Note: the JDs below are cross-company generalizations of similar roles, not the original text or an excerpt of any single real JD.
1. ByteDance · Doubao LLM · Inference Engine Engineer
- Team: Doubao LLM inference team, based in Beijing / Shanghai / Singapore
- Core responsibilities:
- Maintain and extend the internal LLM inference engine (vLLM-based / in-house hybrid), supporting business lines such as the Doubao app, Coze, and the Feishu AI assistant;
- Integrate new model architectures (MoE, long context, multimodal) and implement optimizations such as PagedAttention / Chunked Prefill;
- Tune throughput and latency on H100 / H800 clusters and contribute to capacity planning for thousand-GPU-scale inference services.
- Must-have skills: production-grade Python + systems-level C++; PyTorch internals; source-reading experience in vLLM or TensorRT-LLM; CUDA fundamentals; Linux performance analysis.
- Bonus items: open-source inference engine contributions (vLLM / SGLang / LMDeploy PRs); exposure to FlashAttention / vAttention / TAP; hands-on distributed inference (TP / PP / EP).
- Level range: 2-2 to 4-1 (new PhD grads start at 2-2; senior at 3-1/3-2; team lead at 4-1); compensation RMB 500k–1.5M/year + equity.
- Map to this site: vLLM, TensorRT-LLM, Batching and Request Scheduling, Distributed Inference (TP/PP).
2. Alibaba · PAI Platform · Inference Optimization Engineer
- Team: Alibaba Cloud PAI (Platform for AI) inference team, based in Hangzhou / Beijing
- Core responsibilities:
- Own performance optimization for PAI-EAS (Elastic Algorithm Service) inference services, covering LLM, CV, and recommendation models;
- Drive the rollout of INT8 / FP8 quantization, CUDA Graph, and kernel fusion on the PAI platform;
- Support capacity assurance for major shopping festivals such as Double 11 and 618.
- Must-have skills: Python + C++; hands-on TensorRT or ONNX Runtime; quantization (PTQ / QAT); CUDA / Triton kernel development; Linux + performance analysis.
- Bonus items: hands-on FlashAttention / FlashDecoding; exposure to BladeDISC / XORunner; large-scale GPU cluster operations experience.
- Level range: P6 to P8; compensation RMB 450k–1.2M/year.
- Map to this site: Model Quantization Fundamentals, Kernel Fusion and Custom Kernels, GPU Architecture and Optimization, Inference Benchmarking in Practice.
3. Tencent · Hunyuan LLM · Inference Services Engineer
- Team: Hunyuan LLM inference team, based in Shenzhen / Shanghai / Beijing
- Core responsibilities:
- Deploy and maintain inference services for the Hunyuan model family (open-source + internal), supporting WeChat, QQ, Tencent Docs, and other businesses;
- Evaluate and integrate multi-framework inference engines (vLLM / TensorRT-LLM / in-house framework);
- Tune performance, plan capacity, and troubleshoot, joining major-campaign support such as the "930" shopping festival.
- Must-have skills: Python + Go / C++; hands-on vLLM / TGI / TensorRT-LLM; Kubernetes + Triton Server; Linux + networking; monitoring & alerting systems.
- Bonus items: distributed inference (TP / PP); CUDA / NCCL tuning; elastic autoscaling system design.
- Level range: levels 9 to 11; compensation RMB 400k–1M/year.
- Map to this site: Model Serving and Orchestration, Triton Inference Server, Batching and Request Scheduling.
4. Huawei Cloud · Ascend · Inference Engineer
- Team: Huawei Cloud Ascend AI computing team, based in Shenzhen / Xi'an / Shanghai / Dongguan
- Core responsibilities:
- Deploy LLM inference on Ascend NPUs (Ascend 910/910B/310P) and maintain the MindIE / CANN inference stack;
- Develop and optimize operators for domestic chips (Ascend C / Ascend CL), benchmarking against CUDA kernels;
- Work with algorithm teams to port open-source LLMs (Qwen, ChatGLM, Baichuan) onto the Ascend ecosystem.
- Must-have skills: C / C++; Ascend C or CUDA; PyTorch; Linux; fundamentals of model quantization and operator optimization.
- Bonus items: MindIE / vLLM-Ascend experience; hand-written Transformer operators; familiarity with Huawei's internal ecosystem.
- Level range: levels 13 to 18; compensation RMB 350k–900k/year. Domestic-chip inference roles are rising fast, with openings growing 50%+ annually after 2025 — a real alternative to the NVIDIA path.
- Map to this site: Model Quantization Fundamentals, Kernel Fusion and Custom Kernels, Hardware Primer, Distributed Inference (TP/PP).
5. Moonshot AI · Inference Engine Engineer
- Team: Moonshot AI infra team, based in Beijing / Shanghai
- Core responsibilities:
- R&D on the LLM inference engine behind Kimi Chat, with dedicated optimization for long context (128k → 2M);
- Design in-house inference scheduling policies for the mixed load of ultra-long prompts and high concurrency;
- Partner with the algorithm team to land new model architectures and speculative decoding.
- Must-have skills: Python + C++; PyTorch internals; source-level experience with vLLM or an in-house engine; deep understanding of KV cache / PagedAttention; CUDA fundamentals.
- Bonus items: long-context optimization (YaRN / NTK / Ring Attention); EAGLE / Medusa speculative decoding; hands-on CUDA Graph.
- Level range: level 5 (senior) to level 7 (core); compensation RMB 600k–1.8M/year + equity. Long context is Moonshot's signature, and this line carries visibly more weight in its JDs than at other companies.
- Map to this site: vLLM, Speculative Decoding and Medusa/EAGLE, Batching and Request Scheduling, The GPU Memory Hierarchy and the Bandwidth Wall.
6. DeepSeek · Inference Systems Engineer
- Team: DeepSeek AI infra, based in Hangzhou / Beijing
- Core responsibilities:
- R&D on inference services for the DeepSeek-V3 / R1 model family (a lightweight in-house framework, friendly to the open-source ecosystem);
- Inference optimization for new architectures such as MLA (Multi-head Latent Attention);
- Inference cost control and open-source community contributions (parts of the DeepSeek C++ inference stack are open-sourced).
- Must-have skills: high-performance C++; Python; PyTorch; CUDA; fundamentals of KV cache / attention optimization.
- Bonus items: hands-on MLA / MoE inference; in-house inference engine experience; open-source contribution record.
- Level range: senior engineer to expert; compensation RMB 500k–1.3M/year + equity. DeepSeek's hiring visibly favors engineers who "can write C++ inference kernels" — C++ carries more weight in its JDs than at other LLM startups.
- Map to this site: Kernel Fusion and Custom Kernels, Computation Graph Optimization, vLLM, GPU Architecture and Optimization.
7. Zhipu AI · Inference & Deployment Engineer
- Team: Zhipu AI infra, based in Beijing
- Core responsibilities:
- Deploy inference services for the ChatGLM / GLM-4 model family, covering the public API and on-premise ToB delivery;
- Adapt to multiple hardware platforms (NVIDIA / domestic chips / CPU), including llama.cpp / OpenVINO paths;
- Lightweight deployment and security compliance for on-premise scenarios.
- Must-have skills: Python + Linux; at least one of vLLM / TensorRT / OpenVINO; Docker + Kubernetes basics; model quantization concepts.
- Bonus items: cross-hardware deployment experience (NVIDIA + Ascend + CPU); CPU inference optimization (AVX-512 / AMX); enterprise on-premise delivery experience.
- Level range: senior to expert; compensation RMB 400k–1M/year + equity.
- Map to this site: OpenVINO and CPU Inference, llama.cpp and GGUF, Model Serving and Orchestration, Mobile Deployment.
8. StepFun / MiniMax / ModelBest · Inference Engineer
Inference JDs at these three LLM startups are highly similar, so they are generalized into one entry:
- Core responsibilities: maintain inference services for in-house models, join inference optimization, launch new models, and interface with algorithm and business teams.
- Must-have skills: Python + C++; PyTorch; at least one of vLLM / TensorRT-LLM; Linux; KV cache / batching concepts.
- Bonus items: CUDA / Triton kernels; quantization; speculative decoding; long-context optimization.
- Level range: levels 3–6; compensation RMB 450k–1.2M/year + equity. The character of startup inference roles is "you do everything": deployment, tuning, kernels, and ops blur together — great for fast growth.
- Map to this site: vLLM, Batching and Request Scheduling, Model Quantization Fundamentals, Tuning and Performance Optimization.
3. Roles at International Companies
JD wording differs a lot between Chinese and international postings, but the skill core is highly similar. Below are generalized JD patterns for inference & deployment roles at six representative companies.
1. OpenAI · Triton Inference Engineer / Training Infra
OpenAI's infra organization splits into Training Infra (pre-training clusters) and Inference Infra (the ChatGPT backend inference service).
- Common requirements: Strong C++/Python/Rust; deep experience with GPU inference (CUDA, TensorRT, Triton); experience building production ML serving systems at scale; understanding of distributed systems and networking.
- Reading the signals: in OpenAI's context,
Tritonrefers both to its own Triton Language (a GPU programming DSL) and to inference-serving architecture. The JD barely mentions "machine learning" requirements — it tests systems and kernels. The bar is among the highest of all inference roles, but the target skill set is exactly the content of this site's core concepts chapter and case-study chapter. - Level / pay: L5 starts around $300k/year; L6 around $400k–500k/year + equity.
2. Anthropic · Inference Optimization Engineer
Anthropic's infra team supports training and inference for the Claude model family; the role leans toward "performance optimization" rather than "system building."
- Common requirements: Experience with low-level GPU programming (CUDA, PTX); familiarity with inference engines (vLLM, TensorRT-LLM, TGI); strong systems engineering background; experience with LLM-specific optimizations (PagedAttention, FlashAttention, speculative decoding).
- Reading the signals:
PTXappears more often in Anthropic's JDs than at other companies, reflecting its high bar for "kernel-level understanding."speculative decodingis a priority direction for Claude inference optimization. - Level / pay: L4 starts around $250k/year; L5 around $350k–450k/year + equity.
3. Meta · PyTorch Executor / Inference Engineer
Meta is PyTorch's birthplace, and its roles are the most specialized: Executor, Dispatcher, Quantization, Compiler (torch.compile / Inductor).
- Common requirements: Deep C++ and Python expertise; strong understanding of PyTorch internals (autograd, dispatch, dispatcher); experience with kernel development (CUDA, Triton, CUTLASS); for inference roles, experience with torch.export / AOTInductor / ExecuTorch.
- Reading the signals: Meta's inference roles specifically prefer engineers who understand PyTorch internals (dispatcher / autograd / eager / compile) — a hard requirement that distinguishes Meta from other companies.
ExecuTorchis Meta's flagship on-device inference project. - Level / pay: E4 around $250k–300k/year; E5 around $350k–450k/year; E6 around $450k–600k/year.
4. NVIDIA · TensorRT / TensorRT-LLM Engineer
NVIDIA's TensorRT and TensorRT-LLM teams are "the people who build the GPU inference engines" — the most hardcore version of this work.
- Common requirements: Strong C++ and CUDA; deep understanding of GPU architecture (Tensor Core, TMA, HBM); experience with TensorRT internals; for TensorRT-LLM roles, experience with PagedAttention / In-Flight Batching / FP8 quantization.
- Reading the signals: the JD demands
Tensor Core / TMA / FP8experience directly, meaning the target is extreme optimization on Hopper / Blackwell. NVIDIA rarely hires engineers who "can call TensorRT"; it hires engineers who "can modify TensorRT." - Level / pay: Senior starts around $200k/year; Staff $400k–550k/year + RSU.
5. Groq · Inference Compiler Engineer
Groq is an LPU (Language Processing Unit) startup whose pitch is "extreme-speed inference" (Llama-70B at 1000+ tokens/s).
- Common requirements: Strong C++/Python; experience with compiler backends (MLIR, XLA, TVM); understanding of tensor compilers; for systems roles, experience with high-throughput serving.
- Reading the signals: Groq's JDs lean compiler (MLIR / XLA / TVM), not the traditional CUDA path.
tensor compileris the core keyword. This is the representative of the "compiler track" among inference & deployment roles — different from the mainstream vLLM / TensorRT path, but with extremely scarce skills. - Level / pay: Staff / Senior around $180k–350k/year + equity.
6. Together AI / Anyscale / Modal · Inference Platform Engineer
Silicon Valley "inference-as-a-service" startups, where roles lean toward SaaS platforms rather than extreme single-model optimization.
- Common requirements: Strong Python/Go; experience with Kubernetes and distributed systems; experience building inference platforms (vLLM, TGI); understanding of multi-tenant serving and routing.
- Reading the signals: the JD keywords are
multi-tenant,routing,serverless— a completely different direction from single-model optimization roles. These roles prefer a hybrid background of "distributed-systems engineer + inference engine experience", with low demands on CUDA / kernels. - Level / pay: Senior around $150k–250k/year + equity.
Quick Reference: English JD Keywords
| Keyword | Common meaning | Chinese equivalent |
|---|---|---|
| Inference engine | Inference engines: vLLM / TensorRT / TGI etc. | inference engine R&D |
| Serving / serving systems | Model serving and production deployment | model serving / deployment |
| Low-level / kernel | Kernel-level development: CUDA / Triton / CUTLASS | kernel engineer |
| Quantization | Quantization: INT8 / FP8 / INT4 | quantization & compression |
| PagedAttention | vLLM's KV cache management mechanism | KV cache optimization |
| Speculative decoding | Speculative decoding: Medusa / EAGLE | speculative / assisted decoding |
| Distributed inference | Distributed inference: TP / PP / EP | large-scale inference |
| Multi-tenant serving | Multi-tenant inference services | inference platform / SaaS |
| Triton | OpenAI Triton Language (GPU DSL) | Triton kernel development |
| torch.compile / Inductor | PyTorch 2.x native AOT compilation | torch.compile optimization |
The "translation skill" for overseas job hunting
The same capability is written very differently in Chinese and English JDs. Where a Chinese JD asks for "familiarity with the vLLM source code," an English JD writes contributed to vLLM or equivalent inference engines. Where a Chinese JD asks for "operator optimization experience," an English JD writes experience with low-level GPU programming and kernel optimization. Before applying overseas, translate your project experience into the English JD's "verb + result" sentence pattern — this is the cross-language application of the core technique in Skills Benchmarking: What to Highlight on Your Resume.
4. Requirement Word-Frequency Table
The table below extracts high-frequency skill words from the Chinese and international JDs above, ordered by empirical frequency.
| Rank | Skill word | Frequency tier | Where it appears |
|---|---|---|---|
| 1 | Python | ★★★★★ near-universal | The language bar for all inference & deployment roles |
| 2 | C++ | ★★★★★ near-universal | Hard bar for engine / kernel / systems roles |
| 3 | PyTorch | ★★★★★ near-universal | The base framework for all roles |
| 4 | Linux | ★★★★☆ high | The hard foundation of deployment and ops |
| 5 | CUDA | ★★★★☆ high | Core for engine / optimization / kernel roles |
| 6 | vLLM / TensorRT-LLM | ★★★★☆ high | The de facto standard for LLM inference engines |
| 7 | Kubernetes | ★★★☆☆ medium | Essential for platform / systems engineers |
| 8 | Quantization (INT8 / FP8) | ★★★☆☆ medium | Core skill for optimization roles |
| 9 | KV cache / PagedAttention | ★★★☆☆ medium | The signature term of LLM inference roles |
| 10 | Triton (DSL) | ★★☆☆☆ rising | The new mainstream for kernel development |
| 11 | Speculative decoding | ★★☆☆☆ rising | A bonus item at startups |
| 12 | FlashAttention | ★★☆☆☆ rising | Standard equipment for LLM optimization roles |
Four readings worth noting:
- Python + C++ bilingual bar: the biggest difference between inference & deployment roles and algorithm roles — C++ appears in 80% of JDs. Reading and writing C++ is a hard filter condition; without it, your resume is screened out at the first pass.
- CUDA is the watershed: candidates who know CUDA earn 30–50% more than those who don't. CUDA doesn't demand mastery — but it does demand "can read it / can modify it / can write a simple reduction kernel."
- vLLM is the de facto standard after 2024: it replaced the 2021 position of ONNX Runtime / TensorRT in LLM scenarios. Knowing the vLLM source is the entry ticket to "engine engineer" roles.
- Triton (DSL) and speculative decoding are the "rising" group: still bonus items in 2024, they began entering hard requirements by 2026. Getting ahead of this curve is first-mover advantage.
5. Mapping Requirements to This Site's Pages
Map the high-frequency words above to this site's learning pages and you get a "find and fill the gaps" navigation:
| High-frequency skill | What the role actually demands | Corresponding page | How to use it |
|---|---|---|---|
| Python | Production-grade Python, not just scripts | Learning Paths: Three Routes and Glossary | Check concept wording; fill fundamentals along the path |
| C++ | Can read / modify C++ source, use modern C++ | vLLM and TensorRT and GPU Inference | Learn C++ through source reading |
| PyTorch | Can debug, can modify the model runner | vLLM | Understand the model runner and forward |
| Linux | Troubleshoot production issues, performance analysis | Inference Benchmarking in Practice | Build Linux hands-on skills with the toolchain |
| CUDA | Can write / modify simple kernels | GPU Architecture and Optimization | Start from the SIMT model |
| vLLM / TensorRT-LLM | Source-level understanding of engines | vLLM, TensorRT-LLM | Read source + hands-on |
| Kubernetes | Multi-replica deployment, canary rollout | Triton Inference Server | Learn K8s through deployment |
| Quantization | INT8 / FP8 / INT4 selection and accuracy evaluation | Model Quantization Fundamentals, Weight-Only Quantization and Mixed Precision | Run comparison experiments after learning the theory |
| KV cache / PagedAttention | Source-level understanding + hands-on optimization | Batching and Request Scheduling | Derive the KV cache memory formula |
| Triton (DSL) | Write custom attention / quantization kernels | Kernel Fusion and Custom Kernels | Learn Triton starting from matmul |
| Speculative decoding | Medusa / EAGLE production experience | Speculative Decoding and Medusa/EAGLE | Read papers + reproduce the demo |
| FlashAttention | Kernel-level understanding + performance tuning | Kernel Fusion and Custom Kernels, GPU Architecture and Optimization | Read papers + measure and compare |
What a single table can't cover — go to the knowledge map
This mapping only covers "word → page," not "you → word." Take the high-frequency words above into Deconstructing JD Knowledge Points and mark each as "can do / can't / half-can" to get a catch-up list that is truly yours. The JD tells you what the market wants; the knowledge map tells you what you're missing — you need both steps.
6. Trend Observations
1. Exploding demand for inference optimization roles
Use the rough metric "the share of algorithm-role JDs containing inference & deployment keywords (vLLM / TensorRT / CUDA / quantization / KV cache)":
2022 ~ 3% (pre-LLM; only CV/recommendation deployment roles)
2023 ~ 12% (vLLM goes viral; LLM inference roles become a category)
2024 ~ 30% (open-source LLMs mature; role count explodes)
2025 ~ 45% (domestic-chip inference roles rise)
2026 ~ 55%+ (inference & deployment becomes "standard background" for algorithm roles)
Basis: publicly visible trends in job descriptions on company career sites
and job platforms. Not precise statistics — magnitudes and direction only.Two details of this trend deserve attention:
- It's stacking, not substitution. "Knowing inference optimization" rarely forms a role by itself; most roles are "original-role skills + inference & deployment skills." Algorithm engineers are now expected to know quantization; platform engineers are expected to know vLLM — traditional skills didn't disappear; they just gained a layer.
- Domestic-chip inference roles are rising. Since 2025, inference openings at Huawei Ascend, Moore Threads, Biren and other vendors have grown 50%+ annually. CUDA engineers have first-mover advantage moving to domestic chips (the architecture concepts transfer; migration cost is 1–3 months) — a real alternative to the NVIDIA path.
2. Three Evolution Directions for Inference Roles
| Direction | Manifestation | What it means for job seekers |
|---|---|---|
| Extreme performance | LLM teams chase "thousand tokens/s" and long context | Kernel / scheduling / speculative-decoding experts are the scarcest and best paid |
| Platformization | Business-line algorithm roles shrink; capabilities consolidate into platform teams | Platform / systems engineer roles grow steadily |
| Edge inference | Phone-side LLMs (Phi-3, Qwen2.5-0.5B) mature | Edge inference roles emerge from zero in 2025–2026 |
A structural judgment worth remembering
In inference & deployment, the model itself is becoming infrastructure (pay-per-call), while performance optimization, scheduling, and quantization become the differentiators. This explains two things: why platform / MLOps roles are multiplying, and why "can modify vLLM / write CUDA kernels" keeps gaining weight in LLM JDs. Those two areas are exactly the core of this site's Batching and Request Scheduling and Model Quantization Fundamentals pages.
3. Three Pragmatic Suggestions for Job Seekers
- Protect the "Python + C++ + Linux" trinity. It is the common denominator of all inference & deployment roles and your ballast against role volatility.
- Make vLLM source your "second language." You don't need to write an inference engine from scratch, but being able to read vLLM's three core modules — scheduler, block manager, model runner — is the standard question in 2026 inference-role interviews; two weeks of close reading is worth it.
- Decide with data, not emotion. Refresh your reading of this page every quarter: go to a job board, read 20 recent JDs for your target role, and re-dissect them with the "hard requirement / bonus item" framework in Section 1. JDs are signals the market sends you; you only need to learn to decode them.
7. Further Reading
- Module Guide and Job Landscape — the panorama of seven inference & deployment roles; position yourself before filling gaps
- Deconstructing JD Knowledge Points — turn JD skill words into your personal L1→L3 catch-up list
- Skills Benchmarking: What to Highlight on Your Resume — rewrite project experience against JDs and tell the story interviewers want to hear
- Interview Question Bank — high-frequency questions and answer frameworks; final self-test
- Learning Paths: Three Routes — the complete main line of back-solving learning from JDs and interview questions
- What Is Inference Acceleration? — the conceptual boundary of inference & deployment roles
- vLLM, TensorRT-LLM — source-level deep reading of engines
- Model Quantization Fundamentals, Batching and Request Scheduling — core concepts of inference optimization
- Glossary — unified terminology, so "half-can" misunderstandings don't happen
References
Reliable channels for obtaining real JDs. All generalizations on this page are based on public information from these channels; defer to official career pages for real-time JDs — third-party platforms may lag:
- Official career sites of Chinese companies: ByteDance (jobs.bytedance.com), Alibaba (talent.alibaba.com), Tencent (careers.tencent.com), Huawei (career.huawei.com), Baidu (talent.baidu.com), Moonshot (moonshot.cn/careers), DeepSeek (deepseek.com), Zhipu AI (zhipuai.cn), MiniMax (minimaxi.com), StepFun (stepfun.com), and other official job boards
- Official career sites of international companies: OpenAI (openai.com/careers), Anthropic (anthropic.com/careers), Meta Careers (careers.meta.com), NVIDIA Careers (nvidia.com/en-us/about-nvidia/careers), Groq (groq.com/careers), Together AI (together.ai/careers), Anyscale (anyscale.com/careers), and more
- Chinese job platforms: Maimai (maimai.cn), BOSS Zhipin (zhipin.com), Liepin (liepin.com) — for observing role distribution and JD wording frequency
- Overseas salary and job statistics: levels.fyi — for role types and compensation magnitudes; not official data
- Open-source repos as a "resume channel": vLLM (github.com/vllm-project/vllm), SGLang (github.com/sgl-project/sglang), LMDeploy (github.com/InternLM/lmdeploy), TensorRT-LLM (github.com/NVIDIA/TensorRT-LLM) — merged PRs are the hardest "bonus-item evidence"
One bottom line
This page does not provide or quote the text of any "company X's 2026 opening for role Y," because such content is both unverifiable and guaranteed to go stale. Defer to JDs published in real time on official channels — treat every JD you read as training data; this page only teaches you how to read them.