Skip to content

Job Description Checklist

At a glance What do real model deployment job descriptions look like? This page compiles inference / MLOps / ML platform roles currently open at 15+ companies in China and abroad, breaks down the core requirements line by line, and marks estimated salary ranges (for reference only, not official).

This page contains time-sensitive content. Data is current as of 2025-12; job listings, rankings, product features, and similar details may have changed since then — verify against the original source before citing.

Job Description Checklist ​

Judging roles only by "requirements" and "salary" without ever reading a real JD is the most common blind spot among job seekers. This page compiles the inference / MLOps / ML platform roles open at 16 companies in China and abroad (as of the data date), breaking down the core requirements (tech stack) and nice-to-haves for each, along with estimated salary ranges.

Disclaimer and freshness

All figures on this page are estimates. Roles, cities, and salaries change constantly — always verify against the original source before citing (official career pages are linked below). Salary notation like ¥30k-60K × 15 means a monthly range in RMB times 15 months (estimated); overseas figures are annual ranges. dataAsOf: 2025-12 is the data collection date.

1. Companies in China ​

1. ByteDance (Doubao / Volcano Engine) ​

  • Role directions: AI Infra engineer, inference optimization engineer (LLM), GPU cluster scheduling engineer
  • Cities: Beijing, Shanghai, Hangzhou, Shenzhen
  • Core requirements: K8s / Docker; GPU scheduling (Kubernetes device plugin / Volcano); inference engines (at least one of vLLM, TensorRT-LLM, Triton); performance profiling (nsys, perf); familiarity with LLM inference
  • Nice-to-haves: CUDA kernel development, RDMA networking, production quantization experience (FP8/INT8)
  • Salary range (estimated): ¥30k-60K × 15 (L4/L5 band; senior levels negotiated separately)
  • Official portal: https://jobs.bytedance.com/

2. Alibaba (Alibaba Cloud PAI / Tongyi) ​

  • Role directions: PAI platform R&D engineer, inference engine engineer (LLMs), MLOps platform engineer
  • Cities: Hangzhou, Beijing, Shanghai
  • Core requirements: Java/Go backend + K8s ecosystem; inference frameworks (Triton, vLLM, MNN, PyTorch Inference); model serving orchestration (KServe / in-house); engineering experience with deployment patterns
  • Nice-to-haves: Kubeflow/Argo pipelines, mixed-workload scheduling, distributed LLM inference
  • Salary range (estimated): ¥30k-55K × 16 (P6/P7 band)
  • Official portal: https://talent.alibaba.com/

3. Tencent (Machine Learning Platform / Hunyuan) ​

  • Role directions: ML platform development, LLM inference optimization, MLOps engineer
  • Cities: Shenzhen, Beijing, Shanghai
  • Core requirements: C++/Go + high-performance services; TensorRT/ONNX Runtime; GPU cluster management; load testing and capacity planning (see Load Testing)
  • Nice-to-haves: Model quantization and pruning, elastic inference, logging/monitoring systems
  • Salary range (estimated): ¥25k-50K × 16 (T9/T10 band)
  • Official portal: https://careers.tencent.com/

4. Baidu (PaddlePaddle / Kunlunxin) ​

  • Role directions: PaddlePaddle inference deployment engineer, Kunlunxin software stack engineer (CANN-like), MLOps platform
  • Cities: Beijing, Shenzhen
  • Core requirements: PaddlePaddle/ONNX ecosystem; operator optimization (C++/CUDA/OpenCL); chip toolchains (npu-smi-style tools); quantization and low-bit deployment
  • Nice-to-haves: In-house chip (Kunlunxin) adaptation experience, edge deployment (Paddle Lite)
  • Salary range (estimated): ¥25k-50K × 15
  • Official portal: https://talent.baidu.com/

5. Meituan (Search & Recommendation Platform) ​

  • Role directions: Search/recommendation model deployment engineer, inference service lead, GPU platform engineer
  • Cities: Beijing, Shanghai
  • Core requirements: High-concurrency services (C++/Go/Java); TensorRT/Triton; serving online recommender models (DeepFM-style); latency optimization and resource governance
  • Nice-to-haves: Dynamic batching, hot model updates, multi-level caching
  • Salary range (estimated): ¥28k-55K × 15.5
  • Official portal: https://zhaopin.meituan.com/

6. Ant Group (AI Platform / Bailing LLM) ​

  • Role directions: AI platform R&D, inference optimization engineer, MLOps engineer
  • Cities: Hangzhou, Shanghai, Beijing
  • Core requirements: K8s and cloud-native; Triton/vLLM; financial-grade reliability (disaster recovery, canary releases, auditing); monitoring and alerting systems
  • Nice-to-haves: Secure inference (federated / confidential computing, see Security), multi-tenant isolation
  • Salary range (estimated): ¥30k-60K × 16
  • Official portal: https://talent.antgroup.com/

7. Huawei (Ascend / Huawei Cloud ModelArts) ​

  • Role directions: Ascend inference engine engineer (MindIE), ModelArts platform engineer, operator library development
  • Cities: Shenzhen, Beijing, Xi'an, Shanghai
  • Core requirements: C++/Python; CANN/MindSpore ecosystem; operator development and tuning (Ascend NPU); LLM inference acceleration (PagedAttention-style)
  • Nice-to-haves: TensorRT experience (transfers well), NPU affinity optimization
  • Salary range (estimated): ¥25k-55K × 14-16 (includes performance pay)
  • Official portal: Ascend community at https://www.hiascend.com/ (hiring goes through https://career.huawei.com/)

8. JD.com / Xiaohongshu (RedNote) / Kuaishou (additional observations) ​

CompanyRole DirectionsCore Requirements at a GlanceCitiesSalary Range (est.)
JD.comLLM inference, smart supply chain MLOpsvLLM, K8s, load testingBeijing¥25k-50K × 15
XiaohongshuRecommendation inference, AI InfraTriton, C++, GPU clustersShanghai, Beijing¥30k-55K × 16
KuaishouInference engines, LLM servingTensorRT, CUDA, quantizationBeijing¥28k-55K × 16

2. International Companies ​

9. NVIDIA (TensorRT / Triton) ​

  • Role directions: Deep Learning Performance Engineer, Triton Inference Server Engineer, TensorRT Software Engineer
  • Cities: Santa Clara, Shanghai, Shenzhen (NVIDIA China R&D)
  • Core requirements: C++/CUDA; deep learning framework internals; performance profiling (Nsight); distributed inference
  • Nice-to-haves: Open-source project experience (commit history in TensorRT-LLM, Triton, or vLLM)
  • Salary range (estimated): $150k-250k (US); China benchmark ¥30k-60K × 15+ (estimated)
  • Official portal: https://www.nvidia.com/en-us/about-nvidia/careers/

10. Microsoft (Azure ML / AI Infra) ​

  • Role directions: Azure ML Engineer, AI Platform Infra, ONNX Runtime Developer
  • Cities: Redmond, Beijing (Microsoft Asia-Pacific R&D), Suzhou
  • Core requirements: Cloud-native (K8s); ONNX Runtime/MLflow; distributed training and serving; productionizing MLOps pipelines
  • Nice-to-haves: Azure/Kubernetes certifications, open-source contributions
  • Salary range (estimated): $140k-230k (US); China ¥30k-55K × 15+ (estimated)
  • Official portal: https://careers.microsoft.com/

11. AWS (SageMaker / Inferentia) ​

  • Role directions: Machine Learning Engineer, SageMaker Platform, Neuron SDK engineer
  • Cities: Seattle, Beijing (AWS China), Shenzhen
  • Core requirements: Python/Java + cloud-native; SageMaker/KServe; model serving and monitoring; observability in practice
  • Nice-to-haves: AWS certifications, in-house chip adaptation (Trainium/Inferentia)
  • Salary range (estimated): $140k-230k (US)
  • Official portal: https://www.amazon.jobs/

12. Google (Vertex AI / JAX serving) ​

  • Role directions: ML Infrastructure Engineer, TPU Serving, Vertex AI Platform
  • Cities: Mountain View, Zurich, London
  • Core requirements: Distributed systems; JAX/TensorFlow serving; large-scale capacity planning (see capacity planning interview questions)
  • Nice-to-haves: TPU experience, site reliability engineering (SRE) background
  • Salary range (estimated): $150k-260k (US)
  • Official portal: https://careers.google.com/

13. Meta (TorchServe / ML Infra) ​

  • Role directions: ML Infra Engineer, TorchServe / PyTorch Serving, large-scale recommendation serving
  • Cities: Menlo Park, London, New York
  • Core requirements: C++/Python; PyTorch ecosystem; high-concurrency low-latency systems; hands-on performance optimization
  • Nice-to-haves: Open-source contributions to PyTorch/TorchServe, FPGA/ASIC co-design experience
  • Salary range (estimated): $150k-270k (US)
  • Official portal: https://www.metacareers.com/

14. Hugging Face (Inference) ​

  • Role directions: Inference Engineer, Inference Endpoints, Optimum (ONNX/TensorRT) engineer
  • Cities: Paris, New York, remote
  • Core requirements: Python + deep learning frameworks; ONNX Runtime/TensorRT/OpenVINO; the Transformers ecosystem; containerization
  • Nice-to-haves: Active in the open-source community; has built deployment tools in wide use
  • Salary range (estimated): €60k-100k (Europe) or $120k-200k (US, estimated)
  • Official portal: https://huggingface.co/careers

15. Anthropic / OpenAI (inference infra, remote-friendly) ​

CompanyRole DirectionsCore Requirements at a GlanceSalary Range (est.)
AnthropicInfra/Performance Engineer, ML Engineer (inference)Large-scale LLM inference, K8s, capacity planning$180k-320k
OpenAIInference Infra, Platform EngineerLLM serving, GPU clusters, performance profiling$180k-350k

16. Databricks (MLOps) ​

  • Role directions: MLOps Engineer, ML Platform Engineer, MosaicML inference engineer
  • Cities: San Francisco, remote
  • Core requirements: MLflow (Databricks' own ecosystem); K8s and GPU clusters; LLM inference serving; model registry and pipelines
  • Nice-to-haves: MLflow contributor, end-to-end experience with LLM fine-tuning and deployment
  • Salary range (estimated): $150k-260k (US)
  • Official portal: https://www.databricks.com/company/careers

3. Common Threads: Skills That Keep Reappearing ​

After deduplicating the JDs from these 16 companies, the most frequently appearing skills group into five layers:

LayerHigh-Frequency KeywordsCompanies Mentioning (est.)Pages on This Site
Containers & schedulingDocker, Kubernetes, GPU scheduling15+deployment patterns, anatomy of an inference system
Inference enginesvLLM, Triton, TensorRT, ONNX Runtime13+framework comparison, Triton case study, vLLM case study
Performance & costLoad testing, latency optimization, capacity planning, quantization12+performance optimization, quantization, load testing
Platform & processCI/CD, model registry, canary releases, monitoring & alerting11+MLOps pipelines, monitoring, rollout process
LLM-specificKV Cache, continuous batching, tensor parallelism, speculative decoding9+LLM inference

Three direct takeaways for job seekers:

  1. K8s + one inference engine + one LLM concept is the "minimum common subset" for nearly every target role — close that gap first.
  2. Quantization / performance optimization is the highest-differentiation bonus: more than half the JDs list it as a core requirement or a plus, yet only a minority of candidates can withstand deep questioning on it in an interview.
  3. Job titles are noise; skill layers are the signal: sort any target role's skills into the five layers above and you can instantly see which layer you're missing. For the layer-by-layer method, see Breaking Down JD Requirements.

How to use this list

Don't apply to everything on the list. Instead: circle 3 "reachable" targets + 2 "stretch" targets, write a tailored resume for each target role (method in Resume Benchmarking), then cross-check the role classification against the job landscape.

Further Reading ​

References ​

  • Official career portals for each company are linked at the end of each entry above (all are official company domains and should remain accessible long-term; the live listing on the official site prevails on any given day)
  • Salary references: levels.fyi (overseas, https://www.levels.fyi/) and aggregated public interview-experience reports from China; all figures are estimates