Appearance
Job Description Checklist
Judging roles only by "requirements" and "salary" without ever reading a real JD is the most common blind spot among job seekers. This page compiles the inference / MLOps / ML platform roles open at 16 companies in China and abroad (as of the data date), breaking down the core requirements (tech stack) and nice-to-haves for each, along with estimated salary ranges.
Disclaimer and freshness
All figures on this page are estimates. Roles, cities, and salaries change constantly — always verify against the original source before citing (official career pages are linked below). Salary notation like ¥30k-60K × 15 means a monthly range in RMB times 15 months (estimated); overseas figures are annual ranges. dataAsOf: 2025-12 is the data collection date.
1. Companies in China
1. ByteDance (Doubao / Volcano Engine)
- Role directions: AI Infra engineer, inference optimization engineer (LLM), GPU cluster scheduling engineer
- Cities: Beijing, Shanghai, Hangzhou, Shenzhen
- Core requirements: K8s / Docker; GPU scheduling (Kubernetes device plugin / Volcano); inference engines (at least one of vLLM, TensorRT-LLM, Triton); performance profiling (nsys, perf); familiarity with LLM inference
- Nice-to-haves: CUDA kernel development, RDMA networking, production quantization experience (FP8/INT8)
- Salary range (estimated): ¥30k-60K × 15 (L4/L5 band; senior levels negotiated separately)
- Official portal: https://jobs.bytedance.com/
2. Alibaba (Alibaba Cloud PAI / Tongyi)
- Role directions: PAI platform R&D engineer, inference engine engineer (LLMs), MLOps platform engineer
- Cities: Hangzhou, Beijing, Shanghai
- Core requirements: Java/Go backend + K8s ecosystem; inference frameworks (Triton, vLLM, MNN, PyTorch Inference); model serving orchestration (KServe / in-house); engineering experience with deployment patterns
- Nice-to-haves: Kubeflow/Argo pipelines, mixed-workload scheduling, distributed LLM inference
- Salary range (estimated): ¥30k-55K × 16 (P6/P7 band)
- Official portal: https://talent.alibaba.com/
3. Tencent (Machine Learning Platform / Hunyuan)
- Role directions: ML platform development, LLM inference optimization, MLOps engineer
- Cities: Shenzhen, Beijing, Shanghai
- Core requirements: C++/Go + high-performance services; TensorRT/ONNX Runtime; GPU cluster management; load testing and capacity planning (see Load Testing)
- Nice-to-haves: Model quantization and pruning, elastic inference, logging/monitoring systems
- Salary range (estimated): ¥25k-50K × 16 (T9/T10 band)
- Official portal: https://careers.tencent.com/
4. Baidu (PaddlePaddle / Kunlunxin)
- Role directions: PaddlePaddle inference deployment engineer, Kunlunxin software stack engineer (CANN-like), MLOps platform
- Cities: Beijing, Shenzhen
- Core requirements: PaddlePaddle/ONNX ecosystem; operator optimization (C++/CUDA/OpenCL); chip toolchains (npu-smi-style tools); quantization and low-bit deployment
- Nice-to-haves: In-house chip (Kunlunxin) adaptation experience, edge deployment (Paddle Lite)
- Salary range (estimated): ¥25k-50K × 15
- Official portal: https://talent.baidu.com/
5. Meituan (Search & Recommendation Platform)
- Role directions: Search/recommendation model deployment engineer, inference service lead, GPU platform engineer
- Cities: Beijing, Shanghai
- Core requirements: High-concurrency services (C++/Go/Java); TensorRT/Triton; serving online recommender models (DeepFM-style); latency optimization and resource governance
- Nice-to-haves: Dynamic batching, hot model updates, multi-level caching
- Salary range (estimated): ¥28k-55K × 15.5
- Official portal: https://zhaopin.meituan.com/
6. Ant Group (AI Platform / Bailing LLM)
- Role directions: AI platform R&D, inference optimization engineer, MLOps engineer
- Cities: Hangzhou, Shanghai, Beijing
- Core requirements: K8s and cloud-native; Triton/vLLM; financial-grade reliability (disaster recovery, canary releases, auditing); monitoring and alerting systems
- Nice-to-haves: Secure inference (federated / confidential computing, see Security), multi-tenant isolation
- Salary range (estimated): ¥30k-60K × 16
- Official portal: https://talent.antgroup.com/
7. Huawei (Ascend / Huawei Cloud ModelArts)
- Role directions: Ascend inference engine engineer (MindIE), ModelArts platform engineer, operator library development
- Cities: Shenzhen, Beijing, Xi'an, Shanghai
- Core requirements: C++/Python; CANN/MindSpore ecosystem; operator development and tuning (Ascend NPU); LLM inference acceleration (PagedAttention-style)
- Nice-to-haves: TensorRT experience (transfers well), NPU affinity optimization
- Salary range (estimated): ¥25k-55K × 14-16 (includes performance pay)
- Official portal: Ascend community at https://www.hiascend.com/ (hiring goes through https://career.huawei.com/)
8. JD.com / Xiaohongshu (RedNote) / Kuaishou (additional observations)
| Company | Role Directions | Core Requirements at a Glance | Cities | Salary Range (est.) |
|---|---|---|---|---|
| JD.com | LLM inference, smart supply chain MLOps | vLLM, K8s, load testing | Beijing | ¥25k-50K × 15 |
| Xiaohongshu | Recommendation inference, AI Infra | Triton, C++, GPU clusters | Shanghai, Beijing | ¥30k-55K × 16 |
| Kuaishou | Inference engines, LLM serving | TensorRT, CUDA, quantization | Beijing | ¥28k-55K × 16 |
- Official portals: JD.com https://zhaopin.jd.com/; Xiaohongshu https://job.xiaohongshu.com/; Kuaishou https://zhaopin.kuaishou.com/
2. International Companies
9. NVIDIA (TensorRT / Triton)
- Role directions: Deep Learning Performance Engineer, Triton Inference Server Engineer, TensorRT Software Engineer
- Cities: Santa Clara, Shanghai, Shenzhen (NVIDIA China R&D)
- Core requirements: C++/CUDA; deep learning framework internals; performance profiling (Nsight); distributed inference
- Nice-to-haves: Open-source project experience (commit history in TensorRT-LLM, Triton, or vLLM)
- Salary range (estimated): $150k-250k (US); China benchmark ¥30k-60K × 15+ (estimated)
- Official portal: https://www.nvidia.com/en-us/about-nvidia/careers/
10. Microsoft (Azure ML / AI Infra)
- Role directions: Azure ML Engineer, AI Platform Infra, ONNX Runtime Developer
- Cities: Redmond, Beijing (Microsoft Asia-Pacific R&D), Suzhou
- Core requirements: Cloud-native (K8s); ONNX Runtime/MLflow; distributed training and serving; productionizing MLOps pipelines
- Nice-to-haves: Azure/Kubernetes certifications, open-source contributions
- Salary range (estimated): $140k-230k (US); China ¥30k-55K × 15+ (estimated)
- Official portal: https://careers.microsoft.com/
11. AWS (SageMaker / Inferentia)
- Role directions: Machine Learning Engineer, SageMaker Platform, Neuron SDK engineer
- Cities: Seattle, Beijing (AWS China), Shenzhen
- Core requirements: Python/Java + cloud-native; SageMaker/KServe; model serving and monitoring; observability in practice
- Nice-to-haves: AWS certifications, in-house chip adaptation (Trainium/Inferentia)
- Salary range (estimated): $140k-230k (US)
- Official portal: https://www.amazon.jobs/
12. Google (Vertex AI / JAX serving)
- Role directions: ML Infrastructure Engineer, TPU Serving, Vertex AI Platform
- Cities: Mountain View, Zurich, London
- Core requirements: Distributed systems; JAX/TensorFlow serving; large-scale capacity planning (see capacity planning interview questions)
- Nice-to-haves: TPU experience, site reliability engineering (SRE) background
- Salary range (estimated): $150k-260k (US)
- Official portal: https://careers.google.com/
13. Meta (TorchServe / ML Infra)
- Role directions: ML Infra Engineer, TorchServe / PyTorch Serving, large-scale recommendation serving
- Cities: Menlo Park, London, New York
- Core requirements: C++/Python; PyTorch ecosystem; high-concurrency low-latency systems; hands-on performance optimization
- Nice-to-haves: Open-source contributions to PyTorch/TorchServe, FPGA/ASIC co-design experience
- Salary range (estimated): $150k-270k (US)
- Official portal: https://www.metacareers.com/
14. Hugging Face (Inference)
- Role directions: Inference Engineer, Inference Endpoints, Optimum (ONNX/TensorRT) engineer
- Cities: Paris, New York, remote
- Core requirements: Python + deep learning frameworks; ONNX Runtime/TensorRT/OpenVINO; the Transformers ecosystem; containerization
- Nice-to-haves: Active in the open-source community; has built deployment tools in wide use
- Salary range (estimated): €60k-100k (Europe) or $120k-200k (US, estimated)
- Official portal: https://huggingface.co/careers
15. Anthropic / OpenAI (inference infra, remote-friendly)
| Company | Role Directions | Core Requirements at a Glance | Salary Range (est.) |
|---|---|---|---|
| Anthropic | Infra/Performance Engineer, ML Engineer (inference) | Large-scale LLM inference, K8s, capacity planning | $180k-320k |
| OpenAI | Inference Infra, Platform Engineer | LLM serving, GPU clusters, performance profiling | $180k-350k |
- Official portals: Anthropic https://www.anthropic.com/careers; OpenAI https://openai.com/careers/
16. Databricks (MLOps)
- Role directions: MLOps Engineer, ML Platform Engineer, MosaicML inference engineer
- Cities: San Francisco, remote
- Core requirements: MLflow (Databricks' own ecosystem); K8s and GPU clusters; LLM inference serving; model registry and pipelines
- Nice-to-haves: MLflow contributor, end-to-end experience with LLM fine-tuning and deployment
- Salary range (estimated): $150k-260k (US)
- Official portal: https://www.databricks.com/company/careers
3. Common Threads: Skills That Keep Reappearing
After deduplicating the JDs from these 16 companies, the most frequently appearing skills group into five layers:
| Layer | High-Frequency Keywords | Companies Mentioning (est.) | Pages on This Site |
|---|---|---|---|
| Containers & scheduling | Docker, Kubernetes, GPU scheduling | 15+ | deployment patterns, anatomy of an inference system |
| Inference engines | vLLM, Triton, TensorRT, ONNX Runtime | 13+ | framework comparison, Triton case study, vLLM case study |
| Performance & cost | Load testing, latency optimization, capacity planning, quantization | 12+ | performance optimization, quantization, load testing |
| Platform & process | CI/CD, model registry, canary releases, monitoring & alerting | 11+ | MLOps pipelines, monitoring, rollout process |
| LLM-specific | KV Cache, continuous batching, tensor parallelism, speculative decoding | 9+ | LLM inference |
Three direct takeaways for job seekers:
- K8s + one inference engine + one LLM concept is the "minimum common subset" for nearly every target role — close that gap first.
- Quantization / performance optimization is the highest-differentiation bonus: more than half the JDs list it as a core requirement or a plus, yet only a minority of candidates can withstand deep questioning on it in an interview.
- Job titles are noise; skill layers are the signal: sort any target role's skills into the five layers above and you can instantly see which layer you're missing. For the layer-by-layer method, see Breaking Down JD Requirements.
How to use this list
Don't apply to everything on the list. Instead: circle 3 "reachable" targets + 2 "stretch" targets, write a tailored resume for each target role (method in Resume Benchmarking), then cross-check the role classification against the job landscape.
Further Reading
- Career: Overview and Job Landscape — responsibilities and trends for the five role types
- Breaking Down JD Requirements — translate keywords into a study list
- Resume Benchmarking: What to Highlight — make your resume hit the JD keywords
- Model Deployment Interview Questions — verify your skills measure up
- Learning Paths: Three Routes — the 2-week job-search sprint roadmap
- Framework Comparison — answer "why choose this engine"
References
- Official career portals for each company are linked at the end of each entry above (all are official company domains and should remain accessible long-term; the live listing on the official site prevails on any given day)
- Salary references: levels.fyi (overseas, https://www.levels.fyi/) and aggregated public interview-experience reports from China; all figures are estimates