Appearance
Module Overview and Job Market Landscape
One-line summary: this page answers "what does the RL job market actually look like?" — which roles are hiring, what each one does, what it demands, and what it pays, and, given that landscape, where you should start in the career module (the four-page usage map is in Section 5).
Look back from 2026 and reinforcement learning has gone from "a direction academia cheers on by itself" to a hard skill running through multiple industry pipelines: LLM alignment, embodied AI, autonomous driving, recommendation and advertising, operations research and scheduling, quantitative trading — every one of these verticals is hiring RL people, and every one wants a different kind of RL. Plenty of applications die not because the candidate can't do RL, but because they never figured out which RL the market is actually buying — that is the first reason this career module exists.
1. What This Module Is For: Read the Market Before You Prepare Yourself
The career module runs on a single principle: read the market before you prepare yourself. Most candidates make the mistake in reverse — they bury themselves in a year of study, then finally open the job boards and discover that nothing they learned matches a single JD. Our answer is an assembly line with four stages:
text
What the market is hiring What the JDs demand How to prepare How you're verified
┌───────────────┐ ┌──────────────────┐ ┌────────────────┐ ┌───────────────┐
│ Job landscape │ │ Skill breakdown │ │ Resume/portfolio│ │ Interview │
│ (this page) │────▶│ knowledge-map │────▶│ resume-analysis │────▶│ question bank │
│ Market map │ │ skill→page map │ │ 4 elements+STAR │ │ top Qs+digs │
└───────────────┘ └──────────────────┘ └────────────────┘ └───────────────┘
▲ │ │ │
│ ▼ ▼ ▼
[jd-list] breaks down [concepts / [practice] [concepts]
the actual JD text of practice pages] portfolio & every knowledge
live openings on this site evaluation point on this siteHow this module relates to the Guide section
Learning Paths answers "in what order should I learn things"; the career module answers "what does the market want, and how do I line myself up with it." The job-search sprint track is a roughly two-week interview-prep route built on the four career pages plus the high-frequency concept pages. The two tracks are two sides of the same coin: career tells you what to learn; the guide tells you the pace to learn it at.
Each of the four pages answers exactly one question:
| Page | Question it answers | Best moment to use it |
|---|---|---|
| This page (index) | What RL roles exist in the market, and which one fits me | At the start — 30 minutes, end to end |
| JD List | What real, live JDs actually demand, line by line | Once you have a concrete target role, compare word by word |
| Skill Breakdown | Which page on this site covers each skill term in the JD — and whether I truly know it | During self-assessment, to produce your personal catch-up list |
| Resume Analysis | How to make your RL ability stand out in a resume and portfolio | 1–2 weeks before you start applying |
| Interview Question Bank | What interviews ask, and how to answer it | The final sprint before interviews |
One important caveat
All role descriptions and salary ranges below swing violently with the market, and they are composite summaries drawn from multiple sources — not the official pay band of any specific opening. For any concrete decision — which company to apply to, what number to ask for — defer to the live listings in the JD List. This site hands you a framework for reading the market, not the market's final answer.
2. The Job Landscape: A Panorama of Ten RL Roles
2.1 The Master Table of Ten Roles
We group the roles that are genuinely hiring RL talent into ten categories. The grouping deliberately follows the role RL plays in the business, not the company — because a single company may have four or five of these ten open at once.
| # | Role | Typical employers | Core responsibilities | Key skill mix | Reference pay (China / year) | Reference pay (overseas / year, TC) |
|---|---|---|---|---|---|---|
| 1 | RL Research Scientist | DeepMind, OpenAI, Meta FAIR, big-tech research labs in China, AI startups | Propose new algorithms and paradigms, run the decisive experiments, publish papers or file patents | Mathematical depth, paper reproduction, large-scale experimentation | ¥600k–1.5M | $250k–600k+ |
| 2 | RL Algorithm Engineer (Applied Scientist / RL Engineer) | Big tech in China, autonomous driving, gaming, quant | Land RL algorithms in a concrete business; tune, deploy, monitor | Algorithms + engineering + business modeling | ¥400k–900k | $220k–450k |
| 3 | LLM Alignment Engineer (RLHF) | LLM companies, cloud providers, AI infra | Full pipeline of SFT / reward model / PPO / DPO; alignment evaluation | RLHF, distributed training, evaluation | ¥500k–1.2M | $250k–500k+ |
| 4 | Robotics Learning Engineer (Embodied AI) | Embodied-AI startups, Zhiyuan Robotics (AgiBot) / Unitree / DJI, humanoid-robot companies | Sim training, Sim2Real, VLA policies, data collection | Continuous control, simulators, RL + imitation | ¥500k–1.2M | $220k–450k |
| 5 | Simulation & Training Platform Engineer (RL Infra) | Big-tech RL teams, embodied AI, autonomous driving | Environment development, parallel training frameworks, data pipelines | HPC, simulators, distributed systems | ¥450k–900k | $200k–400k |
| 6 | Recommendation / Ads / Search Strategy Algorithm (Decision Intelligence) | ByteDance, Alibaba, Tencent, Meituan | Ranking, bidding, exploration strategies, offline evaluation | Contextual bandits, offline RL, causality | ¥400k–900k | $200k–380k |
| 7 | Decision Optimization / OR & Scheduling Engineer | Logistics, manufacturing, cloud resource scheduling, supply chain | Sequential decisions for inventory, scheduling, bin packing, routing | MDP modeling, OR algorithms, RL hybrids | ¥350k–800k | $180k–350k |
| 8 | Autonomous Driving Decision & Planning | Huawei, Baidu, Pony.ai, Haomo.AI, Waymo-type companies | Hybrid imitation + RL for behavioral decision-making and motion planning | Planning, simulation testing, safety constraints | ¥400k–1M | $220k–420k |
| 9 | Quantitative Trading / Execution Optimization Researcher | Quant private funds, brokerage prop-trading desks | RL for order execution, market making, and portfolio construction | Formulating finance as MDPs, backtesting, risk control | ¥600k–1.5M+ | $250k–600k+ |
| 10 | Game AI Engineer | Major game studios, game-AI startups | NPC decision-making, competitive AI, RL for content generation | Game environments, multi-agent, adversarial play | ¥300k–700k | $180k–320k |
How to read this table: pay figures are composite market estimates for 2026; China numbers include cash plus equity, while overseas numbers are total compensation (TC: base + bonus + equity). For the same role, the gap between a top-tier company and the rest can exceed 2x. For concrete numbers, defer to the live listings in the JD List.
2.2 Three "Archetype Roles" — Most Jobs Are Their Variants
Ten roles look like a lot, but they all descend from three archetypes; understand the archetypes and the variants take care of themselves:
- Research (archetype one): the goal is "build something nobody has built before." Maps to #1, the research-flavored variant of #3, and the "policy innovation" variant of #4. What gets tested: math and papers.
- Applied (archetype two): the goal is "make RL produce measurable gains in a concrete business." Maps to #2, #6, #7, #8, #9, and #10. What gets tested: modeling ability + engineering delivery.
- Platform (archetype three): the goal is "make RL train at scale in the first place." Maps to #5, plus the infra roles inside #2/#3. What gets tested: distributed systems and engineering.
Why we cut it this way
The same title "RL engineer" means "get SAC from the paper running on a robot" at company A, "write the distributed sampling framework" at company B, and "align bidding targets with the ops team" at company C. Before applying, judge which archetype a JD belongs to, then organize your resume accordingly — that is the single most repeated message of this career module.
2.3 Three Directional Trends (the 2026 View)
- RLHF roles are going from rare to mainstream: alignment is now a standard stage of LLM training; nearly every LLM company has an alignment team, and demand is spilling over to cloud providers and vertical industries (domain alignment for healthcare and financial customer service). The RL bar for these roles is actually not high — PPO alone covers most of it — but distributed training and evaluation skills are critical.
- Embodied AI has pushed robotics learning into the spotlight: humanoid robots, dexterous hands, and VLA (vision-language-action) models are seeing peak demand for both funding and talent. What these roles lack most are people who can write RL algorithms and wrangle simulators at the same time.
- Traditional RL roles have changed their name to "decision intelligence": the frequency of bandit, offline RL, and causal-inference keywords in JDs for recommendation, advertising, risk control, and scheduling keeps climbing. Many of these jobs are RL in everything but name — when applying, don't search for just the words "reinforcement learning."
3. Skill Radar: The Five Capability Axes
Whatever the role, RL talent can be gauged on five capability axes. These five are the master coordinates we reach for over and over when dissecting JDs, writing resumes, and preparing for interviews:
| Axis | Representative skills | How interviews test it | On this site |
|---|---|---|---|
| Theory | MDPs, Bellman equations, convergence intuition, the algorithm family tree | Deriving formulas by hand, comparison questions, "why" follow-ups | concepts module in full |
| Code | PyTorch, algorithm implementation, debugging, experiment scripts | Writing pseudocode live, code review, fixing a bug on the spot | practice module and framework selection |
| Math | Probability, expectation, KL, convex optimization, statistics | Deriving the advantage, explaining GAE / entropy | math-primer |
| Engineering | Simulators, distributed training, data pipelines, evaluation systems | How you'd build the environment, run experiments, ensure reproducibility | evaluation-in-practice and build-your-own |
| Business | Modeling a business problem as an MDP, designing rewards, measuring business impact | Scenario design questions, case questions | reward-engineering and the case-study pages |
Score yourself 1–5 on each axis and you instantly get a mismatch map against every role:
text
Skill self-assessment example (a candidate with 2 years of engineering experience, switching into RL):
Theory Code Math Eng. Business
Research Scientist (research) 5 4 5 3 1 ← math/theory gap
RL Engineer (applied) 3 5 3 4 3 ← theory/math are the weak spots
Alignment Engineer (applied) 3 4 3 5 3 ← RL theory could use a top-up
Robotics Learning (applied) 4 4 3 4 2 ← math/simulation could use a top-up
Simulation Platform (platform) 2 5 2 5 2 ← low theory bar, high engineering barHow to read this radar
A radar chart is not "higher average is better" — it's "the shape must match the role." Research roles want a theory/math spike, platform roles a code/engineering spike, applied roles an engineering/business spike. The three most common reasons for failing an interview map to three axis dips: research roles fail you on math, applied roles on engineering delivery, and every role on business modeling.
4. A Three-Dimension Decision Checklist: Which Direction to Pick
Don't pick from ten roles on gut feel. Score one round on each of the three dimensions below, then close the loop with the decision matrix.
4.1 Dimension 1: Your Background
| Background trait | Directions it favors | Why |
|---|---|---|
| MS/PhD with RL publications | Research scientist, alignment research, robotics research | Publications are the first hard currency for research roles |
| BS/MS from an engineering background (backend / distributed) | Simulation platform, RL infra, alignment engineering | Platform roles weigh engineering heavily |
| MS with recommendation / search / risk-control experience | Decision intelligence, ads strategy | Business experience transfers directly; MDP modeling is a bonus |
| PhD in operations research / optimization | Decision optimization, quant, scheduling | Mathematical modeling is the hard currency |
| Robotics / control background | Robotics learning, autonomous driving decision-making | Continuous control and simulation experience map one-to-one |
4.2 Dimension 2: Your Interest
- You love "proving an idea works" and enjoy being cited → research roles. The catch: the pressure of not publishing, and long experiment cycles.
- You love "shipping something for real and watching the metric move" → applied roles. The trade-off: business problems are rarely elegant, and RL is often not the best solution.
- You love "building platforms and squeezing ten thousand GPUs" → platform roles. The trade-off: far from the algorithms, and years of infra work can dull your algorithmic instincts.
Interest and reality often collide
A common tragedy: a theory-loving master's graduate lands on a pure business team and writes SQL to analyze retention all day; a hands-on engineer joins a research group and can't produce a decent experiment in three months. At the interview, ask about the team's composition and "what will I actually do in my first three months" — far more useful than the job title.
4.3 Dimension 3: Market Trends
| Trend (2026 view) | Roles that benefit | Risks |
|---|---|---|
| LLM alignment becomes standard | Alignment engineers, RLHF | The bar keeps dropping, competition is thick, and the dividend window is narrowing |
| Embodied AI funding boom | Robotics learning, simulation platform | Bubble risk; survival of some companies is uncertain |
| Decision intelligence seeps into recsys / risk control | Strategy algorithms, offline RL | The titles don't say "RL," so these roles are hard to spot |
| Autonomous driving converges on "planning-and-control safety" | Decision & planning | Only the top companies are still hiring aggressively |
| RL for reasoning (R1, RLVR) | Research roles, alignment roles | A paper-driven hotspot; real adoption hinges on inference cost |
4.4 The Decision Matrix
Once you've scored all three dimensions, close the loop with this matrix:
text
Background × interest × market trend — combined judgment
────────────────────────────────────────────────────────────
Background suits research × loves proving ideas × watching alignment/R1
→ Top pick: LLM alignment research / Research Scientist (alignment)
Background suits applied × loves shipping × watching decision intelligence
→ Top pick: strategy algorithm engineer / recsys · risk control · scheduling
Strong engineering background × loves building platforms × watching embodied AI
→ Top pick: simulation platform engineer / RL infra
Strong math background × loves modeling × watching quant / OR
→ Top pick: quant execution / decision optimization
────────────────────────────────────────────────────────────
If none of the four stands out: enter through #2 RL Algorithm Engineer (applied) —
the most forgiving on-ramp, since it demands no single axis to an extreme.5. Usage Map for the Five Pages
Finally, here is the full usage order of the career module's pages (five including this one), and two typical ways to run it:
Usage A: job-search sprint (pair with the job-search sprint track in guide/paths)
text
Day 1 Finish this page → lock in 1–2 target roles
Day 2 Read the [JD List](/career/jd-list) → pull 5 target JDs and compare line by line
Days 3–5 Use the [Skill Breakdown](/career/knowledge-map) to build a catch-up list → clear all P0 topics
Days 6–7 Revise the resume per [Resume Analysis](/career/resume-analysis) + fill gaps with a portfolio
Days 8–14 Self-test with the [Interview Question Bank](/career/interview-questions) → run the top questions three timesUsage B: long-term preparation (the slow track)
text
Start with this page to build a mental map of the market → follow the [Learning Paths](/guide/paths) systematic track
Every time you finish a concept module, come back and check the [Skill Breakdown](/career/knowledge-map)
Confirm "which JD skill term does this knowledge point map to" → align as you learn| Page | Deliverable | Depends on | Time budget |
|---|---|---|---|
| This page | A shortlist of 1–2 target roles | None | 30 minutes |
| JD List | Requirement excerpts from 5 target JDs | This page | 1–2 hours |
| Skill Breakdown | Your personal catch-up list | JD List | 1–2 hours |
| Resume Analysis | A revised resume + a portfolio README | Skill Breakdown | Half a day |
| Interview Question Bank | A checklist of high-frequency questions passed in self-test | Skill Breakdown | The week before interviews |
Further Reading
- Learning Paths: three tracks — the job-search sprint runs on the career module; this page shows the overall rhythm.
- JD List — breakdowns and word-frequency stats of real, live JDs; the next page to enter.
- Skill Breakdown — map JD skill terms to "fluent / half-fluent / can't" and generate a catch-up list.
- Resume Analysis — the four elements + STAR: how to present RL ability on a resume.
- Interview Question Bank — answer frameworks and follow-up predictions for high-frequency interview questions.
References
- levels.fyi (compensation database, 2026): https://www.levels.fyi
- LinkedIn Jobs (global job search): https://www.linkedin.com/jobs
- Glassdoor (job and compensation reviews): https://www.glassdoor.com
- BOSS Zhipin (job search in China): https://www.zhipin.com
- Lagou (internet-industry jobs): https://www.lagou.com
- Liepin (mid-to-senior roles): https://www.liepin.com
- OpenAI Careers (alignment / research roles): https://openai.com/careers
- DeepMind Careers (research / engineering roles): https://deepmind.google/about/careers/
- Anthropic Careers (alignment / research roles): https://www.anthropic.com/careers
- ByteDance Careers: https://jobs.bytedance.com
- Tencent Careers: https://careers.tencent.com
- Alibaba Careers: https://talent.alibaba.com
- Baidu Careers: https://talent.baidu.com