Appearance
JD List
One-line summary: this page translates "who is the market actually hiring" into a list — line-by-line breakdowns of JDs for six representative roles, word-frequency stats on requirements, skill-mix profiles, and a reading method for "what to look at first in a JD."
1. Sources and Data Freshness (dataAsOf: 2026-08)
Data freshness disclaimer
Every JD breakdown on this page is based on a composite sample of live listings publicly visible around August 2026 on mainstream job platforms and company career sites, in China and abroad, then condensed into "typical JD profiles" (not the verbatim text of any single company — identifying details have been anonymized and merged). Requirements, pay, and HC (headcount) can change at any moment:
- Always defer to what you see on the day you apply on BOSS Zhipin, Lagou, Liepin, LinkedIn Jobs, Glassdoor, and company career pages;
- The value of this page is that it gives you a method for reading JDs: knowing which words are hard gates, which are filler, and which are the signals the team actually cares about — not serving as a job database;
- Word-frequency stats are approximations over the sampled listings (10–20 public JDs per role), not a census.
Before reading any RL-related JD, do three things (this page runs on that framework throughout):
- Mark the hard gates: degree, years of experience, must-know frameworks/algorithms — if you don't clear them, don't force the application; save the time.
- Mark the responsibility verbs: is it "develop new algorithms" (research), "deploy and tune" (applied), or "build training platforms" (platform) — classify the role type.
- Mark the signal words: detail words that recur in a JD (e.g., "multiple seeds," "reward hacking," "PPO clip") usually hint at what interviewers will actually ask — that's the input to the Skill Breakdown.
2. JD Breakdowns for Six Representative Roles
The six roles below cover the overwhelming majority of the RL hiring market. Each gets four lines: "typical responsibilities (composite) → hard requirements → nice-to-haves → how to read the interview signals."
2.1 Role 1: RL Research Scientist (LLM company / research org)
| Dimension | Content (composite of key points from multiple public JDs) |
|---|---|
| Responsibilities | ① Explore next-generation RL training paradigms (alignment, reasoning enhancement, world models); ② turn ideas into reproducible experiments and write papers / technical reports; ③ work with engineering teams to move new methods into the production pipeline; ④ track and reproduce frontier papers. |
| Hard requirements | PhD, or MS plus a first-author top-conference paper; solid mathematical grounding (probability, optimization, information theory); expert in PyTorch/JAX; experience training at scale (hundreds of GPUs or more); fluent at reading and reproducing papers. |
| Nice-to-haves | Hands-on RLHF/PPO at scale; experience with reasoning models (R1-style); experience with world models (DreamerV3, TD-MPC2); publications at NeurIPS/ICML/ICLR. |
| Reading the signals | "PhD or top-conference paper" is a hard gate — without a paper, you basically won't get the interview; expect heavy testing of math foundations and policy gradient derivations, plus the critical-thinking follow-up "what's your take on this paper?" |
2.2 Role 2: RL Algorithm Engineer (Applied RL Engineer, big tech / autonomous driving / gaming)
| Dimension | Content (composite of key points from multiple public JDs) |
|---|---|
| Responsibilities | ① Model business problems as MDPs / bandits and design rewards; ② train and tune with SB3 or in-house frameworks (PPO/SAC/DQN); ③ build the evaluation system and produce reproducible comparison results; ④ ship, monitor, roll back, and align on targets with product/ops. |
| Hard requirements | MS or above, 2–5 years of algorithm experience; solid PyTorch and common RL algorithm implementations; at least one complete production launch; basic data structures and engineering discipline (code review, unit tests). |
| Nice-to-haves | Hands-on offline RL or bandits; large-scale parallel training experience; simulator customization experience; deep understanding of the domain (games / trading / recommendation). |
| Reading the signals | JDs of this type keep repeating "production metrics" and "deep business understanding" — interviewers' biggest fear is a candidate who can only run demos. Your resume must carry the four elements "environment–algorithm–metrics–reproduction" (see Resume Analysis). |
2.3 Role 3: LLM Alignment Engineer (RLHF)
| Dimension | Content (composite of key points from multiple public JDs) |
|---|---|
| Responsibilities | ① Build SFT / preference data pipelines; ② train reward models (Bradley-Terry) and govern data quality; ③ fine-tune LLMs with PPO/DPO, handling KL blow-ups and reward over-optimization; ④ build alignment eval sets (helpfulness / harmlessness / factuality) and monitor them continuously. |
| Hard requirements | MS or above; expert PyTorch and distributed training (DeepSpeed/Megatron); understand the three-stage RLHF pipeline; experience training models of 7B parameters or larger. |
| Nice-to-haves | Hands-on DPO and its variants; experience with reasoning enhancement (RLVR, process rewards); evaluation-system experience; have read the InstructGPT/DPO papers. |
| Reading the signals | Interviews love deep-dive alignment questions like "why does PPO need a KL penalty?" and "the reward score went up — why did the model get worse?"; expect live testing on distributed-training engineering details too. |
2.4 Role 4: Robotics Learning Engineer (Embodied AI)
| Dimension | Content (composite of key points from multiple public JDs) |
|---|---|
| Responsibilities | ① Train policies in simulators like MuJoCo/Isaac/Genesis and transfer to real robots; ② build teleoperation and data-collection pipelines; ③ train VLA (vision-language-action) models or classic RL policies (SAC/PPO, Dreamer-style); ④ close the Sim2Real gap (domain randomization, system identification). |
| Hard requirements | MS or above; know at least one simulator; solid PyTorch; hands-on continuous control (SAC/PPO); able to debug a training pipeline independently. |
| Nice-to-haves | Real-robot deployment experience; Isaac Lab / RoboMimic / Diffusion Policy experience; world model experience; C++/ROS. |
| Reading the signals | "Real robot," "transfer," and "data efficiency" are the high-frequency words — expect questions on Sim2Real pitfalls and sample efficiency; hands-on tasks often take the form of a live diagnosis of "the training won't converge." |
2.5 Role 5: Recommendation / Ads / Search Strategy Algorithm (Decision Intelligence / Bandit)
| Dimension | Content (composite of key points from multiple public JDs) |
|---|---|
| Responsibilities | ① Use contextual bandits / RL to optimize ranking, bidding, and traffic allocation; ② design and run online experiments and exploration strategies (ε-greedy, UCB, TS); ③ build offline evaluation and replay systems; ④ define long-term metrics with the business side (retention, GMV, latency). |
| Hard requirements | MS or above; solid statistics and causal inference; fluent Python/Spark; experience in at least one of recsys / ads / search; know A/B testing methodology. |
| Nice-to-haves | Hands-on multi-armed bandits or offline RL; large-scale feature engineering; understand delayed feedback and position bias. |
| Reading the signals | These roles can hire perfectly well without ever writing "RL" — but an RL/bandit background is a clear plus. Interviews test the engineering side of exploration–exploitation (how to keep exploration costs in check) and hand you business modeling questions on the spot. |
2.6 Role 6: Autonomous Driving Decision & Planning (Behavior / Motion Planning)
| Dimension | Content (composite of key points from multiple public JDs) |
|---|---|
| Responsibilities | ① Design rule + learning hybrid policies for the behavioral decision layer (lane changes, merging, yielding); ② train and evaluate decision policies in closed-loop simulation; ③ handle safety constraints and edge cases; ④ collaborate with planning-control, prediction, and simulation teams. |
| Hard requirements | MS or above; foundations in planning and control (MPC, sampling-based planning); familiar with simulation testing and replay; fluent C++/Python. |
| Nice-to-haves | Imitation learning + RL hybrid experience; constrained RL or safe RL experience; large-scale simulation platform development. |
| Reading the signals | "Safety," "edge case," and "closed-loop evaluation" are the core words; expect pragmatic questions like "who catches the fall when the RL policy fails," and the difference between open-loop and closed-loop evaluation. |
About the four roles not covered
These six were chosen because they show up most in the sample. The other four (simulation platform infra, decision optimization / scheduling, quant execution, game AI) are positioned in Module Overview and Job Market Landscape: their JDs are similar in shape — platform roles amplify "responsibilities ②/③," optimization and quant roles are variants of Role 5, and game AI is Role 2 plus multi-agent.
3. High-Frequency Requirements: Word-Frequency Table
After running word-frequency statistics over the sampled JDs, we rank the RL-related hard-skill terms by occurrence, in three tiers: mandatory (appears in >70% of JDs), common (40–70%), and occasional (<40%).
| Tier | Skill term | Roles where it appears | Page on this site |
|---|---|---|---|
| Mandatory | PyTorch (or JAX) | All | framework-comparison |
| Mandatory | Deep learning fundamentals (CNN/Transformer, backpropagation) | All | what-is-rl and related sections |
| Mandatory | Distributed / large-scale training | Roles 1/3/4 | build-your-own |
| Mandatory | MDP / Bellman equations | All (basic questions) | markov-decision-process |
| Mandatory | Policy gradient / PPO | Roles 1/2/3/4 | policy-gradient |
| Common | RLHF / DPO / reward models | Roles 1/3 | rlhf |
| Common | SAC / TD3 / continuous control | Roles 2/4/6 | actor-critic |
| Common | Simulators (MuJoCo/Isaac/Gymnasium) | Roles 2/4/6 | datasets-tools |
| Common | Evaluation systems (seeds, learning curves, offline evaluation) | Roles 2/3/5 | evaluation-in-practice |
| Common | Exploration & exploitation (ε-greedy, UCB, TS) | Roles 2/5 | exploration-exploitation |
| Common | Reward design / reward hacking | Roles 2/3/4 | reward-engineering |
| Occasional | Offline RL (CQL/IQL) | Roles 2/5 | offline-rl |
| Occasional | Multi-agent (MARL) | Role 2 / gaming | multi-agent |
| Occasional | World models / model-based | Roles 1/4 | model-based |
| Occasional | Multi-armed bandits / contextual bandits | Role 5 | bandits |
| Occasional | Imitation learning / behavior cloning | Roles 4/6 | rl-vs-other-paradigms |
Frequency ≠ weight
High frequency only means "everyone loves writing it," not "it weighs heavily in interviews." What really separates candidates are the words in the common and occasional tiers that interviewers dig into: PPO's clip mechanism, SAC's entropy, the RLHF KL penalty, the engineering cost of exploration — these are the high-frequency exam topics in the Interview Question Bank. Don't spend your time on "having read the names of 20 algorithms"; build three levels of follow-up-proof depth on 4–5 algorithms.
4. Skill-Mix Profiles: Theory × Engineering × Business
Here are the six roles plotted by their emphasis on three axes (scores 1–5, relative values):
| Role | Theory (algorithms / math) | Engineering (systems / parallel / shipping) | Business (domain modeling / metrics) |
|---|---|---|---|
| Research Scientist | 5 | 4 | 2 |
| RL Algorithm Engineer | 3 | 4 | 4 |
| LLM Alignment Engineer | 3 | 5 | 3 |
| Robotics Learning Engineer | 4 | 4 | 3 |
| Recsys/Ads Strategy | 2 | 3 | 5 |
| AD Decision & Planning | 3 | 4 | 4 |
How to read it:
- No role scores 5 on all three axes — every role is "one axis spiking." When interviewers judge fit, they look at whether your spike lines up with the role's spike.
- The lowest-theory roles (recsys/ads) are paradoxically the easiest for RL people to land offers in: most competitors don't know RL, so your multi-armed bandit and offline RL knowledge reads as a scarce signal.
- The engineering axis never dips below 3 — the universal market signal of 2026: RL roles now demand "can run large-scale training, can ship, can reproduce" more than "knows many algorithms."
5. For Job Seekers: From JD to Action
5.1 The Five-Minute JD Scan
text
Minute 1 Hard gates — do you clear the bar (degree / years / must-have frameworks)? If not, move on
Minute 2 Classify the responsibility verbs: research / ship / platform → identify the archetype
Minute 3 Mark the signal words (e.g., "PPO clip," "multiple seeds," "reward hacking," "Isaac")
Minute 4 Score the JD's RL hard skills against the frequency table above; check overlap with your strongest axis
Minute 5 Decide: apply / apply after a resume tweak / drop it (saves you 3 hours)5.2 Three Common Mistakes When Reading a JD
- Mistake 1: reading the title, not the responsibilities. The same "RL engineer" is a research role on team A and a platform role on team B. Classifying by responsibility verbs is far more reliable than by title.
- Mistake 2: treating "nice-to-haves" as hard gates. JDs routinely fill an entire line with bonus items (world models, R1, VLA — want them all), while interviewers actually expect you to have one or two of them. The hard gates are the filter.
- Mistake 3: ignoring the "soft signal words." "Strong communication" and "fast learner" are filler, but phrases like "deep business understanding," "owns production metrics," and "ownership" hint at what the team genuinely cares about — mirror them in the right places of your resume.
5.3 The Action Loop After Reading a JD
Dissecting a JD is not about "understanding" it — it's about producing action items: copy out every RL skill term in the JD and feed them to Breaking Down JD Skills to map them to pages and a fluent/not self-assessment, producing a catch-up list. The pipeline's entry point is the target role you locked in at Module Overview and Job Market Landscape; its exit is the self-test checklist from the Interview Question Bank — the five pages are one machine; don't read just one of them.
Further Reading
- Skill Breakdown — map the skill terms dissected on this page to site pages and your personal catch-up list.
- Learning Paths: three tracks — the job-search sprint runs on the career module; this page lays out the two-week rhythm.
- Module Overview and Job Market Landscape — back to the role panorama and the three-dimension decision checklist to confirm your target role.
- Resume Analysis — how to land the "environment–algorithm–metrics–reproduction" that JDs care most about on your resume.
References
- BOSS Zhipin (live listings in China): https://www.zhipin.com
- Lagou (internet-industry jobs): https://www.lagou.com
- Liepin (mid-to-senior roles): https://www.liepin.com
- LinkedIn Jobs (global listings): https://www.linkedin.com/jobs
- Glassdoor (job and company reviews): https://www.glassdoor.com
- levels.fyi (overseas total-compensation reference): https://www.levels.fyi
- OpenAI Careers: https://openai.com/careers
- DeepMind Careers: https://deepmind.google/about/careers/
- Anthropic Careers: https://www.anthropic.com/careers
- ByteDance Careers: https://jobs.bytedance.com
- Tencent Careers: https://careers.tencent.com
- Alibaba Careers: https://talent.alibaba.com