Appearance
Breaking Down JD Skills
One-line summary: this page translates every skill term in a JD into a concrete page and a concrete exam topic on this site — the master mapping table for 30+ terms, a four-quadrant self-assessment (fluent / half-fluent / can't / not needed), a personal catch-up list template, and a weight ranking of high-frequency exam topics.
The skill terms you copied from the JD List are still just a pile of nouns. This page does three things: translate them into knowledge points → make you assess yourself honestly → produce a prioritized catch-up list. That list is your battle map for the next two weeks of study, and it feeds the Interview Question Bank.
1. The Master Mapping Table: Skill Terms → Pages (30+ Terms)
The table below groups the 30+ most frequent JD skill terms into four categories. For each term it lists what it actually tests and the page on this site that covers it (with the Glossary as a quick-lookup entry).
Category A: RL Theory (the Main Battleground in Interviews)
| JD skill term | What it actually tests | Page on this site |
|---|---|---|
| MDP / Markov decision process | The five-tuple, the Markov property, discounted return, policies and value functions | markov-decision-process |
| Bellman equations | The recursions for state value / action value, the principle of optimality | markov-decision-process |
| Q-learning | Off-policy TD updates, convergence intuition | value-based |
| SARSA | On-policy TD, the differences from Q-learning | value-based |
| TD / temporal-difference learning | TD(0) updates, the TD error, bootstrapping, MC vs TD | value-based |
| DQN and variants | Experience replay, target networks, Double/Dueling/Rainbow | value-based |
| Policy gradient | Intuition for the policy gradient theorem, REINFORCE, baseline/advantage | policy-gradient |
| PPO | The clipped objective, why it's stable, the relationship to TRPO | policy-gradient |
| GAE | Generalized advantage estimation, the meaning of λ | policy-gradient |
| Actor-Critic | The division of labor between the two networks, A2C/A3C | actor-critic |
| SAC | The maximum-entropy objective, automatic temperature tuning, why it excels at continuous control | actor-critic |
| TD3 / DDPG | Continuous actions, twin Q networks, delayed updates | actor-critic |
| Exploration & exploitation | ε-greedy, UCB, Thompson sampling, entropy regularization | exploration-exploitation |
| Multi-armed bandits / contextual bandits | Regret, UCB1, posterior sampling, deploying in recommendation | bandits |
| Reward design / reward hacking | Reward shaping, sparse rewards, cases of gaming the metric | reward-engineering |
| Offline RL | OOD actions, value overestimation, CQL/IQL, when not to use it | offline-rl |
| Multi-agent | Non-stationarity, CTDE, MADDPG/QMIX | multi-agent |
| World models / model-based | MPC, Dreamer, TD-MPC, model error | model-based |
Category B: RLHF / LLM Alignment (Hottest in 2026)
| JD skill term | What it actually tests | Page on this site |
|---|---|---|
| The three RLHF stages | SFT→RM→PPO, why SFT alone is not enough | rlhf |
| Reward models / Bradley-Terry | Training on preference pairs, score calibration | rlhf |
| KL penalty | KL against the reference model, why it has to be there | rlhf |
| Reward over-optimization / Goodhart | The alignment tax: scores rise, quality falls | rlhf |
| DPO | The closed-form preference objective, trade-offs vs PPO | rlhf |
| Alignment evaluation | Helpfulness / harmlessness / factuality, eval sets | rlhf and evaluation-benchmarks |
| SFT & preference data pipelines | Data quality, dedup, quality tiering | LLM alignment case study |
| Reasoning-enhancement RL (RLVR, R1) | Verifiable rewards, process rewards | frontier papers |
Category C: Deep Learning & Coding
| JD skill term | What it actually tests | Page on this site |
|---|---|---|
| PyTorch / JAX | Tensors, autograd, modular implementation | framework-comparison |
| Distributed training | DP/PP/TP, DeepSpeed, gradient synchronization | build-your-own |
| Backpropagation / optimizers | The chain rule, Adam, vanishing gradients | math-primer |
| Transformer / attention | Self-attention, positional encoding (must-know for alignment roles) | rlhf background sections |
| Experiment management | Configs (Hydra), logging, W&B/TensorBoard | evaluation-in-practice |
| Code discipline | Unit tests, code review, CI | build-your-own |
| Gymnasium API | reset/step, obs/action spaces | gymnasium-tutorial |
Category D: Engineering & Simulation, Business & Math
| JD skill term | What it actually tests | Page on this site |
|---|---|---|
| MuJoCo / Isaac / simulators | Physics engines, domain randomization, Sim2Real | datasets-tools and robotics-sim2real |
| Distributed sampling / parallel training | Environment parallelism, vectorization, GPU sampling | build-your-own |
| Evaluation protocol | Fixed seeds, multiple runs, learning curves, IQR | evaluation-in-practice |
| Benchmarks (Atari/MuJoCo/Procgen) | The environment landscape and their respective pitfalls | evaluation-benchmarks |
| Probability & expectation | Expectation, conditional expectation, variance, KL divergence | math-primer |
| Optimization basics | Convex optimization, stochastic gradients, Lagrange multipliers | math-primer |
| Business modeling | Writing business metrics as MDPs / rewards, evaluating business impact offline | reward-engineering and the case-study pages |
| Statistics & causality (strategy roles) | A/B testing, bias, confidence intervals | bandits, the deployment sections |
| C++ / ROS (robotics roles) | Real-robot pipelines, communication | robotics-sim2real |
How to use the mapping table
This table is the index from "JD term → site page," and the patch panel that wires this module to the three big content modules — concepts, practice, and papers. Whenever you copy an unfamiliar term out of a JD, check here first to see which category it belongs to and what it tests, then decide how much time to invest.
2. The Four-Quadrant Self-Assessment: Fluent / Half-Fluent / Can't / Don't Need
Run an honest four-quadrant classification on every term in the mapping table. The honesty of this classification directly determines the quality of your catch-up list — overestimating yourself is the number-one cause of interview blowups.
| Quadrant | Criterion | Action |
|---|---|---|
| Fluent | Survives three levels of follow-up (what → why → edge cases/pitfalls) and can write the core formulas by hand | Work it into resume project details; focus on showcasing |
| Half-fluent | Have heard of it, can give the gist, but choke on formulas/details as soon as someone digs | The main battlefield of catch-up: re-read the matching page + self-test with the question bank |
| Can't, but needed | Explicitly required by the JD or high-frequency, and you haven't studied it at all | Study systematically: enter the matching module in Learning Paths |
| Not needed | Occasional, a nice-to-have, and unrelated to your target direction | Ignore for now; park it in "for later" |
Doing the Four-Quadrant Self-Assessment
text
Step 1 Copy every skill term from your target JD in the [JD List](/career/jd-list) (about 15–30 terms)
Step 2 Tag each term with a quadrant (fluent / half-fluent / can't / not needed) and write it into the table below
Step 3 For every "half-fluent" and "can't" term, look up the matching page in the mapping table
Step 4 Generate the catch-up list (template in the next section)
Step 5 Re-assess every time you finish one: half-fluent → fluent is the only way to count it doneTwo traps in self-assessment
- Trap one: mistaking "read the title" for "fluent." The criterion is not "I've seen this algorithm" but "I can survive three levels of follow-up." Use the follow-up lists in the Interview Question Bank as a checkup — far more reliable than gut feel.
- Trap two: conflating "fluent" with "can implement fluently." Being able to derive the formulas doesn't mean you can write them in PyTorch; interviews often include a handwritten-pseudocode round. For anything tagged "fluent," at least get one matching implementation running in gymnasium-tutorial.
3. Catch-Up List Template
Here is a copy-ready catch-up list template. Priorities are marked P0/P1/P2: P0 = a must-have term for your target role that you currently can't or are half-fluent in; P1 = a common term; P2 = a nice-to-have.
| # | Skill term | Quadrant | Page | Priority | Where exactly are you stuck | Planned action | Deadline | Re-assessment |
|---|---|---|---|---|---|---|---|---|
| 1 | PPO | Half-fluent | policy-gradient | P0 | Knows the clip, can't explain why it's stable | Re-read the page + write pseudocode by hand + run it in SB3 | This Saturday | Pending |
| 2 | The three RLHF stages | Can't, but needed | rlhf | P0 | Haven't studied it at all | Read systematically + work through the LLM alignment case | Next week | Pending |
| 3 | Reward design | Half-fluent | reward-engineering | P1 | Only knows the term "reward hacking" | Read the case collection + do two design exercises | Within two weeks | Pending |
| 4 | SAC | Fluent | actor-critic | P1 | Can derive the entropy objective | Write it into a resume project and prep for follow-ups | — | Fluent |
Rules for filling in the template:
- Current-status notes must be concrete: for "half-fluent," pin down "where exactly do I get stuck" — the formula? the implementation? the why? — otherwise there's no target to aim at during catch-up.
- Planned actions must land on specific pages of this site: every "can't" needs an explicit page or module entry point, not empty words like "go study it."
- Deadlines must obey the job-search timeline: along the job-search sprint track in Learning Paths, all P0 items should be cleared within a week.
4. Weights of High-Frequency Exam Topics
Combining the word-frequency stats from the JD List with the probability of questions showing up in interviews, here is a weight ranking of exam topics. This is the answer to "what to learn first when time is short."
| Weight | Topic | Typical form | Page |
|---|---|---|---|
| ★★★★★ | MDP and Bellman equations | Hand derivations, concept follow-ups — near-guaranteed | markov-decision-process |
| ★★★★★ | MC vs TD, Q-learning vs SARSA | Comparison-question regulars | value-based |
| ★★★★★ | PPO (motivation for clip, mechanism, pitfalls) | The highest of the high-frequency | policy-gradient |
| ★★★★☆ | DQN engineering tricks (replay / target networks) | Guaranteed in deep RL | value-based |
| ★★★★☆ | The three RLHF stages and the KL penalty | Guaranteed for LLM roles | rlhf |
| ★★★★☆ | Exploration & exploitation | Applied + conceptual questions | exploration-exploitation |
| ★★★☆☆ | SAC maximum entropy and continuous control | Guaranteed for robotics/control roles | actor-critic |
| ★★★☆☆ | The Actor-Critic architecture | Shows up bundled with PPO | actor-critic |
| ★★★☆☆ | Reward-design scenario questions | Common for business roles | reward-engineering |
| ★★☆☆☆ | Offline RL / bandits | A plus for strategy roles | offline-rl, bandits |
| ★★☆☆☆ | Multi-agent / world models | A plus for research roles | multi-agent, model-based |
Weights drift by role
The above is the "whole-market average weight." Concrete roles drift noticeably: robotics roles push SAC/Sim2Real to ★★★★★, LLM roles push RLHF/DPO to ★★★★★, and strategy roles push bandits/causality to ★★★★★. Adjust the weights to your target role during self-assessment instead of memorizing this average table.
5. Connecting to the Interview Question Bank
The finish line of the catch-up list is the Interview Question Bank. The way to connect them is learning driven by questions:
- Every time you finish a P0 topic, find the matching questions in the question bank and run one round with the "answer framework + follow-up prediction";
- Use the question bank's "self-test checklist" (a 30-question tick list) as the acceptance criterion for the catch-up list: count something "fluent" only when every box is ticked;
- Questions you can't answer in the question bank route back to the mapping table for re-tagging — self-assessment is a loop:
text
JD skill terms → mapping table → four-quadrant self-assessment → catch-up list (P0/P1/P2) → question-bank self-test
↑ │
└────────────────────────── can't answer → re-assess ──────────────────────────────────────────┘The catch-up list this page produces is the only bridge between the target role you locked in at Module Overview and Job Market Landscape and your actual gaps. The other end of that bridge is Resume Analysis — once you've caught up, write the "fluent" items into the resume's four elements.
Further Reading
- JD List — where the skill terms come from: JD breakdowns and word-frequency stats for six roles.
- Interview Question Bank — the acceptance criterion for the catch-up list: high-frequency questions + answer frameworks + self-test checklist.
- Markov Decision Process (MDP) — the shared foundation under almost every skill term; the first of the P0 topics.
- Value Learning — the full thread of Q-learning/SARSA/DQN; the main battlefield of comparison questions.
- Policy Gradient — the home of PPO: the clip mechanism, GAE, REINFORCE.
- RLHF and Alignment with Human Feedback — where every exam topic for LLM alignment roles lives.
- Glossary — 60+ quick-reference entries; keep it open while self-assessing.
References
- Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction, 2nd ed. — the authoritative source for the theory exam topics, free online: http://incompleteideas.net/book/the-book-2nd.html
- OpenAI Spinning Up in Deep RL — algorithm cheat sheets and code; a reference implementation to check your PPO/SAC against: https://spinningup.openai.com
- Hugging Face Deep RL Course — concepts plus companion code: https://huggingface.co/learn/deep-rl-course
- Gymnasium official docs — the environment API and the benchmark environment list: https://gymnasium.farama.org
- Stable-Baselines3 docs and code: https://github.com/DLR-RM/stable-baselines3