Appearance
Decision-Making in Autonomous Driving
In one sentence: this page pinpoints where reinforcement learning actually sits in the autonomous-driving stack — it doesn't run the perception layer; it plays a supporting role in the "behavior decision" and "planning" layers. It also explains why the industry mainstream is imitation learning + rules + search rather than end-to-end RL; how safety constraints (constrained RL, safe sets, redundant guardrails) reshape the problem; and it closes with an honest analysis of "why RL isn't yet running L4 self-driving."
1. The Autonomous-Driving Decision Stack and Where RL Fits
1.1 The Full Software Stack
The software aboard an L4 self-driving vehicle divides roughly into five layers:
text
Perception Prediction Decision & Planning Control
Detection / seg / occ. Trajectory prediction Behavior decision (change lane? Steering / throttle / brake
for other vehicles/peds overtake?) │
│ │ │ │
CNN/Transformer Probabilistic traj. Behavior layer + motion planning PID/MPC
(deep learning dominant) (deep learning + rules) (rules / search / optimization / RL)(classical control)There are two entry points for RL:
- Behavior decision layer: selects discrete maneuvers at the macro level — "change lane, follow, overtake, yield." This is RL's most natural landing spot: discrete actions, definable rewards, and a clear decision objective.
- Motion planning layer: generates smooth spatio-temporal trajectories under the behavior layer's constraints. Here the mainstream is search and optimization (RRT, lattice planners, gradient-based optimization), with RL tried as a way to "generate initial solutions" or to output trajectories end to end.
TIP
Focus on "which layer should RL replace" rather than "RL replaces the whole stack." The industry consensus: perception belongs to supervised learning, and control belongs to classical control theory. Only the behavior decision layer and parts of the planning layer offer RL any cost-effective opening. This aligns exactly with the "Boundaries of RL" section on the RL vs. Adjacent Paradigms page.
1.2 Why "End-to-End RL" Is So Seductive Yet So Hard
NVIDIA's DAVE-2 (Bojarski et al., 2016) showed that end-to-end imitation learning (behavioral cloning) mapping camera pixels → steering angle could drive hundreds of meters in simple scenes. Yet end-to-end approaches have never become the L4 mainstream, for these reasons:
| The selling point | The real price |
|---|---|
| No hand-built modules or feature engineering | A black box — safety certification becomes impossible |
| Directly data-driven | The data needed to "cover every long-tail scenario" is effectively infinite |
| End-to-end optimizable | Gradients get diluted along a long chain; errors accumulate between modules |
| One model solves everything | When a crash happens, you can't tell which step went wrong |
This tension between modularity and end-to-end is the main thread for understanding autonomous-driving technical routes. The industry consensus is a hybrid architecture: modular as the backbone, with imitation learning and RL assisting inside individual components.
2. RL in Hierarchical Planning: the Behavior Decision Layer
2.1 The Hierarchy of the Decision Layer
text
Global navigation (routing) → Behavior decision → Motion planning → Tracking control
Macro path from A to B "Choosing behaviors" Generate concrete Execute
at intersections / in lanes trajectories
(A* / lane network) (FSM / rules / RL) (sampling + opt. / search) (PID/MPC)The behavior decision layer takes "ego state + intent + scene semantics" as input and outputs a behavior intent for the next stretch (e.g. ChangeLaneLeft, FollowLane, YieldToPedestrian). This layer is a natural MDP:
text
State s : ego speed/position, target lane, neighboring-lane vehicle states, traffic lights, right-of-way
Action a : discrete intents — change lane / hold / yield / speed up / slow down
Reward r : task progress (arrival) − comfort penalty − safety-constraint penalty
Transition : the environment (reactions of other traffic participants)2.2 RL's Strengths and Struggles at the Behavior Layer
Strengths: compared with a hand-written finite state machine, RL is better at "long-tail interactions" — contested intersections, yield negotiations, unprotected left turns: the scenarios where rules can never be written exhaustively.
Struggles:
- Reward functions are hard to write. "Good driving" is a weighted blend of soft constraints (safety, comfort, efficiency, courtesy), and a slightly wrong weight skews the behavior — the most textbook instance of "the reward is the specification" from the reward engineering page.
- Safety cannot be guaranteed by "learning" alone: a learned policy carries no theoretical guarantees on corner cases it has never seen.
2.3 A Complete MDP Example: the Lane-Change Decision
Concretize behavior-decision RL as a "lane change" decision:
text
State s : ego speed/position, distance to and speed of the leading/trailing vehicles in the target
lane, proximity to the next intersection, user intent (go straight / turn left)
Action a: hold lane / change left / change right / decelerate and yield / accelerate
Reward r: +1 for arriving on time → − comfort penalty (harsh acceleration / hard braking)
− safety penalty (passing too close to obstacles)
− efficiency penalty (crawling below the speed limit for too long)
Terminal: navigation goal completed, or collision (huge negative reward)
Challenges:
- Surrounding vehicles are "dynamic obstacles"; transitions depend on other agents' behavior (non-stationary)
- Timing a lane change is adversarial (the car behind may accelerate to block you) — no hand-written rule produces a good policy
- Reward weights (safety vs. efficiency) define the personality: safety-biased → it yields forever, too timid to ever go;
efficiency-biased → it cuts in at willEvery difficulty in this example maps to an RL subproblem: non-stationarity → multi-agent RL; weight drift → reward engineering; evaluation → evaluation & benchmarks. Writing a decision problem down as an MDP first is a basic required skill for autonomous-driving RL research.
3. Imitation Learning First: the Limits of Behavior Cloning and the Hybrid Route
3.1 Why Autonomous Driving Started With Imitation Learning
L4 fleet logs (e.g., Waymo's tens of millions of road-test miles) provide massive expert demonstrations: the "correct actions" taken by human drivers or safety drivers in real traffic. So the industry naturally started with behavior cloning (BC) — plain supervised learning of state → action.
BC has three fatal limitations:
| Limitation | Mechanism | Consequence |
|---|---|---|
| Distribution shift (compounding error) | Training states come from expert trajectories; deployment states come from the agent's own mistakes | Small errors snowball into big accidents |
| Imitates but can't reason counterfactually | It never learns "what would happen if I took a different action" | No response to unseen scenes |
| Reward information is lost | It learns only what the expert did, not why | Unable to trade off safety vs. efficiency |
ChauffeurNet (Bansal et al., 2018, from Waymo) is precisely a cure for these root problems: beyond imitation learning, it introduced closed-loop training — placing the agent in simulated traffic and adding "errors at future timesteps" as extra supervision, forcing the policy to learn error correction instead of memorization.
3.2 The RL–Imitation Hybrid: the Industry's Mainstream Recipe
What works best in practice is "BC initialization + RL fine-tuning":
text
Step 1 Large-scale BC: train the initial policy from fleet logs (quickly acquire decent driving behavior)
Step 2 RL fine-tuning: use reward signals in simulation to optimize against "hard scenarios"
(learn trade-offs and long-tail responses)
Step 3 Safety guardrails: RL policy outputs must pass constraint checks (see the next section)One public end-to-end case is Learning to Drive in a Day (Kendall et al., 2019): first BC on just 10 minutes of single-camera driving data, then RL fine-tuning (PPO) in a mix of simulation and the real car, achieving end-to-end driving on a closed loop. It proved the "BC backbone + RL fine-tune" pipeline viable, and also exposed its limits (a single scene, no complex interaction). In real fleets, scenario selection for fine-tuning and the safety case remain open problems.
WARNING
The biggest pitfall of mixing imitation learning with RL is reward–demonstration conflict. If the reward function over-weights "comfort," RL fine-tuning pushes the policy toward "cautious slow driving," drifting away from the efficiency of the human demonstrations. The fix is to write the demonstration constraint (e.g., the KL distance to the BC policy) into the reward, keeping fine-tuning from drifting — an approach isomorphic to the "KL penalty against deviation" trick in RLHF.
3.3 Data Engineering: Fleet Logs, Simulation, and Augmentation
How much data the imitation + RL hybrid pipeline can consume determines its ceiling. Data stands on three legs:
| Source | Content | Quality | Volume |
|---|---|---|---|
| Real fleet logs | Complete driving by safety drivers and human drivers | Most realistic | Limited (road miles are expensive) |
| Simulated traffic flows | Parameterized scenarios (lane changes, yielding, congestion) | Controllable, labelable | Unlimited (just add GPUs) |
| Adversarial synthetic scenes | Hand-crafted corner cases | Rarest but most valuable | Generated on demand |
Key point: the "long tail" (rare scenarios) means real logs alone will never be enough — simulation synthesis is mandatory. But synthetic scenarios must be "plausible"; otherwise the policy learns the "traffic rules of the simulator" rather than the "traffic rules of the real world." This is the driving-flavored version of "the data distribution determines policy generalization" from the offline RL page.
4. Safety: Constrained RL, Safe Sets, and Redundant Guardrails
4.1 Treat Safety as a Constraint, Not a Reward
Writing "safety" into the reward has a classic problem in autonomous driving: a reward is a scalar-weighted sum, so safety becomes commensurable with everything else. A rule like "never trade any collision risk for arriving one minute earlier" cannot be expressed in a scalar reward — yet engineers will never permit such a trade. The right move is to treat safety as a hard constraint:
text
Constrained optimization objective:
maximize E[efficiency reward + comfort reward]
subject to P(collision) = 0; min distance between ego vehicle and obstacles ≥ d_min, etc.
Corresponding methodology:
- Constrained RL: e.g., CPO (Achiam et al., 2017) bakes constraints into the policy update
- Safe sets and control barrier functions (CBF): keep the state inside the safe set, in real time
- Redundant guardrails: RL/search outputs must pass a rule checker; violations fall back to safe behavior4.2 The "Three Layers of Safety" in Engineering
text
┌ Layer 1 Decision layer (RL / search / rules): outputs behavior intent
├ Layer 2 Planning constraints: trajectory must satisfy dynamic feasibility, collision-freedom, and comfort bounds (enforced by the optimization layer)
└ Layer 3 Execution guardrails: control-layer speed limits, steering limits, AEB emergency braking (independent of any decision component)Key insight: whether the decision layer is RL or rules, layers 2 and 3 must exist, and neither trusts layer 1. This engineering discipline is shared with robot RL (see the safety guardrails in Robot Control and Sim2Real).
4.3 The Constrained-RL Toolbox
If the decision layer does use RL, how do safety constraints enter training? Four mainstream families:
| Method | Idea | Characteristics |
|---|---|---|
| Lagrangian methods | Fold constraints into the reward as penalties; dual variables auto-tune the penalty weights | Simple and general, but the weights are hard to converge |
| Constrained policy optimization (CPO) | Directly constrain satisfaction during the policy update | Solid theory, complex implementation |
| Safe sets / control barrier functions (CBF) | Monitor actions in real time; project violations back into the safe set | Deterministic, but needs a model |
| Blocklists / rule filters | Hard-code forbidden actions (e.g., "driving against traffic") | Crudest, and the most reliable |
The engineering conclusion: the first two are for training time; the last two are for deployment time. No matter how much academia champions CPO, CBF/rule filtering at deployment is the non-negotiable last line of defense.
5. Real-World Cases: the Waymo and Tesla Technical Routes
5.1 Waymo: Search + Imitation First, RL Assisting
Waymo (formerly Google's self-driving project) released ChauffeurNet in 2018 and has continued to publish its architecture since. Public information gives this technical profile:
| Module | Waymo's approach |
|---|---|
| Perception | Deep neural networks (LiDAR + camera fusion, occupancy grids) |
| Prediction | Multimodal trajectory prediction (the Transformer family) |
| Planning | Rule-based search + trajectory optimization as the mainstay; RL assists by proposing "behavior candidates for hard scenarios" |
| Evaluation | Carcraft large-scale simulation (millions of virtual miles per day) + road tests |
Worth noting: Waymo's official papers and blog repeatedly stress that "safety comes from redundancy and verification, not from any single model." Its decision layer has never claimed "RL in charge."
5.2 Tesla: Data-Driven Shadow Mode and Occupancy Networks
Tesla follows a "pure vision + data-driven" route and revealed more details at AI Day 2022:
- Perception layer: Occupancy Networks turn the eight camera feeds into 3D grid-occupancy predictions, replacing part of the traditional "object detection + tracking" pipeline.
- Planning layer: FSD's planning leans on data-driven trajectory candidates + cost-function evaluation (optimization) rather than end-to-end RL; policy selection still keeps plenty of hand-written rules and constraints.
- Shadow mode: every car in the fleet continuously records "what would have happened had the system taken a given action," using the offline data pool as a "free simulator" for training and validation — the largest industrial-scale practice of offline-RL thinking (see offline RL).
5.3 A Table: RL Content Across Three Typical Routes
| Company / Project | Perception | Decision / Planning | RL's role |
|---|---|---|---|
| Waymo | Deep learning | Rules + search + optimization | Assisting (behavior candidates for hard scenarios, evaluation) |
| Tesla FSD | Deep learning (occupancy networks) | Data-driven candidates + cost optimization | Shadow mode / offline data application |
| NVIDIA DAVE-2 | End-to-end CNN | (end-to-end) | No RL, pure imitation |
| Academia (Kendall 2019 et al.) | Monocular camera | End-to-end policy | RL fine-tuning of demonstrations |
INFO
Low "RL content" doesn't mean the industry undervalues RL. Autonomous driving hands RL two special hard problems: learning under safety constraints, and using offline data at massive scale — the former gave rise to constrained-RL research, the latter handed offline RL its largest real-world dataset. Both directions are far from mature.
6. "Why Hasn't RL Grabbed the Wheel in L4 Self-Driving?": an Honest Analysis
6.1 Six Structural Reasons
| Reason | Explanation |
|---|---|
| Long-tail data | The "rare but fatal" scenarios an effective policy needs appear extremely infrequently in data; RL's trial-and-error sampling simply can't reach them |
| Safety certification | L4 demands "provable safety"; RL policies offer no probabilistic guarantees, and regulators and insurers won't accept a black box |
| Rewards are hard to write | The weights of safety / comfort / efficiency are product decisions, not optimizable variables; see reward engineering |
| Evaluation is hard | The real environment can't be sufficiently explored through trial and error, and the fidelity of simulation-based evaluation is itself a big pitfall (see the next section) |
| Interpretability and liability | Accident attribution requires an auditable explanation of "why it drove that way" |
| Cost and data access | Real road-test data is expensive; RL's sample demands don't match the physical world |
6.2 A Comparison Worth Remembering
Compare AlphaGo's success conditions (see AlphaGo and Monte Carlo Tree Search): a perfect model, free simulation, and a clear win/loss signal. Autonomous driving fails all three:
- The model is imperfect: traffic is an open world that can't be enumerated.
- Simulation is not free: no simulator captures the full complexity of real physics and human intent.
- There is no clear win or loss: driving has no "winning" — only "completing the task safely."
The conclusion is not "RL is useless," but "RL's applicability is narrower than advertised." Autonomous driving is precisely the best counter-example to the "when should I use RL" decision framework on the RL vs. Adjacent Paradigms page.
7. Simulation Testing and Evaluation: Closed-Loop vs. Open-Loop
7.1 The Lie of Open-Loop Evaluation
- Open-loop evaluation: replay historical trajectories and check whether the model's predicted action at a given moment matches the expert's. It's simple, but it hides distribution shift — a small deviation doesn't count as an error in open loop, yet in closed loop it snowballs into an accident.
- Closed-loop evaluation: put the model into a simulated traffic flow, let it drive itself, and measure task completion rate, accident rate, and violation rate. This is the industry standard.
| Evaluation method | Pros | Cons |
|---|---|---|
| Open-loop (log replay) | Cheap, reproducible | Weakly correlated with real deployment performance |
| Closed-loop simulation (Carcraft, MATSim, etc.) | Enables large-scale stress testing | Simulation-fidelity shortfalls mislead judgment |
| Road tests | Real | Sparse samples, extremely costly, long tail hard to cover |
7.2 The Fidelity Trap of Closed-Loop Simulation
The core risk of closed-loop simulation is "safe in simulation ≠ safe in reality": if the simulator's traffic participants behave too "gently," the policy learns aggressive behavior and fails the moment it hits the real road. This is the same problem in a different guise as the robotics field's Sim2Real gap. The industry countermeasure is adversarial simulation: actively generating hard scenarios (cut-ins, jaywalking pedestrians) to test the policy, rather than evaluating only on the natural distribution.
WARNING
Another easily overlooked pitfall in autonomous-driving evaluation: making product decisions with open-loop metrics. No "accuracy/similarity" open-loop metric can answer the question "can it go on the road?" The "ship or not" question can only be answered by safety statistics from closed-loop simulation plus road tests. This is consistent with the principle "the evaluation protocol determines the credibility of the conclusion" on the evaluation & benchmarks page.
7.3 The Metric Hierarchy for Simulation Evaluation
A car that "passes in simulation" rests on a whole layered system of metrics — don't fixate on a single number:
| Level | Metrics | Notes |
|---|---|---|
| Safety | Accident rate, emergency-braking trigger rate, safe-distance violations | Top priority |
| Task | Task completion rate, average travel time | Compared only after safety is satisfied |
| Comfort | Acceleration / hard-turn frequency, passenger-experience model | Affects product quality |
| Efficiency | Average speed, intersection waiting time | Trades off against comfort |
| Robustness | Variance across seeds, adversarial-scenario pass rate | Guards against "lucky runs" |
Reporting discipline: safety metrics come before everything else. If safety fails, no matter how pretty the task/efficiency numbers are, shipping is forbidden. This ordering is itself "evaluation as product decision" — exactly the "don't decide with one metric" refrain from the evaluation & benchmarks page.
8. Research Frontiers: RL in Autonomous Driving
A few directions worth following for anyone who wants to do RL in autonomous driving (plus a reminder of which ones are merely "narratives"):
| Direction | Content | Maturity |
|---|---|---|
| Safe constrained RL | Engineer CPO/CBF so that RL decisions carry safety guarantees | Active research, little deployment |
| Offline RL on fleet logs | Train and evaluate policies from shadow-mode data | Industrial practice exists (Tesla et al.) |
| End-to-end RL (perception → control) | Learn driving directly from pixels; highly exploratory | Research-grade; not adopted in production |
| World model + planning | Learn a driving-scene model, then plan inside it (model-based) | Frontier; see model-based RL |
| Multi-agent interaction | Model other cars as agents rather than static obstacles | Academic direction, little industrial use |
In one sentence: autonomous driving's "output" for RL is mostly methodology and infrastructure (simulation, evaluation, safety, offline data), not a road-ready end-to-end RL policy. Learning and investing with this expectation is far more pragmatic than chasing the "RL replaces the whole driving stack" narrative.
From Academia to Industry: Three Attempts at End-to-End RL Driving
A retrospective of representative "end-to-end RL driving" work in academia shows where this route actually ends:
| Work | Year | Approach | Result | Where the limit lies |
|---|---|---|---|---|
| NVIDIA DAVE-2 | 2016 | Camera → steering (pure BC) | Drives hundreds of meters in simple scenes | Distribution shift, no safety case |
| Learning to Drive in a Day | 2019 | BC pretraining + PPO fine-tuning | End-to-end driving on a closed loop | Single scene, no interaction |
| Imitation + RL hybrids (various OEMs) | 2020+ | BC backbone, simulation RL fine-tune, guardrails at the end | Usable within local components | Has not replaced the modular architecture |
A conclusion worth remembering: a decade down the end-to-end road, its most stable output is "end-to-end perception" (vision directly producing learnable representations), while "end-to-end decision-making" has never reached mass production. The reason is not that RL is too weak — it is that autonomous driving's safety-liability structure (see Section 7) dooms the "unexplainable end-to-end black box" to fail the safety case. This judgment holds for any "RL takes over a safety-critical system" proposal.
9. Summary: What an RL Engineer Can Do in Autonomous Driving
For RL engineers who want to enter autonomous driving, here is a pragmatic list of landing spots:
- Behavior-decision RL: train lane-change / intersection decisions in simulation (CARLA, MetaDrive, etc.), and study "RL under safety constraints."
- Offline RL: use massive fleet logs for policy training and evaluation — autonomous driving may be offline RL's largest real-world stage.
- Simulators and scenario generation: building adversarial scenarios and improving simulation fidelity are scarce industry skills.
- Hybrid pipelines: the engineering practice of BC initialization + RL fine-tuning + guardrails.
Don't sink your energy into storylines like "end-to-end RL replaces the whole stack" — that is a research-and-exploration topic, not an engineering topic you can land on day one after joining.
Further Reading
- Reward Engineering — "the reward is the specification," and how soft safety constraints distort behavior.
- Model-Based RL and World Models — how driving simulators relate to learned world models, and the model-error trap.
- Offline RL — the methodology behind Tesla's shadow mode.
- RL vs. Adjacent Paradigms — the full "when should I use RL" decision framework.
- Common Pitfalls & Antipatterns — concrete forms of open-loop-evaluation cheating and environment bugs in autonomous driving.
- Robot Control and Sim2Real — the robotics version of the same guardrail and sim-transfer problems.
References
- Bojarski, M., et al. (2016). End to End Learning for Self-Driving Cars. arXiv:1604.07316. (NVIDIA DAVE-2)
- Bansal, M., et al. (2018). ChauffeurNet: Learning to Drive by Imitating the Best and Synthesizing the Worst. arXiv:1812.03079.
- Kendall, A., et al. (2019). Learning to Drive in a Day. ICRA 2019. (arXiv:1807.00412)
- Achiam, J., et al. (2017). Constrained Policy Optimization. arXiv:1705.10528.
- Tesla AI Day 2022 (September 30, 2022) official presentation and blog, which publicized the occupancy network and planning architecture.
- Waymo official blog, "The Waymo Driver's Technical Approach" (continuously updated since 2020), and the ChauffeurNet technical report.
- Dosovitskiy, A., et al. (2017). CARLA: An Open Urban Driving Simulator. CoRL 2017. (open-source driving simulator)
- Grigorescu, S., et al. (2020). A Survey of Deep Learning Techniques for Autonomous Driving. Journal of Field Robotics, 37(3), 362–386.