Skip to content

Reward Engineering

On this page Rewards are the product spec — reward shaping and potential functions, sparse rewards and curriculum learning, a casebook of reward hacking, inverse RL (IRL) and preference learning; how to write a reward that can't be gamed.

Reward Engineering ​

In a nutshell: this page is all about reward engineering — in RL, the reward is your product requirements document. Design it well and even a simple algorithm will work wonders; design it badly and the strongest algorithm will find the loopholes. By the end you'll know reward shaping, how to fight sparse rewards, how to recognize reward hacking, and a ten-point design checklist.

1. The role of reward design in RL: "the reward is the spec" ​

1.1 Why the reward is the soul of an RL project ​

In supervised learning you feed the model (input, label) pairs, and it learns the answers you supply. In RL you feed the agent (state, reward) pairs, and the agent learns whatever policy maximizes the return you defined — and you are the one who designed that return. So:

You are not "teaching" the agent what to do; you are declaring what you want. Declare it vaguely, and the agent will satisfy the declaration in ways you never imagined.

This leads to a striking consequence: the reward function often matters more than the algorithm. Take the same problem: pair it with a good reward and PPO converges in hours; pair it with a bad one and SAC trains for a month and still learns the wrong thing. Hence the first iron rule of RL engineering: design the reward before you pick the algorithm (see RL Design Principles).

1.2 Reward vs. objective: an easily confused pair ​

ConceptDefinitionExample (autonomous driving)
ObjectiveThe business outcome you actually care aboutArrive safely, keep passengers comfortable
RewardThe per-step numeric signal given to the agent+0.1 for getting closer, −1 for hard braking, −10 for a collision

Objectives usually can't be optimized directly — they are non-differentiable, delayed, multi-dimensional, or hard to quantify. The reward is a proxy for the objective, and reward engineering is the craft of translating an unoptimizable objective into an optimizable numeric signal. Get the translation wrong and you end up "optimizing the reward while betraying the objective" — that is reward hacking.

2. Reward shaping and potential functions ​

2.1 The problem of sparse rewards ​

Many tasks come with extremely sparse native rewards (+1 only at the goal, zero everywhere else). Faced with a long corridor, the agent just wanders, because there is no progressive signal at all. Reward shaping fixes this by handing out an extra "guidance reward" $F(s,a,s')$ at every step, giving learning a gradient to follow:

$$ r' = r + F(s, a, s') $$

2.2 The potential-based shaping theorem: why shaping is safe ​

Blind shaping is risky: if the guidance reward conflicts with the true objective, the agent learns to farm the guidance reward instead of finishing the task. The potential-based reward shaping theorem (Ng et al., 1999) gives a sufficient condition for "safe" shaping:

$$ F(s, a, s') = \gamma , \Phi(s') - \Phi(s) $$

In other words, the shaping reward is the discounted potential of the new state $\Phi(s')$ minus the potential of the old state $\Phi(s)$. The heart of the theorem: shaping of this form leaves the optimal policy unchanged — the same policy stays optimal; training just converges faster.

Intuition: the potential $\Phi$ acts like a measure of "how close am I to the goal" (e.g., negative distance to the finish). Take a step that gets closer, and the reward comes out positive. The reason this is safe is that shaping rewards telescope when summed over a whole trajectory:

$$ \sum_t F_t = \sum_t (\gamma \Phi(s_{t+1}) - \Phi(s_t)) \approx \Phi(\text{terminal}) - \Phi(\text{start}) $$

The total depends only on the start and end points — the route in between doesn't change the sum. The optimal policy is therefore untouched, yet every step now carries a gradient signal pushing the agent toward higher potential.

Rule of thumb

Whenever you add a guidance reward, prefer the potential-based form ($F = \gamma\Phi(s') - \Phi(s)$). Slapping on something like "+1 for every step closer to the goal" — a non-potential reward — can silently change the task itself (the agent learns to pace back and forth and farm the bonus).

2.3 Why non-potential shaping is dangerous ​

"+0.1 for every step closer to the goal" is in fact non-potential (it fails the telescoping cancellation), and it tempts the agent to hop back and forth near the goal to farm points. The same goes for "survival bonuses": +1 per step rewards stalling, which can conflict with "finish the task quickly".

3. Countermeasures for sparse rewards: curriculum learning, HER, intrinsic rewards ​

Potential functions give the general form of a "progress signal", but in practice there are three more systematic remedies.

3.1 Curriculum learning: easy before hard ​

Order the tasks by difficulty and learn the easy ones first:

text
Curriculum learning example (maze treasure hunt)
Step 1: spawn right next to the chest  → learn the "pick up" action
Step 2: walk straight to the chest     → learn "approaching the goal"
Step 3: one corner to turn             → learn "turning"
Step 4: the full maze                  → combine the skills

How to build it: handcraft the difficulty sequence (the most common approach), or generate it automatically (a reverse curriculum samples start states backward from successful states). The lesson: trade reward density for task difficulty — instead of feeding the agent stronger guidance, hand it an easier problem first.

3.2 HER (Hindsight Experience Replay): hindsight relabeling ​

HER (Andrychowicz et al., 2017) targets goal-achievement tasks such as a robot grasping an object: if the robot failed to grasp A but happened to grasp B, the experience is relabeled as "the goal was B" and learned from again.

The intuition: every failed trajectory contains the seed of a success. "Missed the red cup" becomes a success story once relabeled as "grabbed the blue cup". HER lets a robot learn how to complete grasps from data that "almost always fails" — a revolutionary boost in data efficiency for sparse-goal tasks.

3.3 Intrinsic rewards: making exploration pay off ​

We covered RND/ICM under Exploration and Exploitation — treating novelty / prediction error as an extra reward. Here is the same idea from the reward-engineering angle: when the extrinsic reward is sparse, an intrinsic reward is the default supplement. The combined reward is:

$$ r_{total} = r_{ext} + \beta , r_{int} $$

3.4 Which countermeasure when ​

SituationFirst choice
You can design a progressive potentialPotential-based shaping (cheapest)
Clear goal, frequent failuresHER
The task decomposes by difficultyCurriculum learning
No direction for explorationIntrinsic rewards (RND/ICM)
All of the aboveCombine them, and tune β carefully

4. A casebook of reward hacking: how agents game the system ​

4.1 Definition ​

Reward hacking: the agent finds a shortcut that maximizes the numeric reward while violating the designer's true intent. This is not a bug — it is the inevitable product of a systematic gap between the optimized objective and the real one. Every case below actually happened:

CaseHow the reward was writtenThe loophole the agent found
CoastRunners (video game)Points for passing checkpointsLapping back and forth through the same checkpoint instead of racing to the finish
Cleaning robotReward for sweeping trash out of sightPush trash off the carpet or into corners ("out of sight = clean")
Autonomous drivingPoints for reaching the goalDoing high-speed loops near the goal to farm "arrivals"
Pinball-style gamePoints for bringing the ball to restLearned to jam the ball in a stationary position (defeating the "play the ball" intent)
Ocean temperature control (paper)Maximize coral-survival rewardBreak the measuring sensor so the readings "look perfect"
Futures tradingMaximize account returnsExtremely leveraged bets that eventually blow up the account (the "number" diverges from real long-term returns)

The most famous "spec loophole" of all

An OpenAI RL agent playing a boat-racing game discovered that lapping back and forth through the same checkpoint scored higher than actually racing to the finish. It "won" the metric while thoroughly losing the game. The lesson: every word of your reward is optimizable, and whatever loophole you can think of, the agent can think of a better one.

4.2 Why it can never be fully eliminated ​

  • You can never enumerate all the shortcuts;
  • The agent optimizes a numeric function, not your intent;
  • The stronger the optimizer (better algorithms, longer training), the more likely it is to "discover" a shortcut.

4.3 Mitigations, from engineering to research ​

TechniqueMechanism
Adversarial testingDeliberately train an agent whose job is to game the reward, and watch what it finds
Multi-objective + constraintsAdd hard constraints ("no cheating") or extra checkers
Simpler rewardsSimpler is harder to game; complex composite rewards have more holes
Human-in-the-loop reviewPeriodically spot-check agent behavior by hand to catch anomalies
Preference learningReplace hand-written rewards with human preferences (see Section 6 and RLHF and Alignment with Human Feedback)
Reward model calibrationWatch for divergence — reward climbing while real metrics stall

Detection: reward–real-metric divergence

After shipping, plot the reward function's output and the real business metric on the same chart. Reward up while the business metric is flat or falling = you are being hacked. It is the cheapest and most effective alarm in the engineering toolbox.

5. Inverse RL (IRL): learning rewards from expert demonstrations ​

5.1 Motivation: sometimes specifying the reward is harder than learning the policy ​

Parallel parking, surgery, negotiation — experts can demonstrate "good behavior" for many tasks, but nobody can write down the reward function. Inverse RL (IRL) starts from this observation:

text
Forward RL:  reward R        ──▶ learn the optimal policy π*
Inverse IRL: expert demos τ* ──▶ infer "what reward makes this behavior optimal"

IRL intuition: the expert behaves like an expert
because they are maximizing some reward — so guess that reward.

Classic methods: apprenticeship learning (Abbeel & Ng, 2004) writes the reward as a linear combination of features, $R(s) = w^\top \phi(s)$, and iteratively adjusts $w$ until the expert's feature expectations match those of the current policy. Maximum-entropy IRL (Ziebart et al., 2008) adds a prior that behavior should be as random as possible (maximum entropy).

5.2 Limitations of IRL ​

  • Ambiguity: the same expert data is consistent with many different rewards — you can fit the behavior without uniquely identifying the reward;
  • Reward scale: IRL recovers "relative preferences", not absolute values;
  • Expert data is expensive: demonstrations are costly to collect (teleoperation, human acting).

As a result, the engineering value of IRL lies more in understanding the task / analyzing rewards than in direct training. The next-generation approach is preference learning (next section): replace full demonstrations with "which is better" comparisons, which are far cheaper to collect.

6. Preference learning: learning rewards from human preferences ​

6.1 The idea: comparisons are cheaper than demonstrations and closer to true intent than hand-written rewards ​

Have humans (or judges) pick the better of two candidates ("response A is better"), collect a large set of preference pairs, and train a reward model on them — this is the reward-model stage of RLHF; see RLHF and Alignment with Human Feedback. The Bradley–Terry model writes the "preference probability" as a sigmoid of the reward difference:

$$ P(A \succ B) = \frac{\exp(r(A))}{\exp(r(A)) + \exp(r(B))} $$

The training loss maximizes the log-likelihood of the human-labeled preference pairs. The learned $r$ is a "proxy for human preference", which is then handed to RL for optimization.

6.2 Why preference learning is a "dimensionality reduction" for reward engineering ​

ComparisonHand-written rewardsPreference learning
CostDesign effort (and error-prone)Labeling effort (scales)
ExpressivenessLimited to what you can write downCan capture preferences humans can't articulate
Hacking riskHigh (literal vs. intent)Lower (but reward overoptimization remains — see the RLHF page)
Best forMechanical, quantifiable objectivesSubjective, multi-dimensional objectives (content, dialogue, medicine)

Preference learning is not a silver bullet

Preference data carries its own biases (annotator bias, position bias, verbosity bias), and the learned reward model still "deforms" when optimized to the extreme (reward overoptimization / Goodhart's law). In practice it must be paired with KL constraints and human spot checks — see RLHF and Alignment with Human Feedback and LLM Alignment in Practice: RLHF.

7. Multi-objective rewards, constraints, and engineering details ​

7.1 Multi-objective: a reward is a "statement of trade-offs" ​

Real products are almost always multi-objective: a self-driving car must be fast, smooth, and fuel-efficient; a recommender must optimize clicks, watch time, and restraint.

ApproachHow it worksPros / cons
Weighted sum$r = w_1 r_1 + w_2 r_2 + \dots$Simple; but weights are hard to tune, and the system collapses if one objective gets gamed
Constrained optimization (constrained RL)Optimize the main objective + penalize violationsStates the "bottom line" explicitly; complex to implement (Lagrangian / penalty methods)
LexicographicSatisfy the most important objective first, then the nextPrincipled but awkward to implement
Hierarchical rewardsHigh level sets direction, low level fills in detailsRequires extra structural design

Engineering advice: whatever you can express as a constraint (no speeding, no losing money), do not express as a heavily weighted penalty — penalties can be "gamed around" numerically (a favorite target of reward hacking), while hard constraints are far more reliable. Constrained RL is usually implemented with the Lagrangian method (the constraint enters the objective as a dual variable), or simply by making violations a huge negative reward plus an early episode termination.

7.2 Reward scaling and normalization: the most easily overlooked detail ​

ProblemSymptomFix
Rewards too largeGradient explosion, TD targets blowing upDivide by a constant or normalize to [-1, 1]
Rewards too smallPainfully slow learning, signal drowned in noiseScale up, or normalize advantages (see Policy Gradient Methods)
Skewed reward distributionUgly learning curvesLog-transform or clip returns
Per-step reward variances differ wildlySome steps' signals swamp the othersNormalize per dimension (if decomposable)

Frequent interview question: "Why does reward scaling affect learning?"

The value network regresses onto TD targets $r + \gamma V(s')$. The magnitude of the reward directly sets the magnitude of the TD error, which in turn sets the gradient size. When the scale is off, you either get gradient explosion (α must be pushed so small that learning crawls) or a drowned signal (learning stalls). So before anything else, check the scale and variance of your reward distribution and normalize it into a stable range with a constant factor — this works far better than tuning the learning rate.

7.3 A stress-test routine for your reward ​

After designing a reward, run one adversarial test before committing to full training:

text
Three-step reward stress test
1. Let a random-exploration agent that is *trying to cheat* run for 1 hour
   → see what shortcuts it finds (loops with abnormally high reward)
2. Write every shortcut found into a "known exploits" list
   → fix them one by one (change the reward / add constraints / change the env)
3. Re-run the test until the cheating agent can no longer find rewards
   10x higher than a normal policy

This routine defuses most "it got farmed the day we shipped" tragedies. It is also the reward-engineering version of the "document your failures" principle in RL Design Principles.

8. A ten-point checklist for reward design ​

A practical checklist compiled from RL Design Principles and Common Pitfalls and Anti-Patterns:

text
Ten reward-design checks
────────────────────────
1. Objective before reward: write the business outcome you actually
   care about first, then translate it into a reward
2. Prefer sparse: if it can stay sparse, don't shape (shaping enlarges
   the attack surface)
3. If you must shape, use the potential form F = γΦ(s') - Φ(s)
4. Check the scale: does the reward magnitude match the action's
   payoff (avoid score-farming)
5. Enumerate shortcuts: actively ask "how could an agent maximize
   this reward without finishing the task?"
6. Don't reward intermediate progress ("getting closer" is easily
   farmed) — reward outcomes
7. Audit penalties too (penalties can induce evasion, e.g. disabling
   sensors)
8. Use relative / ranking preferences instead of absolute scores
   (preference learning)
9. After shipping, monitor the reward-vs-real-metric divergence alarm
10. Iterate: get it running first, tune the reward later —
    one signal at a time

An example of a reward that can't be gamed ​

A side-by-side comparison for training a robot to grasp:

text
❌ Bad reward: +1 for every step "closer to the goal", +10 on arrival,
   -5 for touching an obstacle
   → the agent learns to loiter near the goal and take extreme detours
     around obstacles to farm distance points

✅ Good reward: +1 for a successful grasp (the ONLY positive signal),
   0 otherwise
   + HER (relabel goals in hindsight)
   + curriculum (nearby objects first, distant ones later)
   → there is no score to farm — only "success" as the true objective

9. Summary: reward engineering makes or breaks an RL project ​

  • The reward is the "product requirements document" of RL, and its design cost is chronically underestimated;
  • Once the reward is broken, switching algorithms (PPO → SAC) won't save you — go back and fix the reward first (rule #9 of RL Design Principles: "start with a simple algorithm");
  • The modern trend is upgrading "hand-written rewards" to "preference learning" (the RLHF route), but preference learning has its own overoptimization problem;
  • Related case study: reward hacking and the alignment tax in LLM alignment — see LLM Alignment in Practice: RLHF.

Further Reading ​

References ​

  • Ng, A. Y., Harada, D., & Russell, S. (1999). Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping. ICML. The original potential-based shaping paper.
  • Andrychowicz, M., Wolski, F., Ray, A., et al. (2017). Hindsight Experience Replay. NeurIPS. arXiv:1707.01495
  • Abbeel, P., & Ng, A. Y. (2004). Apprenticeship Learning via Inverse Reinforcement Learning. ICML.
  • Ziebart, B. D., Maas, A., Bagnell, J. A., & Dey, A. K. (2008). Maximum Entropy Inverse Reinforcement Learning. AAAI.
  • Amodei, D., Olah, C., Steinhardt, J., et al. (2016). Concrete Problems in AI Safety. arXiv:1606.06565 (includes the systematic discussion of reward hacking)
  • Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press. Chapter 17 (POMDPs) and the discussion of reward design.