Reward function design #6

Closed
opened 2026-08-15 22:05:29 +02:00 by SirStone · 1 comment
Owner

Question

What reward signal drives learning? Per-tick, per-round, or both?

Available signals:

  • Energy delta per tick (damage dealt − damage taken − firing cost − wall/ram penalties)
  • Round outcome (win=+1, lose=−1, or proportional to survival time)
  • Bullet hit confirmation (+reward per hit)
  • Game points (official scoring, but only available end of round)

Shaping options:

  • Per-tick energy delta encourages aggressive play (gain energy by hitting)
  • Per-tick survival bonus encourages passive play (tension with aggression)
  • Damage-dealt-only reward ignores defense
  • Combined: per-tick shaped reward + end-of-round outcome bonus

What balance produces a bot that learns both offense and defense?

## Question What reward signal drives learning? Per-tick, per-round, or both? Available signals: - Energy delta per tick (damage dealt − damage taken − firing cost − wall/ram penalties) - Round outcome (win=+1, lose=−1, or proportional to survival time) - Bullet hit confirmation (+reward per hit) - Game points (official scoring, but only available end of round) Shaping options: - Per-tick energy delta encourages aggressive play (gain energy by hitting) - Per-tick survival bonus encourages passive play (tension with aggression) - Damage-dealt-only reward ignores defense - Combined: per-tick shaped reward + end-of-round outcome bonus What balance produces a bot that learns both offense and defense?
SirStone added the wayfinder:grilling label 2026-08-15 22:05:29 +02:00
SirStone self-assigned this 2026-08-16 08:54:09 +02:00
Author
Owner

Resolution

Per-tick reward:
reward_tick = my_energy_delta - enemy_energy_delta

  • Raw values, no scaling or normalization
  • No extra wall penalty (already captured in energy delta)
  • No special ram signal (energy delta + round score covers it)
  • No aiming/bearing shaping (let the bot discover targeting strategies — leading the target > pointing at current position)

Per-round reward:
reward_round = round_score / 100.0

  • Uses the game's actual scoring system (bullet damage + ram damage + kill bonuses + survival + last survivor), not binary win/lose
  • Aligns the RL objective with what actually determines battle winners (accumulated score across rounds, not round wins)
  • Fixed divisor; tunable later as a hyperparameter

Discount: gamma = 0.99, GAE lambda = 0.95 (standard PPO defaults)

Design philosophy: Minimal shaping. Energy delta is the dense bootstrapping signal; round score keeps the bot honest about what wins battles. Emergent tactics (ramming as a finisher, leading targets) over hand-coded incentives.

## Resolution **Per-tick reward:** `reward_tick = my_energy_delta - enemy_energy_delta` - Raw values, no scaling or normalization - No extra wall penalty (already captured in energy delta) - No special ram signal (energy delta + round score covers it) - No aiming/bearing shaping (let the bot discover targeting strategies — leading the target > pointing at current position) **Per-round reward:** `reward_round = round_score / 100.0` - Uses the game's actual scoring system (bullet damage + ram damage + kill bonuses + survival + last survivor), not binary win/lose - Aligns the RL objective with what actually determines battle winners (accumulated score across rounds, not round wins) - Fixed divisor; tunable later as a hyperparameter **Discount:** gamma = 0.99, GAE lambda = 0.95 (standard PPO defaults) **Design philosophy:** Minimal shaping. Energy delta is the dense bootstrapping signal; round score keeps the bot honest about what wins battles. Emergent tactics (ramming as a finisher, leading targets) over hand-coded incentives.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#6