Reward module #44

Closed
opened 2026-08-20 23:28:17 +02:00 by SirStone · 1 comment
Owner

Parent

#37

What to build

Reward computation from game events plus running mean/std normalization.

Raw reward formula:

  • Damage inflicted: +(4p + 2(p-1)) where p = fire power
  • Damage received: -(4p_enemy + 2(p_enemy - 1))
  • Wall hit: -5.0 per tick of impact
  • Wasted shot (bullet missed): -0.1 × p
  • Win: +20.0 / Loss: -10.0

Running normalization: maintain running mean and variance, normalize rewards to approximately zero-mean unit-variance before feeding to SAC. This is critical because SAC's entropy term is O(1) and rewards must match that scale.

Acceptance criteria

  • computeReward() returns correct values for each event type per the formula
  • Running normalization converges toward zero-mean unit-variance over many samples
  • Normalization handles the cold-start case (first few samples) without NaN or division by zero
  • test_rewards.nim passes

Blocked by

  • #38 (SAC_LSTM_Bot: project scaffold)
## Parent #37 ## What to build Reward computation from game events plus running mean/std normalization. Raw reward formula: - Damage inflicted: `+(4p + 2(p-1))` where p = fire power - Damage received: `-(4p_enemy + 2(p_enemy - 1))` - Wall hit: `-5.0` per tick of impact - Wasted shot (bullet missed): `-0.1 × p` - Win: `+20.0` / Loss: `-10.0` Running normalization: maintain running mean and variance, normalize rewards to approximately zero-mean unit-variance before feeding to SAC. This is critical because SAC's entropy term is O(1) and rewards must match that scale. ## Acceptance criteria - [ ] `computeReward()` returns correct values for each event type per the formula - [ ] Running normalization converges toward zero-mean unit-variance over many samples - [ ] Normalization handles the cold-start case (first few samples) without NaN or division by zero - [ ] `test_rewards.nim` passes ## Blocked by - #38 (SAC_LSTM_Bot: project scaffold)
SirStone added the ready-for-agent label 2026-08-20 23:28:17 +02:00
Author
Owner

Reward module complete — all event types, Welford running normalization, cold-start safe. All tests pass.

Reward module complete — all event types, Welford running normalization, cold-start safe. All tests pass.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#44