Reward + trajectory + GAE + PPO training #16

Closed
opened 2026-08-16 15:02:23 +02:00 by SirStone · 1 comment
Owner

Parent

#13

What to build

Close the learning loop: compute rewards, collect trajectories, calculate advantages, and train the networks via PPO after each round. After this slice, the bot improves across rounds within a single session.

End-to-end slice:

  • Per-tick reward: reward_tick = my_energy_delta - enemy_energy_delta. Raw values, no scaling.
  • Per-round reward: reward_round = round_score / 100.0 added to the last tick's reward.
  • Trajectory buffer: Seq of (state: Tensor, action: Tensor, log_prob: float32, reward: float32, value: float32) tuples. Append each tick during the round. Clear after training.
  • GAE computation: gamma=0.99, lambda=0.95. Terminal value=0 on natural round end (death/win). Bootstrap with critic(last_state) on abnormal interruption.
  • PPO update: 4 epochs, mini-batches of 64 (shuffled each epoch). Clipped surrogate objective (epsilon=0.2). Entropy bonus (coefficient=0.01). Combined loss = actor_loss + 0.5×critic_loss - entropy_bonus.
  • Adam optimizer: lr=3e-4, separate instances for actor and critic. beta1=0.9, beta2=0.999, epsilon=1e-8.
  • Gradient clipping: max global norm 0.5. Hand-roll if Arraymancer doesn't support it natively.
  • Bot integration: on each tick, store transition in buffer. On round end (onRoundEnded), run training pass synchronously (background thread comes in slice 4). Swap weights in place.
  • Assert-based tests: hand-calculated short trajectory → expected GAE advantages; known energy deltas → expected reward.

Acceptance criteria

  • Per-tick reward computed from energy deltas
  • Per-round reward added from game score
  • Trajectory buffer collects full round of transitions
  • GAE produces advantages and returns for the trajectory
  • PPO clipped surrogate loss computed correctly
  • Entropy bonus included in loss
  • Adam updates both actor and critic networks
  • Gradient clipping enforced (max norm 0.5)
  • Bot visibly changes behavior across rounds (not stuck in initial random policy)
  • Assert-based test file passes

Blocked by

  • #14 (Network forward pass → bot moves)
  • #15 (Enemy tracker + full state vector)
## Parent #13 ## What to build Close the learning loop: compute rewards, collect trajectories, calculate advantages, and train the networks via PPO after each round. After this slice, the bot improves across rounds within a single session. End-to-end slice: - **Per-tick reward**: `reward_tick = my_energy_delta - enemy_energy_delta`. Raw values, no scaling. - **Per-round reward**: `reward_round = round_score / 100.0` added to the last tick's reward. - **Trajectory buffer**: Seq of `(state: Tensor, action: Tensor, log_prob: float32, reward: float32, value: float32)` tuples. Append each tick during the round. Clear after training. - **GAE computation**: gamma=0.99, lambda=0.95. Terminal value=0 on natural round end (death/win). Bootstrap with critic(last_state) on abnormal interruption. - **PPO update**: 4 epochs, mini-batches of 64 (shuffled each epoch). Clipped surrogate objective (epsilon=0.2). Entropy bonus (coefficient=0.01). Combined loss = actor_loss + 0.5×critic_loss - entropy_bonus. - **Adam optimizer**: lr=3e-4, separate instances for actor and critic. beta1=0.9, beta2=0.999, epsilon=1e-8. - **Gradient clipping**: max global norm 0.5. Hand-roll if Arraymancer doesn't support it natively. - **Bot integration**: on each tick, store transition in buffer. On round end (`onRoundEnded`), run training pass synchronously (background thread comes in slice 4). Swap weights in place. - **Assert-based tests**: hand-calculated short trajectory → expected GAE advantages; known energy deltas → expected reward. ## Acceptance criteria - [ ] Per-tick reward computed from energy deltas - [ ] Per-round reward added from game score - [ ] Trajectory buffer collects full round of transitions - [ ] GAE produces advantages and returns for the trajectory - [ ] PPO clipped surrogate loss computed correctly - [ ] Entropy bonus included in loss - [ ] Adam updates both actor and critic networks - [ ] Gradient clipping enforced (max norm 0.5) - [ ] Bot visibly changes behavior across rounds (not stuck in initial random policy) - [ ] Assert-based test file passes ## Blocked by - #14 (Network forward pass → bot moves) - #15 (Enemy tracker + full state vector)
SirStone added the ready-for-agent label 2026-08-16 15:02:23 +02:00
Author
Owner

Resolved in f27b023. Per-tick energy delta reward + per-round score/100, trajectory buffer, GAE (γ=0.99, λ=0.95), PPO clipped surrogate (ε=0.2) with entropy bonus (0.01), Adam lr=3e-4, gradient clipping max norm 0.5. Bot improves across rounds.

Resolved in f27b023. Per-tick energy delta reward + per-round score/100, trajectory buffer, GAE (γ=0.99, λ=0.95), PPO clipped surrogate (ε=0.2) with entropy bonus (0.01), Adam lr=3e-4, gradient clipping max norm 0.5. Bot improves across rounds.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#16