Reward + trajectory + GAE + PPO training #16
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Parent
#13
What to build
Close the learning loop: compute rewards, collect trajectories, calculate advantages, and train the networks via PPO after each round. After this slice, the bot improves across rounds within a single session.
End-to-end slice:
reward_tick = my_energy_delta - enemy_energy_delta. Raw values, no scaling.reward_round = round_score / 100.0added to the last tick's reward.(state: Tensor, action: Tensor, log_prob: float32, reward: float32, value: float32)tuples. Append each tick during the round. Clear after training.onRoundEnded), run training pass synchronously (background thread comes in slice 4). Swap weights in place.Acceptance criteria
Blocked by
Resolved in
f27b023. Per-tick energy delta reward + per-round score/100, trajectory buffer, GAE (γ=0.99, λ=0.95), PPO clipped surrogate (ε=0.2) with entropy bonus (0.01), Adam lr=3e-4, gradient clipping max norm 0.5. Bot improves across rounds.