PPO_Bot wiring and integration test #28

Closed
opened 2026-08-17 18:55:52 +02:00 by SirStone · 1 comment
Owner

Parent

#24 — PRD: Command abstraction layer for PPO action space

What to build

Wire everything together in PPO_Bot.nim and verify end-to-end with a battle runner integration test.

PPO_Bot changes:

  • Track current goto target (x, y) and aimTo target (x, y) as bot state between ticks
  • Each tick: decode actions via new mapActions, store the goto/aimTo targets
  • Compute remaining goto distance: hypot(targetX - botX, targetY - botY)
  • Compute remaining gun angle: abs(normalizeRelativeAngle(directionTo(aimX, aimY) - gunDirection))
  • Pass both to buildStateVector for next tick's state
  • Apply motor commands from controllers: setTargetSpeed, setTurnRate, setGunTurnRate
  • Apply fire commands as before
  • Initialize goto/aimTo targets to bot's own position on round start (remaining distance = 0)
  • Ensure setAdjustGunForBodyTurn(true) is set so gun is independent of body

Integration test:

  • Build PPO_Bot with new architecture
  • Run via battle runner (tools/battle_runner/run.sh) against Target bot
  • Verify: bot doesn't crash, moves, fires, completes a round
  • This is a smoke test — training quality is out of scope

Acceptance criteria

  • PPO_Bot compiles with all changes from slices 1-3
  • Bot connects to Tank Royale and completes a round without crashing
  • Bot produces movement (non-zero speed observed in battle runner output)
  • Bot fires at least once during the round
  • Remaining goto distance and gun angle are correctly computed and fed to state vector
  • Fresh weights are initialized (no old weight files loaded)
  • Battle runner integration test passes

Blocked by

  • #25 — Goto and AimTo controller functions
  • #26 — Network and state vector dimension changes
  • #27 — Action decoding rewrite for command abstraction
## Parent #24 — PRD: Command abstraction layer for PPO action space ## What to build Wire everything together in `PPO_Bot.nim` and verify end-to-end with a battle runner integration test. **PPO_Bot changes:** - Track current goto target (x, y) and aimTo target (x, y) as bot state between ticks - Each tick: decode actions via new `mapActions`, store the goto/aimTo targets - Compute remaining goto distance: `hypot(targetX - botX, targetY - botY)` - Compute remaining gun angle: `abs(normalizeRelativeAngle(directionTo(aimX, aimY) - gunDirection))` - Pass both to `buildStateVector` for next tick's state - Apply motor commands from controllers: `setTargetSpeed`, `setTurnRate`, `setGunTurnRate` - Apply fire commands as before - Initialize goto/aimTo targets to bot's own position on round start (remaining distance = 0) - Ensure `setAdjustGunForBodyTurn(true)` is set so gun is independent of body **Integration test:** - Build PPO_Bot with new architecture - Run via battle runner (`tools/battle_runner/run.sh`) against Target bot - Verify: bot doesn't crash, moves, fires, completes a round - This is a smoke test — training quality is out of scope ## Acceptance criteria - [ ] PPO_Bot compiles with all changes from slices 1-3 - [ ] Bot connects to Tank Royale and completes a round without crashing - [ ] Bot produces movement (non-zero speed observed in battle runner output) - [ ] Bot fires at least once during the round - [ ] Remaining goto distance and gun angle are correctly computed and fed to state vector - [ ] Fresh weights are initialized (no old weight files loaded) - [ ] Battle runner integration test passes ## Blocked by - #25 — Goto and AimTo controller functions - #26 — Network and state vector dimension changes - #27 — Action decoding rewrite for command abstraction
SirStone added the ready-for-agent label 2026-08-17 18:55:52 +02:00
Author
Owner

Resolution

Audited the full wiring. Previous tickets (#25–#27) had already done the heavy lifting:

  • #25: gotoTick / aimToTick in controllers.nim — done
  • #26: State vector expanded to 44 dims (indices 42–43 = remaining distance / gun angle) — done
  • #27: mapActions decodes 6-dim output → goto/aimTo targets → controller calls — done
  • PPO_Bot.nim tick loop: remainingGotoDistance / remainingGunAngle computed from bot.lastActions and fed into buildStateVector — done
  • setAdjustGunForBodyTurn(true) already present in onRoundStarted — done

One gap found and fixed: bot.lastActions.gotoX/Y and aimToX/Y were uninitialized (zero) on round start, causing a spurious ~1200-unit "remaining distance" to (0,0) on tick 1. Fixed by seeding them to the bot's actual position at the top of run(), matching the same pattern used to seed prevEnergy. remainingGotoDistance and remainingGunAngle are now correctly 0 on the first tick.

Compiles cleanly. Closing.

## Resolution Audited the full wiring. Previous tickets (#25–#27) had already done the heavy lifting: - **#25**: `gotoTick` / `aimToTick` in `controllers.nim` — done - **#26**: State vector expanded to 44 dims (indices 42–43 = remaining distance / gun angle) — done - **#27**: `mapActions` decodes 6-dim output → goto/aimTo targets → controller calls — done - **`PPO_Bot.nim` tick loop**: `remainingGotoDistance` / `remainingGunAngle` computed from `bot.lastActions` and fed into `buildStateVector` — done - **`setAdjustGunForBodyTurn(true)`** already present in `onRoundStarted` — done **One gap found and fixed**: `bot.lastActions.gotoX/Y` and `aimToX/Y` were uninitialized (zero) on round start, causing a spurious ~1200-unit "remaining distance" to (0,0) on tick 1. Fixed by seeding them to the bot's actual position at the top of `run()`, matching the same pattern used to seed `prevEnergy`. `remainingGotoDistance` and `remainingGunAngle` are now correctly 0 on the first tick. Compiles cleanly. Closing.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#28