Network and state vector dimension changes #26

Closed
opened 2026-08-17 18:55:27 +02:00 by SirStone · 1 comment
Owner

Parent

#24 — PRD: Command abstraction layer for PPO action space

What to build

Mechanical dimension changes across the network and state vector modules to support the new 6-dim action space and 44-dim state input.

Network changes (in network.nim):

  • Actor output layer: 5 → 6
  • Learnable logStd: 5 → 6
  • Input layer: 42 → 44
  • Discard all saved weights (shapes change) — delete or move existing weight files

State vector changes (in state_vector.nim):

  • Append 2 new inputs at the end of the state vector:
    • Remaining distance to goto target, normalized by battlefield diagonal → [0, 1]
    • Remaining gun angle to aimTo target, normalized by 180° → [0, 1]
  • buildStateVector needs to accept these two values as parameters (they come from PPO_Bot at runtime)
  • Update dimension constant from 42 to 44

Training changes (in training.nim):

  • Update any hardcoded 5 references in log-prob computation and dLogP_dMean allocation to 6

Acceptance criteria

  • actorForward produces 6-dim output tensor
  • criticForward accepts 44-dim input tensor
  • logStd is 6-dimensional
  • buildStateVector returns 44-dim tensor
  • State dims 42 and 43 contain normalized remaining goto distance and gun angle
  • Old weight files are removed or archived
  • computeLogProb works with 6-dim actions
  • Bot compiles and network initializes without errors

Blocked by

None — can start immediately

## Parent #24 — PRD: Command abstraction layer for PPO action space ## What to build Mechanical dimension changes across the network and state vector modules to support the new 6-dim action space and 44-dim state input. **Network changes** (in `network.nim`): - Actor output layer: 5 → 6 - Learnable `logStd`: 5 → 6 - Input layer: 42 → 44 - Discard all saved weights (shapes change) — delete or move existing weight files **State vector changes** (in `state_vector.nim`): - Append 2 new inputs at the end of the state vector: - Remaining distance to goto target, normalized by battlefield diagonal → [0, 1] - Remaining gun angle to aimTo target, normalized by 180° → [0, 1] - `buildStateVector` needs to accept these two values as parameters (they come from PPO_Bot at runtime) - Update dimension constant from 42 to 44 **Training changes** (in `training.nim`): - Update any hardcoded `5` references in log-prob computation and `dLogP_dMean` allocation to `6` ## Acceptance criteria - [ ] `actorForward` produces 6-dim output tensor - [ ] `criticForward` accepts 44-dim input tensor - [ ] `logStd` is 6-dimensional - [ ] `buildStateVector` returns 44-dim tensor - [ ] State dims 42 and 43 contain normalized remaining goto distance and gun angle - [ ] Old weight files are removed or archived - [ ] `computeLogProb` works with 6-dim actions - [ ] Bot compiles and network initializes without errors ## Blocked by None — can start immediately
SirStone added the ready-for-agent label 2026-08-17 18:55:27 +02:00
Author
Owner

Implemented in commit c713585 on branch research/goto-controller.

Implemented in commit c713585 on branch research/goto-controller.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#26