Network forward pass → bot moves via neural network #14

Closed
opened 2026-08-16 15:01:49 +02:00 by SirStone · 1 comment
Owner

Parent

#13

What to build

Wire a working neural network into PPO_Bot so the bot moves, turns, and fires based on network output — even though weights are random and state is minimal.

End-to-end slice:

  • Actor network: 42→64→tanh→64→tanh→5 linear outputs. Separate learnable log-std parameter vector (5 floats, initialized to 0). Floor log-std at -3.0.
  • Critic network: 42→64→tanh→64→tanh→1 linear output (unbounded).
  • Action mapping: network output → bot commands. targetSpeed (tanh×8), turnRate (tanh × speed-dependent max, where max = 10 - 0.75×|speed|), gunTurnRate (tanh×20), fireDecision (tanh, threshold at 0), firePower (sigmoid×2.9+0.1). Suppress fire when gun heat > 0.
  • Stochastic policy: sample actions from Normal(mean=network_output, std=exp(log_std)). Return log_prob alongside actions for later PPO use.
  • Minimal bot glue: in the run loop, construct a dummy/hardcoded 42-float state tensor (zeros or constants — real state comes in slice 2), forward through actor, apply actions via setTargetSpeed, setTurnRate, setGunTurnRate, setFire/setFirePower.
  • Assert-based tests in a runnable test file: network forward produces correct output shapes, action values fall in valid ranges, gun heat masking works, log-std floor enforced.

All RL logic must operate on plain Tensor[float32] — no bot API types cross the seam. The bot glue is thin assignment statements only.

Acceptance criteria

  • Actor network forward pass: 42-float input → 5-float output
  • Critic network forward pass: 42-float input → 1-float output
  • Action mapping produces values in valid ranges (speed [-8,8], turnRate within speed-dependent bounds, gunTurnRate [-20,20], firePower [0.1,3.0])
  • Fire suppressed when gun heat > 0
  • Stochastic sampling works and returns log_prob
  • Log-std floored at -3.0
  • Bot compiles, runs, and visibly moves/turns/fires in a Tank Royale match
  • Assert-based test file passes

Blocked by

None — can start immediately

## Parent #13 ## What to build Wire a working neural network into PPO_Bot so the bot moves, turns, and fires based on network output — even though weights are random and state is minimal. End-to-end slice: - **Actor network**: 42→64→tanh→64→tanh→5 linear outputs. Separate learnable log-std parameter vector (5 floats, initialized to 0). Floor log-std at -3.0. - **Critic network**: 42→64→tanh→64→tanh→1 linear output (unbounded). - **Action mapping**: network output → bot commands. targetSpeed (tanh×8), turnRate (tanh × speed-dependent max, where max = 10 - 0.75×|speed|), gunTurnRate (tanh×20), fireDecision (tanh, threshold at 0), firePower (sigmoid×2.9+0.1). Suppress fire when gun heat > 0. - **Stochastic policy**: sample actions from Normal(mean=network_output, std=exp(log_std)). Return log_prob alongside actions for later PPO use. - **Minimal bot glue**: in the `run` loop, construct a dummy/hardcoded 42-float state tensor (zeros or constants — real state comes in slice 2), forward through actor, apply actions via `setTargetSpeed`, `setTurnRate`, `setGunTurnRate`, `setFire`/`setFirePower`. - **Assert-based tests** in a runnable test file: network forward produces correct output shapes, action values fall in valid ranges, gun heat masking works, log-std floor enforced. All RL logic must operate on plain Tensor[float32] — no bot API types cross the seam. The bot glue is thin assignment statements only. ## Acceptance criteria - [ ] Actor network forward pass: 42-float input → 5-float output - [ ] Critic network forward pass: 42-float input → 1-float output - [ ] Action mapping produces values in valid ranges (speed [-8,8], turnRate within speed-dependent bounds, gunTurnRate [-20,20], firePower [0.1,3.0]) - [ ] Fire suppressed when gun heat > 0 - [ ] Stochastic sampling works and returns log_prob - [ ] Log-std floored at -3.0 - [ ] Bot compiles, runs, and visibly moves/turns/fires in a Tank Royale match - [ ] Assert-based test file passes ## Blocked by None — can start immediately
SirStone added the ready-for-agent label 2026-08-16 15:01:49 +02:00
Author
Owner

Resolved in f27b023. Actor (42→64→64→5) and critic (42→64→64→1) forward pass, stochastic policy with learnable log-std (floored at -3.0), action mapping with speed-dependent turn scaling and gun heat masking. Tests in test_state.

Resolved in f27b023. Actor (42→64→64→5) and critic (42→64→64→1) forward pass, stochastic policy with learnable log-std (floored at -3.0), action mapping with speed-dependent turn scaling and gun heat masking. Tests in test_state.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#14