SirStone
  • Joined on 2025-05-04
SirStone opened issue SirStone/SirRoboGarage#12 2026-08-16 12:56:47 +02:00
Learning rate and optimizer hyperparameters
SirStone closed issue SirStone/SirRoboGarage#10 2026-08-16 12:54:53 +02:00
Exploration noise
SirStone commented on issue SirStone/SirRoboGarage#10 2026-08-16 12:54:48 +02:00
Exploration noise

Resolution

No explicit exploration noise. PPO's stochastic Gaussian policy (learned log-std per action dimension, decided in #9) provides exploration naturally. Entropy bonus in the PPO…

SirStone closed issue SirStone/SirRoboGarage#9 2026-08-16 12:19:24 +02:00
Network architecture
SirStone commented on issue SirStone/SirRoboGarage#9 2026-08-16 12:19:20 +02:00
Network architecture

Resolution

Structure: Separate actor and critic networks (no shared trunk). ~8K parameters total + 5 log-std floats.

Actor:

  • Input: 42 floats (state vector from #4)
  • Hidden: 64 →…
SirStone opened issue SirStone/SirRoboGarage#11 2026-08-16 11:51:33 +02:00
Bot scaffold and build setup
SirStone opened issue SirStone/SirRoboGarage#10 2026-08-16 11:51:25 +02:00
Exploration noise
SirStone opened issue SirStone/SirRoboGarage#9 2026-08-16 11:51:22 +02:00
Network architecture
SirStone closed issue SirStone/SirRoboGarage#8 2026-08-16 09:58:32 +02:00
Training loop design
SirStone commented on issue SirStone/SirRoboGarage#8 2026-08-16 09:58:22 +02:00
Training loop design

Resolution

Episode buffer: Seq of (state: Tensor, action: Tensor, log_prob: float32, reward: float32, value: float32) tuples. Append during round, clear after training. Single episode,…

SirStone closed issue SirStone/SirRoboGarage#6 2026-08-16 09:14:32 +02:00
Reward function design
SirStone commented on issue SirStone/SirRoboGarage#6 2026-08-16 09:14:31 +02:00
Reward function design

Resolution

Per-tick reward: reward_tick = my_energy_delta - enemy_energy_delta

  • Raw values, no scaling or normalization
  • No extra wall penalty (already captured in energy delta)
  • No…
SirStone closed issue SirStone/SirRoboGarage#5 2026-08-15 22:58:45 +02:00
Action space mapping
SirStone commented on issue SirStone/SirRoboGarage#5 2026-08-15 22:58:44 +02:00
Action space mapping

Resolution

5D continuous action space

SirStone closed issue SirStone/SirRoboGarage#4 2026-08-15 22:47:56 +02:00
State vector design
SirStone commented on issue SirStone/SirRoboGarage#4 2026-08-15 22:47:55 +02:00
State vector design

Resolution

State vector: 42 floats, all normalized to ~[-1, 1]

Current tick (14 floats):

SirStone closed issue SirStone/SirRoboGarage#7 2026-08-15 22:31:28 +02:00
Enemy tracker design
SirStone commented on issue SirStone/SirRoboGarage#7 2026-08-15 22:31:23 +02:00
Enemy tracker design

Resolution

Radar strategy (1v1)

  • Decoupled radar — tracker fully owns setRadarTurnRate, compensating for body/gun rotation.
  • Round start: full-speed sweep (45°/tick) until…
SirStone closed issue SirStone/SirRoboGarage#2 2026-08-15 22:17:15 +02:00
Arraymancer viability spike
SirStone commented on issue SirStone/SirRoboGarage#2 2026-08-15 22:17:02 +02:00
Arraymancer viability spike

Resolution: Arraymancer 0.7.33 compiles and trains on Nim 2.2.4 + NixOS. XOR MLP trains in 0.10s, all predictions correct. NixOS requires LD_LIBRARY_PATH for OpenBLAS (standard flake.nix fix).…