Commit Graph

  • 4258d364b9 feat(SAC_LSTM_Bot): project scaffold (#38) SirStone 2026-08-20 23:33:43 +02:00
  • f88580b157 feat(radar_lock): standalone reusable module (#39) SirStone 2026-08-20 23:33:40 +02:00
  • ca3e3d2272 tune(PPO_Bot): logStd=-2.0 (std≈0.135), entropy=0, ceiling=-1.0 SirStone 2026-08-20 15:28:15 +02:00
  • 0d35646dc9 feat(PPO_Bot): deterministic eval + fix logStd warm-start SirStone 2026-08-20 15:18:07 +02:00
  • c834d2cbee fix(PPO_Bot): round_counter always written after increment SirStone 2026-08-20 15:11:12 +02:00
  • fedab54bc0 feat(PPO_Bot): multi-round transition accumulation (UPDATE_INTERVAL=10) SirStone 2026-08-20 15:06:00 +02:00
  • 82eeb53e5c tune(training): 20-round chunks, 50 eval rounds, 30k total rounds SirStone 2026-08-20 14:38:09 +02:00
  • 12624d3069 feat(PPO_Bot): bot-relative bullets + scan staleness (STATE_DIM=57) SirStone 2026-08-20 14:27:28 +02:00
  • 6ad51148f4 fix(PPO_Bot): SIGSEGV crash fixes + static buffers for thread safety SirStone 2026-08-20 14:23:08 +02:00
  • 75e32e3315 PPO_Bot Fire campaign: 3787/3789 wins (99.95%); frozen eval 500/500 SirStone 2026-08-19 04:00:47 +02:00
  • db99153f65 cap terminal reward scale (bounded score bonus) and restore single-battle campaigns SirStone 2026-08-19 03:50:14 +02:00
  • da2f825ad8 fix(training): keep run.sh looping across 60-round battle chunks SirStone 2026-08-19 03:36:20 +02:00
  • 186e005a96 fix(training): cap battles at 60 rounds to bound round-end reward scale SirStone 2026-08-19 03:32:31 +02:00
  • a4e830531b fix(botapi): static SVG + intent buffers to kill cross-thread heap realloc SirStone 2026-08-19 03:32:28 +02:00
  • 64697f917e fix(botapi): static event queue storage + end-of-battle train wait SirStone 2026-08-19 03:12:22 +02:00
  • 766b9e03ee feat(PPO_Bot): enemy-centered action space + reward shaping for 100% vs Target SirStone 2026-08-18 17:40:59 +02:00
  • 0b17430735 feat(PPO_Bot): persist Adam optimizer state and round counter across restarts (#35) SirStone 2026-08-18 15:40:54 +02:00
  • cdde60d79f feat(PPO_Bot): command abstraction layer — goto/aimTo controllers (#24) SirStone 2026-08-17 19:08:31 +02:00
  • 56e0b306c9 docs(research): goto controller algorithm for issue #20 SirStone 2026-08-17 17:04:57 +02:00
  • e5609a7d9b fix(PPO_Bot): radar lock — use enemy_tracker width-lock, fix arctan2 arg order SirStone 2026-08-17 11:52:36 +02:00
  • a8ee2a86e3 feat(PPO_Bot): show training progress in game UI SirStone 2026-08-16 16:48:08 +02:00
  • bbc9e51166 fix(PPO_Bot): state vector bearing uses game coords (north=0° CW) SirStone 2026-08-16 16:40:28 +02:00
  • fd22535f5b fix(PPO_Bot): radar lock oscillation bug — arctan2 arg order wrong for Tank Royale coords SirStone 2026-08-16 16:38:15 +02:00
  • f27b0238f0 feat(PPO_Bot): full PPO RL implementation (#14, #15, #16, #17) SirStone 2026-08-16 16:34:37 +02:00
  • aea0724d3a fix(PPO_Bot): radar oscillation, Adam persistence, checkpoint order, channel race SirStone 2026-08-16 15:42:06 +02:00
  • 473d67f644 feat(PPO_Bot): weight persistence + background training (#17) SirStone 2026-08-16 15:35:42 +02:00
  • eadd177d3b feat(PPO_Bot): reward + trajectory + GAE + PPO training (#16) SirStone 2026-08-16 15:27:15 +02:00
  • 588c9ebc2f feat(PPO_Bot): enemy tracker + 42-float state vector (#15) SirStone 2026-08-16 15:12:13 +02:00
  • aa4bc77068 feat(PPO_Bot): network forward pass + action mapping (#14) SirStone 2026-08-16 15:10:46 +02:00
  • 30cda871cc research: RL algorithm choice — recommend PPO for Tank Royale bot SirStone 2026-08-15 22:07:47 +02:00
  • 64f73dd413 chore: add agent skills configuration SirStone 2026-08-15 19:26:43 +02:00