-
4258d364b9
feat(SAC_LSTM_Bot): project scaffold (#38)
SirStone
2026-08-20 23:33:43 +02:00
-
-
-
f88580b157
feat(radar_lock): standalone reusable module (#39)
SirStone
2026-08-20 23:33:40 +02:00
-
-
ca3e3d2272
tune(PPO_Bot): logStd=-2.0 (std≈0.135), entropy=0, ceiling=-1.0
SirStone
2026-08-20 15:28:15 +02:00
-
0d35646dc9
feat(PPO_Bot): deterministic eval + fix logStd warm-start
SirStone
2026-08-20 15:18:07 +02:00
-
c834d2cbee
fix(PPO_Bot): round_counter always written after increment
SirStone
2026-08-20 15:11:12 +02:00
-
fedab54bc0
feat(PPO_Bot): multi-round transition accumulation (UPDATE_INTERVAL=10)
SirStone
2026-08-20 15:06:00 +02:00
-
82eeb53e5c
tune(training): 20-round chunks, 50 eval rounds, 30k total rounds
SirStone
2026-08-20 14:38:09 +02:00
-
12624d3069
feat(PPO_Bot): bot-relative bullets + scan staleness (STATE_DIM=57)
SirStone
2026-08-20 14:27:28 +02:00
-
6ad51148f4
fix(PPO_Bot): SIGSEGV crash fixes + static buffers for thread safety
SirStone
2026-08-20 14:23:08 +02:00
-
75e32e3315
PPO_Bot Fire campaign: 3787/3789 wins (99.95%); frozen eval 500/500
SirStone
2026-08-19 04:00:47 +02:00
-
db99153f65
cap terminal reward scale (bounded score bonus) and restore single-battle campaigns
SirStone
2026-08-19 03:50:14 +02:00
-
da2f825ad8
fix(training): keep run.sh looping across 60-round battle chunks
SirStone
2026-08-19 03:36:20 +02:00
-
186e005a96
fix(training): cap battles at 60 rounds to bound round-end reward scale
SirStone
2026-08-19 03:32:31 +02:00
-
a4e830531b
fix(botapi): static SVG + intent buffers to kill cross-thread heap realloc
SirStone
2026-08-19 03:32:28 +02:00
-
64697f917e
fix(botapi): static event queue storage + end-of-battle train wait
SirStone
2026-08-19 03:12:22 +02:00
-
766b9e03ee
feat(PPO_Bot): enemy-centered action space + reward shaping for 100% vs Target
SirStone
2026-08-18 17:40:59 +02:00
-
0b17430735
feat(PPO_Bot): persist Adam optimizer state and round counter across restarts (#35)
SirStone
2026-08-18 15:40:54 +02:00
-
cdde60d79f
feat(PPO_Bot): command abstraction layer — goto/aimTo controllers (#24)
SirStone
2026-08-17 19:08:31 +02:00
-
56e0b306c9
docs(research): goto controller algorithm for issue #20
SirStone
2026-08-17 17:04:57 +02:00
-
e5609a7d9b
fix(PPO_Bot): radar lock — use enemy_tracker width-lock, fix arctan2 arg order
SirStone
2026-08-17 11:52:36 +02:00
-
a8ee2a86e3
feat(PPO_Bot): show training progress in game UI
SirStone
2026-08-16 16:48:08 +02:00
-
bbc9e51166
fix(PPO_Bot): state vector bearing uses game coords (north=0° CW)
SirStone
2026-08-16 16:40:28 +02:00
-
fd22535f5b
fix(PPO_Bot): radar lock oscillation bug — arctan2 arg order wrong for Tank Royale coords
SirStone
2026-08-16 16:38:15 +02:00
-
f27b0238f0
feat(PPO_Bot): full PPO RL implementation (#14, #15, #16, #17)
SirStone
2026-08-16 16:34:37 +02:00
-
aea0724d3a
fix(PPO_Bot): radar oscillation, Adam persistence, checkpoint order, channel race
SirStone
2026-08-16 15:42:06 +02:00
-
473d67f644
feat(PPO_Bot): weight persistence + background training (#17)
SirStone
2026-08-16 15:35:42 +02:00
-
eadd177d3b
feat(PPO_Bot): reward + trajectory + GAE + PPO training (#16)
SirStone
2026-08-16 15:27:15 +02:00
-
588c9ebc2f
feat(PPO_Bot): enemy tracker + 42-float state vector (#15)
SirStone
2026-08-16 15:12:13 +02:00
-
aa4bc77068
feat(PPO_Bot): network forward pass + action mapping (#14)
SirStone
2026-08-16 15:10:46 +02:00
-
30cda871cc
research: RL algorithm choice — recommend PPO for Tank Royale bot
SirStone
2026-08-15 22:07:47 +02:00
-
-
64f73dd413
chore: add agent skills configuration
SirStone
2026-08-15 19:26:43 +02:00