diff --git a/.gitignore b/.gitignore index 50723d9..6005ba6 100644 --- a/.gitignore +++ b/.gitignore @@ -22,5 +22,16 @@ __pycache__/ devbox.json devbox.lock +# Compiled bot binaries (no extension, bot-specific directories) +OscillatorBot_garage/OscillatorBot +SAC_LSTM_Bot_garage/SAC_LSTM_Bot + +# Java compiled classes +*.class + +# Log and snapshot directories +**/logs/ +**/snapshots/ + # Git worktrees worktrees/ diff --git a/SAC_LSTM_Bot_garage/docs/bozza_requisiti.md b/SAC_LSTM_Bot_garage/docs/bozza_requisiti.md deleted file mode 100644 index a91d3c4..0000000 --- a/SAC_LSTM_Bot_garage/docs/bozza_requisiti.md +++ /dev/null @@ -1,98 +0,0 @@ -# Implementazione Agente DRL Recurrent-SAC (SAC-GRU) Nativo per Robocode TankRoyale - - Bot nativo basato sull'algoritmo **Recurrent Soft Actor-Critic (SAC-GRU)** addestrato per combattere nell'ambiente **Robocode TankRoyale**. - -Il bot deve connettersi direttamente al server WebSocket di TankRoyale tramite un'interfaccia client API nativa, senza l'uso di middleware o socket intermediari extra. Il sistema deve garantire il rispetto del limite rigido di **30 ms per turno** imposto dal simulatore, eseguendo l'inferenza e l'addestramento in modo concorrente all'interno dello stesso processo. - ---- - -## 1. Architettura Multi-Thread e Concorrenza (30 ms Tick Constraint) - -Per evitare il fenomeno degli "skipped turns", l'applicazione deve essere suddivisa in due thread principali coordinati nello stesso processo: - -1. **Thread 1: Realtime Async Event Loop (WebSocket Client)** - * Riceve i messaggi di evento dal server TankRoyale via WebSocket ad ogni tick. - * Estragga e normalizza il vettore di stato a 17 dimensioni ($S_t$). - * Esegue l'inferenza rapida dell'Actor ($< 2 \text{ ms}$) fornendo lo stato $S_t$ e lo stato nascosto corrente della GRU ($h_{t-1}$). - * Mappa le azioni restituite e invia immediatamente il `BotIntent` al server WebSocket. - * Invia la transizione $(S_t, A_t, R_t, S_{t+1}, D_t)$ a una coda/canale thread-safe (`Thread-safe Channel`). - -2. **Thread 2: Background Trainer Thread (SAC-GRU Engine)** - * Preleva le transizioni dal canale e le accumula in un **Sequential Replay Buffer**. - * Quando il buffer contiene un numero sufficiente di esperienze, estrae mini-batch di sequenze temporali. - * Esegue l'addestramento in background (Forward/Backward Pass dell'Actor, dei Dual Critic e dell'Alpha autotuning) sfruttando l'accelerazione hardware (GPU/CPU). - * Aggiorna periodicamente i pesi della rete Actor usata dal Thread 1 in modo thread-safe (es. scambio atomico di puntatori o mutua esclusione leggera). - ---- - -## 2. Specifiche dell'Ambiente DRL (POMDP) - -### 2.1 Vettore di Stato ($S \in \mathbb{R}^{17}$, Normalizzato in $[-1, 1]$) - -* `s[0]`: Posizione X propria ($X / \text{width}$) -* `s[1]`: Posizione Y propria ($Y / \text{height}$) -* `s[2]`: Orientamento scafo ($[-\pi, \pi] / \pi$) -* `s[3]`: Velocità lineare propria ($[-8, 8] / 8$) -* `s[4]`: Orientamento cannone ($[-\pi, \pi] / \pi$) -* `s[5]`: Orientamento radar ($[-\pi, \pi] / \pi$) -* `s[6]`: Temperatura del cannone ($[0, 3] / 3$) -* `s[7]`: Energia propria ($[0, 100] / 100$) -* `s[8]`: Distanza dal muro NORD ($(\text{height} - Y) / \text{height}$) -* `s[9]`: Distanza dal muro SUD ($Y / \text{height}$) -* `s[10]`: Distanza dal muro EST ($(\text{width} - X) / \text{width}$) -* `s[11]`: Distanza dal muro OVEST ($X / \text{width}$) -* `s[12]`: Ultima distanza rilevata del nemico ($[0, \text{max\_dist}] / \text{max\_dist}$) -* `s[13]`: Angolo relativo (bearing) del nemico ($[-\pi, \pi] / \pi$) -* `s[14]`: Orientamento del nemico ($[-\pi, \pi] / \pi$) -* `s[15]`: Velocità del nemico ($[-8, 8] / 8$) -* `s[16]`: Energia residua del nemico ($[0, 100] / 100$) - -### 2.2 Vettore delle Azioni Continuo ($A \in \mathbb{R}^4$, Output $[-1, 1]$) - -* `a[0]`: Rotazione scafo $\rightarrow$ Mappato su $[-10^\circ, +10^\circ]$ per tick. -* `a[1]`: Traslazione $\rightarrow$ Mappato su $[-8, +8]$ px/tick. -* `a[2]`: Rotazione cannone $\rightarrow$ Mappato su $[-20^\circ, +20^\circ]$ per tick. -* `a[3]`: Potenza di sparo $\rightarrow$ $\text{ReLU}(a_3) \times 3.0$ (Spara solo se $> 0.1$). - -### 2.3 Reward Function - -$$R_t = R_{\text{danno\_inflitto}} - R_{\text{danno\_subito}} + R_{\text{vittoria/sconfitta}} - R_{\text{muri}} - R_{\text{sparo\_vuoto}}$$ - -* Danno inflitto: $+ (4p + 2(p - 1))$ con $p \le 3$. -* Danno subito: $- (4p_{\text{nemico}} + 2(p_{\text{nemico}} - 1))$. -* Muri: $-5.0$ per tick di impatto. -* Sparo a vuoto: $-0.1 \times p$. -* Vittoria/Sconfitta: $+20.0$ / $-10.0$. - ---- - -## 3. Modelli Neurali e Meccanismi di Addestramento - -Tutte le reti devono integrare una cella ricorsiva **GRU (Gated Recurrent Unit)** per gestire la parziale osservabilità dell'arena (radar in rotazione): - -* **Recurrent Actor Network $\pi_\phi(a_t | s_{:t}, h_{t-1})$:** - * Feature Extractor: Linear/Dense ($17 \rightarrow 128$) + Attivazione - * Memory: GRU Layer (Input: $128$, Hidden: $128$) - * Heads: Linear/Dense ($128 \rightarrow 64$) $\rightarrow$ Outputs: $\mu \in \mathbb{R}^4$ e $\log \sigma \in \mathbb{R}^4$ (Clamped $[-20, 2]$) - * Sampling: Reparameterization Trick con squashing $\tanh$. - -* **Recurrent Dual-Critic Networks $Q_{\theta_1, \theta_2}(s_{:t}, a_{:t}, h_{t-1})$:** - * Fusion: Concatenazione di Stato e Azione ($17 + 4 = 21$) - * Feature Extractor: Linear/Dense ($21 \rightarrow 128$) + Attivazione - * Memory: GRU Layer (Input: $128$, Hidden: $128$) - * Q-Head: Linear/Dense ($128 \rightarrow 64$) $\rightarrow$ Output Q-Value scalare. - -* **Sequential Replay Buffer & Burn-in Strategy:** - * Campionamento di sequenze temporali contigue di lunghezza $L = 24$. - * **Burn-in ($L_{\text{burn}} = 8$ step):** I primi 8 step vengono usati unicamente per aggiornare lo stato nascosto $h$ della GRU, senza calcolo di loss o backpropagation. - * **Training ($L_{\text{train}} = 16$ step):** I successivi 16 step calcolano le loss dell'Actor, dei Critic e l'autotuning del parametro di temperatura $\alpha$ (con target $H_{\text{target}} = -4.0$). - ---- - -## 4. Requisiti di Strutturazione del Codice - -Fornisci il codice sorgente completo, modulare e pronto all'uso articolato nelle seguenti componenti: - -1. **Model Definitions:** Classi/Strutture per Actor, Critic e Target Networks basate su GRU. -2. **Sequential Replay Buffer:** Struttura dati thread-safe con supporto per estrazione di sequenze e gestione Burn-in. -3. **SAC-GRU Trainer Engine:** Algoritmo di aggiornamento, calcolo loss, gradient clipping ($1.0$) e soft update ($\tau = 0.005$). diff --git a/tools/training_runner/CURRICULUM_STATE.md b/tools/training_runner/CURRICULUM_STATE.md deleted file mode 100644 index 0f2b9a6..0000000 --- a/tools/training_runner/CURRICULUM_STATE.md +++ /dev/null @@ -1,187 +0,0 @@ -# PPO Curriculum — Running State File -Chunk 1 of larger campaign. Source of truth for resuming. -Bots (easy->hard order): Fire, MyFirstDroid, Target, MyFirstLeader, Crazy, MyFirstBot, PaintingBot, Corners. -Protocol: BLOCK (train to N+200) -> GATE (frozen eval to N+350, PASS >=148/150) -> snapshot best__r. -Maintenance from 3rd mastered bot on: frozen eval 100 rounds per mastered bot; bar >=97/100; else replay. -GATE bar: >=148/150. MAINT bar: >=97/100. Max 3 attempts per bot. - -## Baseline (verified pre-chunk) -- round_counter.txt = 10905 -- Weights proven (frozen eval): Fire/Target/Crazy/MyFirstDroid/MyFirstLeader 100%, - MyFirstBot/PaintingBot 99%, Corners 93%, VelocityBot 10%, TrackFire/RamFire/SpinBot/Walls 0%, - MyFirstTeam 0% (1v5, flagged). -- Bot launchers chmod-fixed. Logs rotate per phase. Snapshots = cp -a of weights/latest. - ---- - -## Bot 1: Fire -- block1: 2026-08-19 16:41, counter 10905 -> 11105 (200 train rounds, 200 train lines, 199W/1L, 0 NaN, no crash) - log: logs/train_Fire_block1_20260819_164245.jsonl -- gate1: 2026-08-19 16:42, eval to 11255 (150 rounds, 0 train lines, 149W) -> PASS >=148 - log: logs/eval_Fire_gate_20260819_164311.jsonl -- SNAPSHOT best_Fire_r11255: MISSING (operator error: snapshots/ dir absent; cp failed, training overwrote weights before retry). RECOVERY below. -- block2 (REDO): 2026-08-19 16:46, counter 11605 -> 11805 (200 train, 200W, 0 NaN). Log: logs/train_Fire_block2_20260819_164631.jsonl -- gate2 (REDO): 2026-08-19 16:47, eval to 11955 (150 rounds, 0 train lines, 150W) -> PASS. Log: logs/eval_Fire_gate2_20260819_164740.jsonl -- SNAPSHOT: snapshots/best_Fire_r11955 (40 files) -- STATUS: PASS (150/150) -- maint (post-Target): 100/100 PASS. Log: logs/eval_Fire_maint_20260819_164849.jsonl - -## Bot 2: MyFirstDroid -- block1: 2026-08-19 16:43, counter 11255 -> 11455 (200 train rounds, 200 train lines, 200W, 0 NaN) - ANOMALY: operator Ctrl-C at round ~127 triggered run.sh crash-restart (">>> Restart #2, 73 rounds remaining"); round sequence in log continuous (11256..11455, no gaps), slice valid. Log: logs/train_MyFirstDroid_block1_20260819_164349.jsonl -- gate1: 2026-08-19 16:44, eval to 11605 (150 rounds, 0 train lines, 150W) -> PASS. Log: logs/eval_MyFirstDroid_gate_20260819_164506.jsonl -- SNAPSHOT: snapshots/best_MyFirstDroid_r11605 (40 files) -- STATUS: PASS (150/150) -- maint (post-Target): 100/100 PASS. Log: logs/eval_MyFirstDroid_maint_20260819_164919.jsonl - -## Bot 3: Target -- block1: 2026-08-19 16:48, counter 11955 -> 12155 (200 train, 200W, 0 NaN). Log: logs/train_Target_block1_20260819_164823.jsonl -- gate1: 2026-08-19 16:49, eval to 12305 (150 rounds, 0 train lines, 150W) -> PASS. Log: logs/eval_Target_gate_20260819_164928.jsonl -- SNAPSHOT: snapshots/best_Target_r12305 (40 files) -- STATUS: PASS (150/150) -- maint (post-self): 100/100 PASS. Log: logs/eval_Target_maint_20260819_165033.jsonl -- MAINTENANCE ROUND 1 (post-Target): Fire 100/100, MyFirstDroid 100/100, Target 100/100. NO regressions. No replays. - -## Bot 4: MyFirstLeader -- block1: 2026-08-19 17:06, counter 12605 -> 12805 (200 train, 200W, 0 NaN). Log: logs/train_MyFirstLeader_block1_20260819_170632.jsonl -- gate1: 2026-08-19 17:07, eval to 12955 (150 rounds, 0 train, 150W) -> PASS. Log: logs/eval_MyFirstLeader_gate_20260819_170720.jsonl -- SNAPSHOT: snapshots/best_MyFirstLeader_r12955 (40 files) -- STATUS: PASS (150/150) -- maint r2 (post-self): 100/100 PASS. Log: logs/eval_MyFirstLeader_maint_20260819_171104.jsonl -- maint r3 (post-Crazy): 100/100 PASS. Log: logs/eval_MyFirstLeader_maint_20260819_171246.jsonl -- MAINTENANCE ROUND 2 (post-Leader): Fire 100, MyFirstDroid 100, Target 100, MyFirstLeader 100 -> NO regressions, no replays. - -## Bot 5: Crazy -- block1: 2026-08-19 17:12, counter 13355 -> 13555 (200 train, 200W, 0 NaN). Log: logs/train_Crazy_block1_20260819_171246.jsonl -- gate1: 2026-08-19 17:13, eval to 13705 (150 rounds, 0 train, 150W) -> PASS. Log: logs/eval_Crazy_gate_20260819_171346.jsonl -- SNAPSHOT: snapshots/best_Crazy_r13705 (40 files) -- STATUS: PASS (150/150) -- maint r3 (post-self): 99/100 PASS. Log: logs/eval_Crazy_maint_20260819_171623.jsonl -- MAINTENANCE ROUND 3 (post-Crazy): Fire 99, MyFirstDroid 100, Target 100, MyFirstLeader 100, Crazy 99 -> NO regressions, no replays. - ANOMALY (logs): mid-run log rename at 17:00:44 (before Droid maint eval finished) split Droid/Target round lines across files; win evidence intact - (Droid 100 game lines win, Target 100 game lines win). Logs: eval_MyFirstDroid_maint_20260819_170044.jsonl (85 round+100 game), - eval_Target_maint_20260819_170055.jsonl (100 game). No crash; counter continuity verified (13805->13905->14005). - -## Bot 6: MyFirstBot -- block1: 2026-08-19 17:18, counter 14205 -> 14405 (200 train, 0W/200L, 0 NaN) — REGRESSION vs baseline 99%; policy loses every round. - Log: logs/train_MyFirstBot_block1_20260819_171827.jsonl -- gate1: eval to 14555: 0/150 FAIL. Log: logs/eval_MyFirstBot_gate1_20260819_171920.jsonl -- block2: 2026-08-19 17:22, counter 14555 -> 14955 (400 train, 0W/400L). Log: logs/train_MyFirstBot_block2_20260819_172215.jsonl -- gate2: eval to 15105: 0/150 FAIL. Log: logs/eval_MyFirstBot_gate2_20260819_172327.jsonl -- block3: 2026-08-19 17:24, counter 15105 -> 15505 (400 train, 0W/400L). Log: logs/train_MyFirstBot_block3_20260819_172448.jsonl -- gate3: eval to 15655: 0/150 FAIL. Log: logs/eval_MyFirstBot_gate3_20260819_172625.jsonl -- STATUS: OPEN (3 attempts, 1000 train + 450 gate rounds, 0 wins total; eval mechanicals verified: 0 train lines in gates, counter advances, no NaN). Likely policy basin shift (aggressive brawler vs stationary out-damager). Snapshot anyway: snapshots/best_MyFirstBot_r15655. -- Chunk 2/3 to revisit (possibly via PaintingBot/Corners transfer). - -## Bot 7: PaintingBot -- CATASTROPHE (previous session overshoot): run.sh ran unbounded from 15655 to 31694 (16039 rounds) - vs PaintingBot; 3/16027 wins (0%). Policy collapsed. Log: logs/train_PaintingBot_catastrophe_20260819_191718.jsonl - (0 NaN, no coredump, but weights destroyed by repeated loss signal). -- RECOVERY: 2026-08-19 19:17 — restored weights from snapshots/best_Crazy_r13705; round_counter stays 31694. - ROOT CAUSE: Crazy-snapshot weights cannot beat PaintingBot (chunk training caused catastrophic forgetting; - baseline 99% was from pre-chunk Fire-campaign weights). best_Fire_r11955 (closest to baseline) needed. -- block1: 2026-08-19 19:18, counter 31694 -> 31894 (200 train rounds from Crazy restore; 0/200 wins) - log: logs/train_PaintingBot_block1_20260819_192027.jsonl -- gate1: 2026-08-19 19:21, eval to 32044 (150 rounds, 0 train lines, 0/150) -> FAIL - log: logs/eval_PaintingBot_gate1_20260819_192129.jsonl -- RESTORE: best_Fire_r11955 (pre-chunk-adjacent, known 99% vs PaintingBot). Counter stays 32044. -- block2: 2026-08-19 19:23, counter 32044 -> 32444 (400 train rounds, 400/400 wins, 0 NaN) - log: logs/train_PaintingBot_block2_20260819_192325.jsonl -- gate2: 2026-08-19 19:24, eval to 32594 (150 rounds, 0 train lines, 146/150) -> FAIL (bar >=148) - log: logs/eval_PaintingBot_gate2_20260819_192401.jsonl -- block3: 2026-08-19 19:24, counter 32594 -> 32994 (400 train rounds, 400/400 wins, 0 NaN) - log: logs/train_PaintingBot_block3_20260819_192507.jsonl -- gate3: 2026-08-19 19:26, eval to 33144 (150 rounds, 0 train lines, 147/150) -> FAIL (bar >=148) - log: logs/eval_PaintingBot_gate3_20260819_192630.jsonl -- STATUS: OPEN (3 attempts: block1 0/200 from Crazy, block2 400/400 from Fire, block3 400/400 from Fire; - gate results 0/150, 146/150, 147/150 — high variance, bar missed by 1-4 in gates 2+3). - Likely issue: gate runs single JVM with Corners-like random-state dependency (initial corner selection for - PaintingBot = random; policy may be sensitive to specific PaintingBot seed/state at eval time). - Snapshot: snapshots/best_PaintingBot_r33144. -- MAINTENANCE (post-OPEN, 2026-08-19 19:27): Fire 100, MyFirstDroid 100, Target 100, MyFirstLeader 100, - Crazy 99 — NO regressions. No replays. - Logs: eval_Fire_maint_20260819_192658.jsonl, eval_MyFirstDroid_maint_20260819_XXXXXX.jsonl, etc. - -## Bot 8: Corners -- NOTE: Corners bot switches corner on death when losing. Single-JVM gate has fixed initial corner; - policy must handle all 4 corner positions. Training (each JVM = new random corner) cycles all 4. -- block1: 2026-08-19 19:29, counter 33644 -> 33844 (200 train rounds from PaintingBot weights; 78/200 wins, - declining trend: 1/50 in last 50). Training degraded policy. - log: logs/train_Corners_block1_20260819_192947.jsonl -- gate1: 2026-08-19 19:30, eval to 33994 (150 rounds, 0 train, 0/150) -> FAIL - log: logs/eval_Corners_gate1_20260819_193050.jsonl -- RESTORE: best_Fire_r11955 (same as PaintingBot recovery). Counter stays 33994. -- block2: 2026-08-19 19:32, counter 33994 -> 34394 (400 train rounds from Fire restore; 231/400 wins, - improving: 0/50 early, 50/50 last 50; 0 NaN). Policy converges during training. - log: logs/train_Corners_block2_20260819_193239.jsonl -- gate2: 2026-08-19 19:32, eval to 34544 (150 rounds, 0 train, 17/150) -> FAIL. - Gate starts wins then fails: 7 wins (favorable initial corner), then losses (unfavorable corners). - log: logs/eval_Corners_gate2_20260819_193510.jsonl -- block3: 2026-08-19 19:35, counter 34544 -> 34944 (400 train rounds; 343/400 wins, 0/50 early, - 50/50 last 50; 0 NaN). log: logs/train_Corners_block3_20260819_193622.jsonl -- gate3: 2026-08-19 19:37, eval to 35094 (150 rounds, 0 train, 148/150) -> PASS (exactly at bar) - log: logs/eval_Corners_gate3_20260819_193705.jsonl -- SNAPSHOT: snapshots/best_Corners_r35094 (40 files) -- STATUS: PASS (148/150, gate3) -- MAINTENANCE ROUND 5 (post-Corners PASS): Fire 100, MyFirstDroid 100, Target 100, MyFirstLeader 100, - Crazy 99 — NO regressions. No replays. - Final counter: 35594. - ---- - -## Chunk 1 Summary -- Mastered: Fire, MyFirstDroid, Target, MyFirstLeader, Crazy, Corners (6 bots) -- OPEN: MyFirstBot, PaintingBot -- Final round_counter: 35594 -- Active weights: best_Corners_r35094 restore point -- Snapshots: best_Fire_r11955, best_MyFirstDroid_r11605, best_Target_r12305, - best_MyFirstLeader_r12955, best_Crazy_r13705, best_PaintingBot_r33144, best_Corners_r35094 - (best_MyFirstBot_r15655 = OPEN snapshot) -- Key findings: chunk training causes forgetting of PaintingBot/Corners — pre-chunk weights - (best_Fire_r11955 as proxy) needed to re-learn these. Corners gate is corner-position-variance - sensitive (single JVM, single random corner for full eval run). - ---- - -## Chunk 2 (session 2026-08-19, agent: claude-code) - -### PaintingBot retry -- Restored best_PaintingBot_r33144 weights -- Training block: 200 rounds, 15.5% win rate (exploration noise; 57% in last 50 rounds) -- Gate eval 1 (frozen, stochastic): 145/150 — FAIL (losses rounds 1-5, cold-start) -- Training block 2: 200 more rounds, 200/200 (100%) in training mode -- Gate eval 2: 147/150 — FAIL (losses rounds 1, 2, 6) -- Gate eval 3: 146/150 — FAIL (losses rounds 1, 2, 5, 6) -- Deterministic eval fix attempted (mean action instead of sampling): 0/150 — REVERTED - - Root cause: policy trained with Gaussian noise; mean action is degenerate - - PPOB_EVAL_ONLY already correct: freezes weights but keeps stochastic sampling -- PaintingBot: OPEN (cold-start EnemyTracker history buffer = ~3-5 early-round losses; 97%+ once warmed up) -- Snapshot: best_PaintingBot_r36603 - -### Maintenance evals (from best_Corners_r35094 restore) -- Fire: 100/100 PASS -- MyFirstDroid: 100/100 PASS -- Target: 100/100 PASS -- MyFirstLeader: 100/100 PASS -- Crazy: 100/100 PASS -- Corners: 93/100 FAIL → replay 200 rounds (200/200 training) → re-eval 100/100 PASS - -### MyFirstBot investigation -- 500 training rounds: 17 wins in rounds 1-18 (carry-over), then 0/482 -- Total across chunks: ~2400 rounds, near-zero learning signal -- Root cause analysis — architectural hard counter: - 1. No bullet-in-flight state representation → can't correlate fire decisions with dodge response - 2. Aim space too tight (±12.5% arena around enemy) → can't aim at predicted post-dodge position - 3. Credit assignment chain too long: fire → bullet travel → hit → dodge → needed different aim - 4. State[15] (enemy turn rate) CAN detect the dodge, but can't ANTICIPATE it - 5. MyFirstBot's key mechanic: 90° perpendicular turn on every bullet hit + constant gun sweep -- MyFirstBot: OPEN (needs state/action space architecture changes, not more training) - -### Chunk 2 Summary -- Mastered: Fire, MyFirstDroid, Target, MyFirstLeader, Crazy, Corners (6/8 — unchanged) -- OPEN: PaintingBot (cold-start variance, 97%+ capable), MyFirstBot (architectural gap) -- Final round_counter: ~38003 -- Active weights: post-maintenance with Corners replay -- Next steps: - - PaintingBot: add EnemyTracker pre-population or warm-up rounds to fix cold-start - - MyFirstBot: add bullet-in-flight features to state vector; widen aim space; or accept as beyond current architecture