chore: gitignore binaries/artifacts, remove drafts and stale files

This commit is contained in:
2026-08-27 18:28:21 +02:00
parent 06ed905033
commit 2097c2e6fa
3 changed files with 11 additions and 285 deletions
+11
View File
@@ -22,5 +22,16 @@ __pycache__/
devbox.json
devbox.lock
# Compiled bot binaries (no extension, bot-specific directories)
OscillatorBot_garage/OscillatorBot
SAC_LSTM_Bot_garage/SAC_LSTM_Bot
# Java compiled classes
*.class
# Log and snapshot directories
**/logs/
**/snapshots/
# Git worktrees
worktrees/
@@ -1,98 +0,0 @@
# Implementazione Agente DRL Recurrent-SAC (SAC-GRU) Nativo per Robocode TankRoyale
Bot nativo basato sull'algoritmo **Recurrent Soft Actor-Critic (SAC-GRU)** addestrato per combattere nell'ambiente **Robocode TankRoyale**.
Il bot deve connettersi direttamente al server WebSocket di TankRoyale tramite un'interfaccia client API nativa, senza l'uso di middleware o socket intermediari extra. Il sistema deve garantire il rispetto del limite rigido di **30 ms per turno** imposto dal simulatore, eseguendo l'inferenza e l'addestramento in modo concorrente all'interno dello stesso processo.
---
## 1. Architettura Multi-Thread e Concorrenza (30 ms Tick Constraint)
Per evitare il fenomeno degli "skipped turns", l'applicazione deve essere suddivisa in due thread principali coordinati nello stesso processo:
1. **Thread 1: Realtime Async Event Loop (WebSocket Client)**
* Riceve i messaggi di evento dal server TankRoyale via WebSocket ad ogni tick.
* Estragga e normalizza il vettore di stato a 17 dimensioni ($S_t$).
* Esegue l'inferenza rapida dell'Actor ($< 2 \text{ ms}$) fornendo lo stato $S_t$ e lo stato nascosto corrente della GRU ($h_{t-1}$).
* Mappa le azioni restituite e invia immediatamente il `BotIntent` al server WebSocket.
* Invia la transizione $(S_t, A_t, R_t, S_{t+1}, D_t)$ a una coda/canale thread-safe (`Thread-safe Channel`).
2. **Thread 2: Background Trainer Thread (SAC-GRU Engine)**
* Preleva le transizioni dal canale e le accumula in un **Sequential Replay Buffer**.
* Quando il buffer contiene un numero sufficiente di esperienze, estrae mini-batch di sequenze temporali.
* Esegue l'addestramento in background (Forward/Backward Pass dell'Actor, dei Dual Critic e dell'Alpha autotuning) sfruttando l'accelerazione hardware (GPU/CPU).
* Aggiorna periodicamente i pesi della rete Actor usata dal Thread 1 in modo thread-safe (es. scambio atomico di puntatori o mutua esclusione leggera).
---
## 2. Specifiche dell'Ambiente DRL (POMDP)
### 2.1 Vettore di Stato ($S \in \mathbb{R}^{17}$, Normalizzato in $[-1, 1]$)
* `s[0]`: Posizione X propria ($X / \text{width}$)
* `s[1]`: Posizione Y propria ($Y / \text{height}$)
* `s[2]`: Orientamento scafo ($[-\pi, \pi] / \pi$)
* `s[3]`: Velocità lineare propria ($[-8, 8] / 8$)
* `s[4]`: Orientamento cannone ($[-\pi, \pi] / \pi$)
* `s[5]`: Orientamento radar ($[-\pi, \pi] / \pi$)
* `s[6]`: Temperatura del cannone ($[0, 3] / 3$)
* `s[7]`: Energia propria ($[0, 100] / 100$)
* `s[8]`: Distanza dal muro NORD ($(\text{height} - Y) / \text{height}$)
* `s[9]`: Distanza dal muro SUD ($Y / \text{height}$)
* `s[10]`: Distanza dal muro EST ($(\text{width} - X) / \text{width}$)
* `s[11]`: Distanza dal muro OVEST ($X / \text{width}$)
* `s[12]`: Ultima distanza rilevata del nemico ($[0, \text{max\_dist}] / \text{max\_dist}$)
* `s[13]`: Angolo relativo (bearing) del nemico ($[-\pi, \pi] / \pi$)
* `s[14]`: Orientamento del nemico ($[-\pi, \pi] / \pi$)
* `s[15]`: Velocità del nemico ($[-8, 8] / 8$)
* `s[16]`: Energia residua del nemico ($[0, 100] / 100$)
### 2.2 Vettore delle Azioni Continuo ($A \in \mathbb{R}^4$, Output $[-1, 1]$)
* `a[0]`: Rotazione scafo $\rightarrow$ Mappato su $[-10^\circ, +10^\circ]$ per tick.
* `a[1]`: Traslazione $\rightarrow$ Mappato su $[-8, +8]$ px/tick.
* `a[2]`: Rotazione cannone $\rightarrow$ Mappato su $[-20^\circ, +20^\circ]$ per tick.
* `a[3]`: Potenza di sparo $\rightarrow$ $\text{ReLU}(a_3) \times 3.0$ (Spara solo se $> 0.1$).
### 2.3 Reward Function
$$R_t = R_{\text{danno\_inflitto}} - R_{\text{danno\_subito}} + R_{\text{vittoria/sconfitta}} - R_{\text{muri}} - R_{\text{sparo\_vuoto}}$$
* Danno inflitto: $+ (4p + 2(p - 1))$ con $p \le 3$.
* Danno subito: $- (4p_{\text{nemico}} + 2(p_{\text{nemico}} - 1))$.
* Muri: $-5.0$ per tick di impatto.
* Sparo a vuoto: $-0.1 \times p$.
* Vittoria/Sconfitta: $+20.0$ / $-10.0$.
---
## 3. Modelli Neurali e Meccanismi di Addestramento
Tutte le reti devono integrare una cella ricorsiva **GRU (Gated Recurrent Unit)** per gestire la parziale osservabilità dell'arena (radar in rotazione):
* **Recurrent Actor Network $\pi_\phi(a_t | s_{:t}, h_{t-1})$:**
* Feature Extractor: Linear/Dense ($17 \rightarrow 128$) + Attivazione
* Memory: GRU Layer (Input: $128$, Hidden: $128$)
* Heads: Linear/Dense ($128 \rightarrow 64$) $\rightarrow$ Outputs: $\mu \in \mathbb{R}^4$ e $\log \sigma \in \mathbb{R}^4$ (Clamped $[-20, 2]$)
* Sampling: Reparameterization Trick con squashing $\tanh$.
* **Recurrent Dual-Critic Networks $Q_{\theta_1, \theta_2}(s_{:t}, a_{:t}, h_{t-1})$:**
* Fusion: Concatenazione di Stato e Azione ($17 + 4 = 21$)
* Feature Extractor: Linear/Dense ($21 \rightarrow 128$) + Attivazione
* Memory: GRU Layer (Input: $128$, Hidden: $128$)
* Q-Head: Linear/Dense ($128 \rightarrow 64$) $\rightarrow$ Output Q-Value scalare.
* **Sequential Replay Buffer & Burn-in Strategy:**
* Campionamento di sequenze temporali contigue di lunghezza $L = 24$.
* **Burn-in ($L_{\text{burn}} = 8$ step):** I primi 8 step vengono usati unicamente per aggiornare lo stato nascosto $h$ della GRU, senza calcolo di loss o backpropagation.
* **Training ($L_{\text{train}} = 16$ step):** I successivi 16 step calcolano le loss dell'Actor, dei Critic e l'autotuning del parametro di temperatura $\alpha$ (con target $H_{\text{target}} = -4.0$).
---
## 4. Requisiti di Strutturazione del Codice
Fornisci il codice sorgente completo, modulare e pronto all'uso articolato nelle seguenti componenti:
1. **Model Definitions:** Classi/Strutture per Actor, Critic e Target Networks basate su GRU.
2. **Sequential Replay Buffer:** Struttura dati thread-safe con supporto per estrazione di sequenze e gestione Burn-in.
3. **SAC-GRU Trainer Engine:** Algoritmo di aggiornamento, calcolo loss, gradient clipping ($1.0$) e soft update ($\tau = 0.005$).
-187
View File
@@ -1,187 +0,0 @@
# PPO Curriculum — Running State File
Chunk 1 of larger campaign. Source of truth for resuming.
Bots (easy->hard order): Fire, MyFirstDroid, Target, MyFirstLeader, Crazy, MyFirstBot, PaintingBot, Corners.
Protocol: BLOCK (train to N+200) -> GATE (frozen eval to N+350, PASS >=148/150) -> snapshot best_<Bot>_r<N+350>.
Maintenance from 3rd mastered bot on: frozen eval 100 rounds per mastered bot; bar >=97/100; else replay.
GATE bar: >=148/150. MAINT bar: >=97/100. Max 3 attempts per bot.
## Baseline (verified pre-chunk)
- round_counter.txt = 10905
- Weights proven (frozen eval): Fire/Target/Crazy/MyFirstDroid/MyFirstLeader 100%,
MyFirstBot/PaintingBot 99%, Corners 93%, VelocityBot 10%, TrackFire/RamFire/SpinBot/Walls 0%,
MyFirstTeam 0% (1v5, flagged).
- Bot launchers chmod-fixed. Logs rotate per phase. Snapshots = cp -a of weights/latest.
---
## Bot 1: Fire
- block1: 2026-08-19 16:41, counter 10905 -> 11105 (200 train rounds, 200 train lines, 199W/1L, 0 NaN, no crash)
log: logs/train_Fire_block1_20260819_164245.jsonl
- gate1: 2026-08-19 16:42, eval to 11255 (150 rounds, 0 train lines, 149W) -> PASS >=148
log: logs/eval_Fire_gate_20260819_164311.jsonl
- SNAPSHOT best_Fire_r11255: MISSING (operator error: snapshots/ dir absent; cp failed, training overwrote weights before retry). RECOVERY below.
- block2 (REDO): 2026-08-19 16:46, counter 11605 -> 11805 (200 train, 200W, 0 NaN). Log: logs/train_Fire_block2_20260819_164631.jsonl
- gate2 (REDO): 2026-08-19 16:47, eval to 11955 (150 rounds, 0 train lines, 150W) -> PASS. Log: logs/eval_Fire_gate2_20260819_164740.jsonl
- SNAPSHOT: snapshots/best_Fire_r11955 (40 files)
- STATUS: PASS (150/150)
- maint (post-Target): 100/100 PASS. Log: logs/eval_Fire_maint_20260819_164849.jsonl
## Bot 2: MyFirstDroid
- block1: 2026-08-19 16:43, counter 11255 -> 11455 (200 train rounds, 200 train lines, 200W, 0 NaN)
ANOMALY: operator Ctrl-C at round ~127 triggered run.sh crash-restart (">>> Restart #2, 73 rounds remaining"); round sequence in log continuous (11256..11455, no gaps), slice valid. Log: logs/train_MyFirstDroid_block1_20260819_164349.jsonl
- gate1: 2026-08-19 16:44, eval to 11605 (150 rounds, 0 train lines, 150W) -> PASS. Log: logs/eval_MyFirstDroid_gate_20260819_164506.jsonl
- SNAPSHOT: snapshots/best_MyFirstDroid_r11605 (40 files)
- STATUS: PASS (150/150)
- maint (post-Target): 100/100 PASS. Log: logs/eval_MyFirstDroid_maint_20260819_164919.jsonl
## Bot 3: Target
- block1: 2026-08-19 16:48, counter 11955 -> 12155 (200 train, 200W, 0 NaN). Log: logs/train_Target_block1_20260819_164823.jsonl
- gate1: 2026-08-19 16:49, eval to 12305 (150 rounds, 0 train lines, 150W) -> PASS. Log: logs/eval_Target_gate_20260819_164928.jsonl
- SNAPSHOT: snapshots/best_Target_r12305 (40 files)
- STATUS: PASS (150/150)
- maint (post-self): 100/100 PASS. Log: logs/eval_Target_maint_20260819_165033.jsonl
- MAINTENANCE ROUND 1 (post-Target): Fire 100/100, MyFirstDroid 100/100, Target 100/100. NO regressions. No replays.
## Bot 4: MyFirstLeader
- block1: 2026-08-19 17:06, counter 12605 -> 12805 (200 train, 200W, 0 NaN). Log: logs/train_MyFirstLeader_block1_20260819_170632.jsonl
- gate1: 2026-08-19 17:07, eval to 12955 (150 rounds, 0 train, 150W) -> PASS. Log: logs/eval_MyFirstLeader_gate_20260819_170720.jsonl
- SNAPSHOT: snapshots/best_MyFirstLeader_r12955 (40 files)
- STATUS: PASS (150/150)
- maint r2 (post-self): 100/100 PASS. Log: logs/eval_MyFirstLeader_maint_20260819_171104.jsonl
- maint r3 (post-Crazy): 100/100 PASS. Log: logs/eval_MyFirstLeader_maint_20260819_171246.jsonl
- MAINTENANCE ROUND 2 (post-Leader): Fire 100, MyFirstDroid 100, Target 100, MyFirstLeader 100 -> NO regressions, no replays.
## Bot 5: Crazy
- block1: 2026-08-19 17:12, counter 13355 -> 13555 (200 train, 200W, 0 NaN). Log: logs/train_Crazy_block1_20260819_171246.jsonl
- gate1: 2026-08-19 17:13, eval to 13705 (150 rounds, 0 train, 150W) -> PASS. Log: logs/eval_Crazy_gate_20260819_171346.jsonl
- SNAPSHOT: snapshots/best_Crazy_r13705 (40 files)
- STATUS: PASS (150/150)
- maint r3 (post-self): 99/100 PASS. Log: logs/eval_Crazy_maint_20260819_171623.jsonl
- MAINTENANCE ROUND 3 (post-Crazy): Fire 99, MyFirstDroid 100, Target 100, MyFirstLeader 100, Crazy 99 -> NO regressions, no replays.
ANOMALY (logs): mid-run log rename at 17:00:44 (before Droid maint eval finished) split Droid/Target round lines across files; win evidence intact
(Droid 100 game lines win, Target 100 game lines win). Logs: eval_MyFirstDroid_maint_20260819_170044.jsonl (85 round+100 game),
eval_Target_maint_20260819_170055.jsonl (100 game). No crash; counter continuity verified (13805->13905->14005).
## Bot 6: MyFirstBot
- block1: 2026-08-19 17:18, counter 14205 -> 14405 (200 train, 0W/200L, 0 NaN) — REGRESSION vs baseline 99%; policy loses every round.
Log: logs/train_MyFirstBot_block1_20260819_171827.jsonl
- gate1: eval to 14555: 0/150 FAIL. Log: logs/eval_MyFirstBot_gate1_20260819_171920.jsonl
- block2: 2026-08-19 17:22, counter 14555 -> 14955 (400 train, 0W/400L). Log: logs/train_MyFirstBot_block2_20260819_172215.jsonl
- gate2: eval to 15105: 0/150 FAIL. Log: logs/eval_MyFirstBot_gate2_20260819_172327.jsonl
- block3: 2026-08-19 17:24, counter 15105 -> 15505 (400 train, 0W/400L). Log: logs/train_MyFirstBot_block3_20260819_172448.jsonl
- gate3: eval to 15655: 0/150 FAIL. Log: logs/eval_MyFirstBot_gate3_20260819_172625.jsonl
- STATUS: OPEN (3 attempts, 1000 train + 450 gate rounds, 0 wins total; eval mechanicals verified: 0 train lines in gates, counter advances, no NaN). Likely policy basin shift (aggressive brawler vs stationary out-damager). Snapshot anyway: snapshots/best_MyFirstBot_r15655.
- Chunk 2/3 to revisit (possibly via PaintingBot/Corners transfer).
## Bot 7: PaintingBot
- CATASTROPHE (previous session overshoot): run.sh ran unbounded from 15655 to 31694 (16039 rounds)
vs PaintingBot; 3/16027 wins (0%). Policy collapsed. Log: logs/train_PaintingBot_catastrophe_20260819_191718.jsonl
(0 NaN, no coredump, but weights destroyed by repeated loss signal).
- RECOVERY: 2026-08-19 19:17 — restored weights from snapshots/best_Crazy_r13705; round_counter stays 31694.
ROOT CAUSE: Crazy-snapshot weights cannot beat PaintingBot (chunk training caused catastrophic forgetting;
baseline 99% was from pre-chunk Fire-campaign weights). best_Fire_r11955 (closest to baseline) needed.
- block1: 2026-08-19 19:18, counter 31694 -> 31894 (200 train rounds from Crazy restore; 0/200 wins)
log: logs/train_PaintingBot_block1_20260819_192027.jsonl
- gate1: 2026-08-19 19:21, eval to 32044 (150 rounds, 0 train lines, 0/150) -> FAIL
log: logs/eval_PaintingBot_gate1_20260819_192129.jsonl
- RESTORE: best_Fire_r11955 (pre-chunk-adjacent, known 99% vs PaintingBot). Counter stays 32044.
- block2: 2026-08-19 19:23, counter 32044 -> 32444 (400 train rounds, 400/400 wins, 0 NaN)
log: logs/train_PaintingBot_block2_20260819_192325.jsonl
- gate2: 2026-08-19 19:24, eval to 32594 (150 rounds, 0 train lines, 146/150) -> FAIL (bar >=148)
log: logs/eval_PaintingBot_gate2_20260819_192401.jsonl
- block3: 2026-08-19 19:24, counter 32594 -> 32994 (400 train rounds, 400/400 wins, 0 NaN)
log: logs/train_PaintingBot_block3_20260819_192507.jsonl
- gate3: 2026-08-19 19:26, eval to 33144 (150 rounds, 0 train lines, 147/150) -> FAIL (bar >=148)
log: logs/eval_PaintingBot_gate3_20260819_192630.jsonl
- STATUS: OPEN (3 attempts: block1 0/200 from Crazy, block2 400/400 from Fire, block3 400/400 from Fire;
gate results 0/150, 146/150, 147/150 — high variance, bar missed by 1-4 in gates 2+3).
Likely issue: gate runs single JVM with Corners-like random-state dependency (initial corner selection for
PaintingBot = random; policy may be sensitive to specific PaintingBot seed/state at eval time).
Snapshot: snapshots/best_PaintingBot_r33144.
- MAINTENANCE (post-OPEN, 2026-08-19 19:27): Fire 100, MyFirstDroid 100, Target 100, MyFirstLeader 100,
Crazy 99 — NO regressions. No replays.
Logs: eval_Fire_maint_20260819_192658.jsonl, eval_MyFirstDroid_maint_20260819_XXXXXX.jsonl, etc.
## Bot 8: Corners
- NOTE: Corners bot switches corner on death when losing. Single-JVM gate has fixed initial corner;
policy must handle all 4 corner positions. Training (each JVM = new random corner) cycles all 4.
- block1: 2026-08-19 19:29, counter 33644 -> 33844 (200 train rounds from PaintingBot weights; 78/200 wins,
declining trend: 1/50 in last 50). Training degraded policy.
log: logs/train_Corners_block1_20260819_192947.jsonl
- gate1: 2026-08-19 19:30, eval to 33994 (150 rounds, 0 train, 0/150) -> FAIL
log: logs/eval_Corners_gate1_20260819_193050.jsonl
- RESTORE: best_Fire_r11955 (same as PaintingBot recovery). Counter stays 33994.
- block2: 2026-08-19 19:32, counter 33994 -> 34394 (400 train rounds from Fire restore; 231/400 wins,
improving: 0/50 early, 50/50 last 50; 0 NaN). Policy converges during training.
log: logs/train_Corners_block2_20260819_193239.jsonl
- gate2: 2026-08-19 19:32, eval to 34544 (150 rounds, 0 train, 17/150) -> FAIL.
Gate starts wins then fails: 7 wins (favorable initial corner), then losses (unfavorable corners).
log: logs/eval_Corners_gate2_20260819_193510.jsonl
- block3: 2026-08-19 19:35, counter 34544 -> 34944 (400 train rounds; 343/400 wins, 0/50 early,
50/50 last 50; 0 NaN). log: logs/train_Corners_block3_20260819_193622.jsonl
- gate3: 2026-08-19 19:37, eval to 35094 (150 rounds, 0 train, 148/150) -> PASS (exactly at bar)
log: logs/eval_Corners_gate3_20260819_193705.jsonl
- SNAPSHOT: snapshots/best_Corners_r35094 (40 files)
- STATUS: PASS (148/150, gate3)
- MAINTENANCE ROUND 5 (post-Corners PASS): Fire 100, MyFirstDroid 100, Target 100, MyFirstLeader 100,
Crazy 99 — NO regressions. No replays.
Final counter: 35594.
---
## Chunk 1 Summary
- Mastered: Fire, MyFirstDroid, Target, MyFirstLeader, Crazy, Corners (6 bots)
- OPEN: MyFirstBot, PaintingBot
- Final round_counter: 35594
- Active weights: best_Corners_r35094 restore point
- Snapshots: best_Fire_r11955, best_MyFirstDroid_r11605, best_Target_r12305,
best_MyFirstLeader_r12955, best_Crazy_r13705, best_PaintingBot_r33144, best_Corners_r35094
(best_MyFirstBot_r15655 = OPEN snapshot)
- Key findings: chunk training causes forgetting of PaintingBot/Corners — pre-chunk weights
(best_Fire_r11955 as proxy) needed to re-learn these. Corners gate is corner-position-variance
sensitive (single JVM, single random corner for full eval run).
---
## Chunk 2 (session 2026-08-19, agent: claude-code)
### PaintingBot retry
- Restored best_PaintingBot_r33144 weights
- Training block: 200 rounds, 15.5% win rate (exploration noise; 57% in last 50 rounds)
- Gate eval 1 (frozen, stochastic): 145/150 — FAIL (losses rounds 1-5, cold-start)
- Training block 2: 200 more rounds, 200/200 (100%) in training mode
- Gate eval 2: 147/150 — FAIL (losses rounds 1, 2, 6)
- Gate eval 3: 146/150 — FAIL (losses rounds 1, 2, 5, 6)
- Deterministic eval fix attempted (mean action instead of sampling): 0/150 — REVERTED
- Root cause: policy trained with Gaussian noise; mean action is degenerate
- PPOB_EVAL_ONLY already correct: freezes weights but keeps stochastic sampling
- PaintingBot: OPEN (cold-start EnemyTracker history buffer = ~3-5 early-round losses; 97%+ once warmed up)
- Snapshot: best_PaintingBot_r36603
### Maintenance evals (from best_Corners_r35094 restore)
- Fire: 100/100 PASS
- MyFirstDroid: 100/100 PASS
- Target: 100/100 PASS
- MyFirstLeader: 100/100 PASS
- Crazy: 100/100 PASS
- Corners: 93/100 FAIL → replay 200 rounds (200/200 training) → re-eval 100/100 PASS
### MyFirstBot investigation
- 500 training rounds: 17 wins in rounds 1-18 (carry-over), then 0/482
- Total across chunks: ~2400 rounds, near-zero learning signal
- Root cause analysis — architectural hard counter:
1. No bullet-in-flight state representation → can't correlate fire decisions with dodge response
2. Aim space too tight (±12.5% arena around enemy) → can't aim at predicted post-dodge position
3. Credit assignment chain too long: fire → bullet travel → hit → dodge → needed different aim
4. State[15] (enemy turn rate) CAN detect the dodge, but can't ANTICIPATE it
5. MyFirstBot's key mechanic: 90° perpendicular turn on every bullet hit + constant gun sweep
- MyFirstBot: OPEN (needs state/action space architecture changes, not more training)
### Chunk 2 Summary
- Mastered: Fire, MyFirstDroid, Target, MyFirstLeader, Crazy, Corners (6/8 — unchanged)
- OPEN: PaintingBot (cold-start variance, 97%+ capable), MyFirstBot (architectural gap)
- Final round_counter: ~38003
- Active weights: post-maintenance with Corners replay
- Next steps:
- PaintingBot: add EnemyTracker pre-population or warm-up rounds to fix cold-start
- MyFirstBot: add bullet-in-flight features to state vector; widen aim space; or accept as beyond current architecture