b509195ee9
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
333 lines
31 KiB
Markdown
333 lines
31 KiB
Markdown
## The story so far, in simple words
|
||
|
||
This project trains a robot tank. It plays many fights against other tanks.
|
||
After each fight it changes itself a little. It keeps the changes that helped it win.
|
||
|
||
Night 1 (run 1) finished without problems. It ran for 14 hours alone. It never crashed.
|
||
It beat an old copy of itself most of the time. It beat Crazy about half the time.
|
||
Three things went badly. First, learning was not stable. Good skill appeared, then disappeared again.
|
||
Second, the saved "best" version came from one lucky perfect score. It was not really its best.
|
||
Third, the bot learned to hide and survive. It almost never shot back.
|
||
|
||
The human approved five fixes. All five were put into the code.
|
||
Run 2 used these fixes. Its error numbers grew far too big. Learning broke.
|
||
We made one speed number smaller. This number sets how fast one part learns.
|
||
Then we dropped the broken progress and started clean. This is run 3. It is running now.
|
||
|
||
Next we watch run 3. One of three doors will open.
|
||
Door 1: it stays steady. We let it run to the end.
|
||
Door 2: the numbers grow too big again. We turn the next speed number down.
|
||
Door 3: it stays steady but still fights badly. We teach aiming as a separate, direct lesson.
|
||
|
||
Updated: 2026-08-23 — this section is refreshed at every major step.
|
||
|
||
## Small dictionary
|
||
|
||
- **training**: the time when the bot plays fights and changes itself to improve. It learns only during training.
|
||
- **battle**: one group of fights against one opponent. The bot restarts between groups.
|
||
- **round**: one single fight. Win it by destroying the enemy tank or outliving it.
|
||
- **chunk**: one work block: a battle of up to 10 rounds, then some learning from it.
|
||
- **eval (test match)**: a test match. The bot does not learn during these. We use them only to measure.
|
||
- **win rate**: how many test matches were won, as a percent. 8 wins in 10 matches = 80%.
|
||
- **checkpoint**: a saved copy of the bot's brain (a zip file). Written every few learning steps.
|
||
- **"best" checkpoint**: the saved copy we currently call best. Run 1 picked one from a lucky score, hence the quotes.
|
||
- **replay buffer**: the bot's memory of past moments: what it saw, did, and received. Learning picks old moments from it.
|
||
- **loss (critic/actor)**: a number saying how wrong the bot's inner guesses are. Lower usually means better. Losses growing huge mean trouble.
|
||
- **alpha**: a dial setting how much the bot tries new moves instead of repeating known good ones.
|
||
- **MA / composite score**: MA is the average of the last few win rates; it smooths luck. Composite is the average of MAs across all test opponents.
|
||
- **twin (SacTwin)**: a frozen copy of our own bot, used as a practice partner. Beating it proves real improvement.
|
||
- **lever**: one numbered change we prepared, waiting for approval. There are levers 1 to 5.
|
||
- **watchman**: a helper who checks the running training at set times and stops it if something breaks.
|
||
|
||
---
|
||
|
||
# Campaign Notebook — campaign-v1 (SAC_LSTM_Bot)
|
||
|
||
Overnight training campaign on branch `research/goto-controller`.
|
||
Companion tickets: config+launch = **#56**, morning verdict = **#57**, map = **#53**.
|
||
Monitoring contract: observability inventory in **#55 comment 461** (9 signals, thresholds, four-way discrimination).
|
||
|
||
## Locked config (campaign-v1)
|
||
|
||
Launch command (tmux session `sac_campaign`, stdout teed to `campaign_stdout.log`):
|
||
|
||
```bash
|
||
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1' \
|
||
SAC_EVAL_OPPONENT=Corners \
|
||
SAC_TOTAL_ROUNDS=25000 \
|
||
SAC_CHUNK_SIZE=10 \
|
||
SAC_EVAL_INTERVAL=2 \
|
||
SAC_EVAL_ROUNDS=10 \
|
||
SAC_MAX_CRASHES=5 \
|
||
SACLSTM_HIDDEN_SIZE=256 \
|
||
SACLSTM_BATCH_SIZE=16 \
|
||
SACLSTM_UTD_RATIO=1 \
|
||
SACLSTM_SAVE_INTERVAL=5 \
|
||
./sac_train.sh 2>&1 | tee -a campaign_stdout.log
|
||
```
|
||
|
||
| Knob | Value | Source / rationale |
|
||
|------|-------|--------------------|
|
||
| `SAC_OPPONENTS` | `Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1` | Working recommendation kept. Twin pinned at **1 not 2**: #54 showed mirror battles end early ⇒ fewer transitions per chunk; weight 2 would starve the replay buffer. |
|
||
| `SAC_EVAL_OPPONENT` | `Corners` | Pinned explicitly (= first pool entry default, #54 note) — removes reorder footgun. |
|
||
| `SAC_TOTAL_ROUNDS` | `25000` | Sized so the **wall-clock ceiling binds first**: measured throughput ~3400 rounds/h early (drops as battles lengthen) ⇒ 2000 would have exhausted in ~1 h. |
|
||
| `SAC_CHUNK_SIZE` | `10` | Harness default. |
|
||
| `SAC_EVAL_INTERVAL` | `2` | Harness default — eval every ~20 rounds. |
|
||
| `SAC_EVAL_ROUNDS` | `10` | Harness default. |
|
||
| `SAC_MAX_CRASHES` | `5` | Harness self-abort; monitor intervenes earlier at ≥3 consecutive crashes (#55). |
|
||
| `SACLSTM_HIDDEN_SIZE` | `256` | Module default (`network.nim`); real capacity vs #49 smoke's 32; under `MaxHidden`=512 cap. |
|
||
| `SACLSTM_BATCH_SIZE` | `16` | Module default (`integration.nim`). |
|
||
| `SACLSTM_UTD_RATIO` | `1` | Module default. |
|
||
| `SACLSTM_SAVE_INTERVAL` | `5` | **Deviation** from default 500 and from #54's "10–20": at hidden 256 a gradient step takes ~1 s and a chunk process fits only ~10–18 steps (see incident below) — interval must sit **inside the per-process step budget**. 5 ⇒ checkpoint every ~5–10 s of active training; IO trivial (10.5 MB zip, atomic replace). |
|
||
| *(not pinned)* | module defaults | `LR_ACTOR/LR_CRITIC/LR_ALPHA=3e-4`, `GAMMA=0.99`, `TAU=0.005`, `TARGET_ENTROPY=-4.0`, `BUFFER_CAPACITY=500000`, `BURN_IN=8`, `TRAIN_WINDOW=16`. |
|
||
|
||
**Budget**: generous wall-clock **ceiling, not a deadline** — originally **T+12 h** from launch 00:21:55 CEST 2026-08-22 ⇒ 12:22 CEST (epoch 1787394115); **extended 2026-08-22 ~07:35 by orchestrator decision on human mandate** ("no deadlines — let the 25000-round budget complete", ~14:35 projected) ⇒ ceiling now **16:30 CEST 2026-08-22 (epoch 1787409000)**, enforced by a hard user-systemd net unit `sac-ceiling-net` (sleeps to the epoch, then kills the tmux session and any straggler harness processes). While HEALTHY per #55 discrimination rules the run continues; a monitor kills the tmux session at the ceiling or on an intervene threshold.
|
||
|
||
**Fresh start**: pre-campaign `weights/` held #49-smoke 32-hidden checkpoints, incompatible with hidden=256. Archived to `weights_smoke49_backup/`; campaign baseline re-established by probe battles vs Corners (random-init hidden-256 checkpoint, `best_score.txt` reset then re-raised to 20 by a genuine eval). Twin regenerated via `./make_twin.sh` from that baseline (md5 `61521cff…` verified seed).
|
||
|
||
## Phase log
|
||
|
||
Every entry below starts with a plain-language first sentence. Technical detail follows for those who want it.
|
||
|
||
- **2026-08-21 23:22** — Claim posted on #56 (comment 471). Config locked, notebook committed (`2f49cb2`).
|
||
- **2026-08-21 23:27** — Smoke weights archived; release build; bootstrap + probe battles vs Corners established a hidden-256 baseline checkpoint (`sac_latest.zip`, 10.5 MB) and twin seed.
|
||
- **2026-08-21 23:32** — Twin regenerated (md5-verified). **Launch attempt 1** (SAVE_INTERVAL=20, TOTAL_ROUNDS=2000): ran 16+ chunks, evals every 2 chunks — but **zero checkpoints persisted** (see incident). Killed 23:48.
|
||
- **2026-08-21 23:52–00:10** — Diagnosis (see incident): interval=1 fired, interval=2/20 never; instrumentation + /proc thread forensics ⇒ per-process step budget ~10–18 at ~1 s/step; save check ran only between drain-burst passes.
|
||
- **2026-08-22 00:12** — Fix: save check moved inside the gradient-step loop, committed `2653671`. Validated: interval=5 save fired ~12 s into a battle.
|
||
- **2026-08-22 00:21:55** — **Launch (final)**: tmux `sac_campaign`, config above. First campaign save on disk at t+54 s; eval #1 on cadence.
|
||
- **2026-08-22 00:30** — HEALTHY checklist passed (see below).
|
||
- **2026-08-22 03:45** — Watch shift 1 (00:34–03:35): liveness flawless (10/10 HEALTHY, zero banners, zip ≤30 s). Learning signal: steady-state eval vs Corners 0–5% with two isolated 10/10 spikes (~01:15) → capability emerged, then lost. Eval-regression intervene threshold fired per #55; intervention DEFERRED to Campaign verdict (#57) — rationale: n=2 evidence, no loss metrics, buffer-loss on restart, run completes ~08:05 anyway. Milestone issue: #58 "Campaign-v1 watch: eval-regression threshold fired — intervention deferred to verdict".
|
||
- **2026-08-22 ~07:15** — Morning audit: policy demonstrably learning off-benchmark (SacTwin 73→100%, Crazy 26→56%) while Corners eval stays ~0–9% with 3 transient 10/10s; wall-clock ceiling extended to let the 25k complete (~14:35 projected); instability-vs-plateau question left to the curve.
|
||
- **2026-08-22 ~07:50 — score:60 anatomy**: Tank Royale survival(50)+last-survivor(10) awarded when opponent dies while we survive; exactly-60 ⇒ zero damage dealt by us that round (opponent self-destructed via wasted shots + 0.1/turn inactivity drain). 3,878 rounds (19%); modal vs SacTwin; vs Corners 986 damageless outlives vs 242 true wins; combined with 43% of rounds being score:0, texture = survivor-not-fighter against walls. RL rewards are event-driven (rewards module), so behavioral evidence, not reward poisoning. Feeds #57 levers: aggression shaping / specialist-vs-generalist.
|
||
- **2026-08-22 ~08:05 — `.part` debris forensics + sweep**: 99 `*.zip.tmp.*.part` (483 MB) are NOT weights.nim debris (that proc uses fixed `.tmp` + finally-cleanup, working); naming matches an external write-temp→rename copier killed mid-write, bursts correlating with kill events; possible culprit: a folder-sync client fighting a file that changes every ~20 s (**human asked to confirm**). Swept with `-mmin +10` age guard (protects in-flight writes); 483 MB freed, real zips untouched — note 2 fresh `.part` reappeared minutes later, copier still active.
|
||
- **2026-08-22 ~08:00 — ceiling defused**: the 12:22 'self-kill' was notebook prose instructing watchmen — never an OS mechanism; rewritten to 16:30 CEST AND armed a real systemd --user net `sac-ceiling-net` firing 16:30:00 (epoch 1787409000); training uninterrupted (counter +99/147 s verified); annotated #56 comment 487.
|
||
- **2026-08-22 ~14:28 — CAMPAIGN ENDED NATURALLY**: banner `>>> training complete: 2500 chunks` after **14 h 07 m** (00:21:55 → ~14:28 CEST); round_counter **38912**; **zero crash banners across the whole run**; the `sac-ceiling-net` backstop never fired — cancelled unneeded. Budget note: the harness loop is **chunk-based** — `SAC_TOTAL_ROUNDS=25000` ÷ `CHUNK_SIZE=10` ⇒ **2500 chunk battles** of ≤10 rounds each; the "25k-rounds" label was a misnomer (the counter also accrues 1287×10 eval rounds and rerun chunks). Final eval vs Corners: 10%.
|
||
- **2026-08-22 ~14:50 — VERDICT posted (→ #57)**: **RETUNE BEFORE SCALING.** Ops layer PROVEN (14 h autonomous, zero crashes, self-healing restarts, natural completion — the harness scales); learning REAL BUT NARROW (within-opponent gains genuine — SacTwin 90.7%, Crazy 26→48% — but specialist-not-generalist, walls untouched; 12 eval spikes ≥8/10 incl. 5×10/10, none retained); benchmark pathology: Corners-only deterministic eval + single-max best gating froze `sac_best.zip` at 01:01:57 on a fluke 10/10. Five code-level levers staged **awaiting human sign-off**; v2 NOT launched. Full rationale: Results below + #57 resolution comment.
|
||
- **2026-08-22 ~14:50 — HYGIENE (campaign over, no live writers — safe)**: `sac-ceiling-net` stopped + `reset-failed` (backstop obsolete); `.part` corpse sweep **48 → 0** (no age guard needed — nothing writes anymore); future-run guard added to `sac_train.sh`: startup `rm -f "$WEIGHTS_DIR"/sac_latest.zip.tmp.*.part "$WEIGHTS_DIR"/sac_latest.zip.tmp` so a SIGKILLed run's libzip modify-path corpses can't accumulate again.
|
||
|
||
## Decision-issue index
|
||
|
||
| Issue | What it decided |
|
||
|-------|-----------------|
|
||
| #37–#48 | Bot built: skeleton, state, actions, rewards, LSTM network, weights, SAC+LSTM training, integration. |
|
||
| #49 | Training harness + smoke run (toy hyperparams: hidden 32). |
|
||
| #54 | Mirror-twin sparring partner; SAVE_INTERVAL persistence rule; eval opponent = first pool entry. |
|
||
| #55 | 9-signal observability inventory; CRASHED/STALLED/SLOW-LEARNER/HEALTHY discriminators; monitor thresholds. |
|
||
| #56 | This campaign: locked config + launch + the save-check fix (`2653671`). |
|
||
| #57 | Morning verdict — consumes this notebook + logs. |
|
||
|
||
## Incidents & checks
|
||
|
||
### Incident 1 — zero checkpoint persistence at production sizes (launch blockers, fixed)
|
||
|
||
**Symptom**: campaign ran 26+ chunk processes across two attempts without a single `sac_latest.zip` update, while rounds/evals flowed normally. #54's rule ("keep `SACLSTM_SAVE_INTERVAL` well below per-chunk gradient-step counts, 10–20 fired in smokes") silently broke at hidden 256.
|
||
|
||
**Diagnosis chain** (all reproducible):
|
||
1. Interval=1 saved within seconds; interval=2 and 20 never saved — through the *same* harness ⇒ not env propagation.
|
||
2. Temporary step instrumentation (bot stderr via a one-line `SAC_LSTM_Bot.sh` redirect — the vendored runner swallows bot stderr, #55 gap S9-adjacent): steps cost **~1.06 s each**; a drain burst queued 53 steps; logging stopped mid-pass while rounds kept completing.
|
||
3. `/proc/<pid>/task` sampling: training thread alive and RUNNING (~13 s CPU per ~40 s process) — not deadlocked, just slow ⇒ **per-process step budget ≈ 10–18 steps**.
|
||
4. The save check lived *between* drain-burst passes; with bursts queueing minutes of steps, `stepCount` never reached `nextSave` before process teardown. Smoke runs masked this: hidden 32 steps were sub-millisecond, so hundreds of steps fit per chunk.
|
||
|
||
**Fix** (commit `2653671`): save check relocated **inside** the step loop (checked every gradient step; `packFull`+`trySend` unchanged). Validated: interval=5 save fires ~12 s into a battle; campaign save fired 54 s after launch.
|
||
|
||
**Config consequences**: `SACLSTM_SAVE_INTERVAL=5` (inside the per-process budget; #54's 10–20 was derived at smoke speeds). `SAC_TOTAL_ROUNDS=25000` (throughput measured ~3400 rounds/h, so 2000 was a 1-hour budget, not an overnight one). OMP_NUM_THREADS=1 tested and **not** needed (hang was step-budget exhaustion, not OpenMP).
|
||
|
||
### Launch health check (t+9 min, 00:30:16) — **HEALTHY** per #55 checklist
|
||
|
||
| Signal | Reading | Verdict |
|
||
|--------|---------|---------|
|
||
| S1 harness stdout | teed to `campaign_stdout.log`; **0** `crash`/`aborted` banners | ✓ |
|
||
| S4 round_counter | 1890, +460 in 9 min (~51 rounds/min) | ✓ advancing |
|
||
| S6 sac_latest.zip mtime | **6 s old**; first save at t+54 s | ✓ fresh |
|
||
| S2 training_log.jsonl | 1492 lines, growing; last ticks=668, plausible | ✓ |
|
||
| S3 eval_log.jsonl | age 3 s (atomic replace); eval every 2 chunks | ✓ on cadence |
|
||
| S5 best_score | 20 (from a genuine campaign-1 eval; non-decreasing) | ✓ |
|
||
| Sampling | 112 chunks: Corners 41 / Crazy 25 / RamFire 15 / SacTwin 16 / Target 15 ≈ weights 3:2:1:1:1 | ✓ plausible |
|
||
| Disk | 418 GB free | ✓ |
|
||
|
||
Early evals 0% vs Corners — expected for a near-random policy minutes in; SLOW-LEARNER watch rule (flat ≥5 evals = watch) applies, never intervene.
|
||
|
||
## Check-in procedure (for monitor sessions)
|
||
|
||
```bash
|
||
tmux capture-pane -p -t sac_campaign | tail -5 # S1: banners, crashes
|
||
cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/round_counter.txt
|
||
stat -c '%Y' ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/sac_latest.zip # age <~600s = training alive
|
||
tail -1 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/training_log.jsonl
|
||
tail -3 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/campaign_stdout.log # eval results / new best
|
||
cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/best_score.txt
|
||
```
|
||
|
||
Intervene per #55 thresholds: ≥3 consecutive `crash #N` banners; ΔS4=0 over ≥15 min; zip mtime >10 min stale while S4 advances (STALLED); disk <1 GB. At the **ceiling (16:30 CEST Aug 22, epoch 1787409000 — extended from 12:22 per human mandate)**: `tmux kill-session -t sac_campaign` if still running (the hard net unit `sac-ceiling-net` fires at the same epoch regardless of monitors) — final state is in `weights/`, logs, and this notebook.
|
||
|
||
## Results (campaign-v1 — filled by #57)
|
||
|
||
**Run**: 2500/2500 chunks · **14 h 07 m** autonomous (00:21:55 → ~14:28 CEST 2026-08-22) · round_counter 38912 · **zero crashes** · graceful banner `>>> training complete: 2500 chunks` · systemd net never fired. Budget was chunk-based (see phase log) — the "25k-rounds" label was a misnomer.
|
||
|
||
**Per-opponent training win rates** (run-3 slice of `training_log.jsonl`):
|
||
|
||
| Opponent | Win rate | Record | Note |
|
||
|----------|----------|--------|------|
|
||
| SacTwin | **90.7%** | 2693/2970 | vs frozen past-self — genuine self-play gain |
|
||
| Crazy | 48.3% | 2965/6140 | doubled from 26% early-run |
|
||
| Target | 9.6% | — | static, barely moved |
|
||
| Corners | 7.2% | — | walls untouched |
|
||
| RamFire | **0%** | 0/2920 | mirrors the PPO-era ladder — ram-class needs dedicated pressure |
|
||
|
||
**Eval vs Corners (pinned benchmark)**: 1287 evals · overall mean **7.7%** · histogram headline: `0/10 = 867 (67%)`, spikes ≥8/10 = **12** (incl. **5× perfect 10/10**) · final eval 10%. Stdout log carries no timestamps; timing reconstructed from file mtimes.
|
||
|
||
**Best-zip paradox**: `sac_best.zip` frozen since **01:01:57** — a single lucky 10/10 at ~round 3.5k wrote `best_score=100`, and no later eval could outrank a perfect score (even genuine ~50%-winrate stretches elsewhere). Best checkpoint = lottery ticket, decoupled from the steady-state policy (which sat at 0–10% vs Corners).
|
||
|
||
**Verdict: RETUNE BEFORE SCALING** — full rationale in #57 resolution comment. Five code-level levers staged for human sign-off: (1) eval rotation across pool + moving-average best gating; (2) reward shaping toward damage/aggression incl. anti-ram signal; (3) training-loss/step metrics logged from the training thread (#55 gap #1); (4) gate `sendTrainingMsg` off in eval mode (#55 gap #3); (5) optional stability knobs (lower LR / entropy coeff) once loss curves exist.
|
||
|
||
---
|
||
|
||
# Campaign Notebook — campaign-v2 (SAC_LSTM_Bot)
|
||
|
||
## Locked config (campaign-v2)
|
||
|
||
Launched verbatim from orchestrator mandate (umbrella ticket #59, all five levers approved & implemented in `a07e530` + `6fc01eb`):
|
||
|
||
```bash
|
||
cd /home/davide/Projects/SirRoboGarage/SAC_LSTM_Bot && \
|
||
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:2,Target:1,SacTwin:1' \
|
||
SAC_EVAL_OPPONENTS='Corners,Crazy,Target' \
|
||
SAC_EVAL_INTERVAL=2 SAC_EVAL_ROUNDS=10 \
|
||
SAC_TOTAL_ROUNDS=25000 SAC_CHUNK_SIZE=10 SAC_MAX_CRASHES=5 \
|
||
SACLSTM_HIDDEN_SIZE=256 SACLSTM_BATCH_SIZE=16 SACLSTM_SAVE_INTERVAL=5 \
|
||
./sac_train.sh 2>&1 | tee -a campaign_v2_stdout.log
|
||
```
|
||
|
||
| Knob | Value | Rationale |
|
||
|------|-------|-----------|
|
||
| `SAC_OPPONENTS` | Corners:3, Crazy:2, **RamFire:2**, Target:1, SacTwin:1 | RamFire bumped 1→2 vs v1: anti-ram shaping (lever 2, `6fc01eb`) needs exposure to fire; without samples there is no gradient signal against ram-class |
|
||
| `SAC_EVAL_OPPONENTS` | Corners,Crazy,Target | Lever-1 rotation set (default); composite = mean of per-opponent MA-5 win rates |
|
||
| `SAC_EVAL_INTERVAL/ROUNDS` | 2 / 10 | Unchanged from v1 cadence |
|
||
| `SAC_TOTAL_ROUNDS/CHUNK_SIZE/MAX_CRASHES` | 25000 / 10 / 5 | Same budget semantics as v1 (chunk-based) |
|
||
| `SACLSTM_HIDDEN_SIZE/BATCH_SIZE/SAVE_INTERVAL` | 256 / 16 / 5 | Architecture + throughput knobs carried over; LR/entropy defaults untouched = **lever-5 conditional posture** (activation decided by loss-curve evidence, not upfront) |
|
||
|
||
## Fresh start & archive
|
||
|
||
v1 state archived intact (notebook references preserved) into `SAC_LSTM_Bot/weights_v1_archive/`: `sac_latest.zip`, `sac_best.zip`, `best_score.txt`, `round_counter.txt`, `training_log.jsonl`, `eval_log.jsonl`, `campaign_stdout.log`, plus the lever-3 smoke leftover `training_metrics.jsonl` (from `src/SAC_LSTM_Bot/`, moved so v2 loss curves start clean for lever-5 reading). Main `weights/` verified empty afterwards ⇒ main bot takes the genuine random-init path (`loadOrInitFull` → `randomFull()`).
|
||
|
||
**Twin reseed**: `make_twin.sh` requires a seed zip, but the fresh-start baseline has none. Generated a fresh random-init checkpoint (hidden=256, alpha=1.0 matching `logAlpha=0`) via a throwaway Nim script against `network.nim`/`weights.nim`, seeded it as `sac_best.zip` transiently, ran `./make_twin.sh`, removed the transient copy. Verified twin dir got byte-identical fresh zips (`cmp` OK; NOT v1 zips) + `round_counter.txt=0`. No script changes needed.
|
||
|
||
## Safety net
|
||
|
||
`systemd-run --user --unit=sac-ceiling-net-v2` armed at launch: sleeps 72000 s then `tmux kill-session -t sac_campaign_v2; sleep 5; pkill -f sac_train.sh`. Unit active at 19:59:40 CEST 2026-08-22, fires **15:59:40 CEST 2026-08-23** (epoch 1787493580).
|
||
|
||
## Launch & health evidence (first ~45 min)
|
||
|
||
Launched 20:00:19 CEST 2026-08-22 (epoch 1787421619), tmux session `sac_campaign_v2`.
|
||
|
||
| Check | Evidence |
|
||
|-------|----------|
|
||
| Round counter advances | `round_counter.txt` 95→100→465 across polls; RunTraining `Counter check passed: N == N` every chunk |
|
||
| Metrics JSONL with scalars | `training_metrics.jsonl` growing (20 lines @ t+45m): full `{epoch, steps, buffer_size, drained, grad_steps, critic_loss, actor_loss, alpha_loss, alpha}` per line |
|
||
| Eval rotation cycles ≥2 | All 3 opponents EVERY cycle: Corners→Crazy→Target ×4+ cycles in `eval_log.jsonl` (10 games each per cycle) |
|
||
| MA files written | `weights/ma_history_{Corners,Crazy,Target}.txt` created at first cycle, appended since |
|
||
| Best-gate on composite only | First write exactly when composite 0.0000 > −1 (missing-file default); later 0% cycles correctly did NOT rewrite (strict improvement enforced) |
|
||
| No transitions during eval windows | Lever-4 gate active by construction (`sendTrainingMsg` drops all msgs under `SACLSTM_EVAL_MODE=1`, unit-tested); metrics epochs cluster at chunk boundaries |
|
||
| Zero crash banners | `grep -c 'crash #'` = 0 through 20 chunks |
|
||
| Sampling distribution plausible | 20 chunks: Corners 10, RamFire 5, Crazy 3, SacTwin 2, Target 0 — within small-n noise of weights (3/2/2/1/1)/9; RamFire already sampled (exposure goal met) |
|
||
|
||
Twin liveness: own `round_counter.txt` advancing, own `sac_latest.zip` updating during SacTwin chunks, own metrics file separate from the main bot's.
|
||
|
||
### Watch items (not blockers)
|
||
|
||
1. **Training-throughput signature**: metrics lines consistently show `steps=1, buffer_size=24 (= burnIn 8 + trainWindow 16, i.e. exact canSample threshold), drained=1` — one gradient step per pass at threshold-crossing moments rather than large drain bursts. Mechanism unexplained by static code read (per-tick sends should yield bigger bursts); v1 learned to its score-60 state under the same integration code without instrumentation, so learning is not obviously broken — but effective grad-steps/hour is THE number to check at first review. This is precisely what lever-3 instrumentation exists to surface.
|
||
2. **Early critic-loss spikes**: two `1.56e16` outliers (t+6:16, t+12:04) amid otherwise sane values (~8–35) — same class as the random-init artifact flagged in #59 phase-2 notes (~1e12 there); expect decay. If persistent past early chunks, feeds the lever-5 decision.
|
||
3. **`.part` corpses**: three `sac_latest.zip.tmp.*.part` files accumulated mid-run (libzip interrupted-write artifact, #57 forensics); harmless — atomic renames keep the main zips valid, startup sweep clears them next restart.
|
||
|
||
## ~21:30 — v2 attempt-1 checkpoint finding
|
||
|
||
Attempt-1's rolling zip was dead ~40 min at birth: `SACLSTM_SAVE_INTERVAL` counts **gradient steps**, and each training pass inside the short battle processes carries exactly **1 step** (the `steps=1, drained=1` signature of watch item 1) ⇒ interval-5 never reached its trigger within a process lifetime — no `sac_latest.zip` roll ever fired despite healthy training. Same class as v1's 23:32 incident, resurfacing through a different seam (per-process step budget vs per-pass step count).
|
||
|
||
**Fix (env-level, no rebuild)**: `SACLSTM_SAVE_INTERVAL=1`, restart 21:09:30 CEST. Saves verified flowing: 21:09 / 21:13 / 21:16 (and still flowing at audit time — `sac_latest.zip` mtime 21:22:28, `sac_best.zip` 21:12:50). The locked-config table above keeps the original attempt-1 launch line for the record; live relaunch differs only in this knob.
|
||
|
||
## ~21:30 — twin-freeze contract corrected
|
||
|
||
`make_twin.sh` twin launcher now exports `SACLSTM_EVAL_MODE=1` (commit `19f34ab`). Without lever-4's gate on the twin side, the **v1 SacTwin had been TRAINING throughout**, not frozen — every mirror battle updated the twin's own weights. Consequence: v1's headline "**90.7% vs twin**" is retroactively an **arms-race win rate** (both policies co-evolving), not a fixed-benchmark score, and the #57 verdict phrasing implying a frozen sparring partner is corrected by this entry. No numbers change; the story does.
|
||
|
||
## ~21:30 — net extended + pacing decision
|
||
|
||
Slowdown attribution complete: **89% of wall clock = eval rotation by design** (`SAC_EVAL_INTERVAL=2` × 3 opponents × 10 rounds each); trainer exonerated at **~500 ticks/s**. Decision recorded in issue **#60**.
|
||
|
||
Ceiling net re-armed to match the extended budget: fires **Thu 2026-08-27 07:59:40 CEST (epoch 1787810380)** — supersedes the 2026-08-23 15:59:40 fire noted under Safety net. Attempt-1 artifacts preserved in `/tmp/v2_attempt1_backup/`.
|
||
|
||
# Campaign v2 — attempt-3 (2026-08-23)
|
||
|
||
## ~06:45 — LEVER-5 TRIGGERED (watchman shift 2, 02:11–06:22 CEST)
|
||
|
||
Evidence-gated activation of the lever-5 conditional posture (`SACLSTM_LR_CRITIC`), human signed off. Trigger evidence:
|
||
|
||
| Signal | Observation |
|
||
|--------|-------------|
|
||
| actor_loss | \|34M\| → \|106M\| monotone (+15M/h), **no plateau** |
|
||
| critic_loss | spikes >1e12 in **100%** of NEW metric records; campaign share 80.2% |
|
||
| Signature | exact `1.5625e16` recurring (= float32 saturation neighborhood) |
|
||
| alpha | decayed 0.951 → 0.711 (entropy collapse under runaway Q scale) |
|
||
| best_score.txt | frozen 00:32 (=83.3333) through 5h50m of composite oscillation 3.3–43.3 |
|
||
| Per-opponent MAs | whipsawing (Crazy 30→100→90→0→60→90→10→0) |
|
||
|
||
## DECISION — LR_CRITIC=1e-4, single knob, fresh start
|
||
|
||
- **Adjudication**: primary suspect is critic/Q value-scale growth. Actor loss inherits Q magnitude through the policy gradient, so the actor explosion is downstream; alpha decay is a *symptom* (entropy temperature chasing a blown value scale), not a cause. Therefore ONE knob moves: `SACLSTM_LR_CRITIC=1e-4` (critic learns slower → Q estimates stay estimable); LR_ACTOR/LR_ALPHA stay at default 3e-4, TARGET_ENTROPY −4. Multi-knob changes would confound attempt-3's attribution.
|
||
- **Fresh start over warm start**: attempt-2's weights carry saturated Q representations; warm-starting them under a new LR would inherit the pathology we're trying to escape (contamination risk). Attempt-2 archived intact, nothing discarded.
|
||
|
||
## MA-wipe root cause FOUND & FIXED (`167bcc4`)
|
||
|
||
Watchman anomaly explained: `ma_history_*.txt` held exactly one leading-space value per file (e.g. `[ 10 ]`, `[ 20 ]`, `[ 0 ]` in the attempt-2 backup) instead of the designed 5-value history, degrading the composite best-gate to last-cycle mean — which fully accounts for the composite whipsaw 3.3–43.3.
|
||
|
||
Root cause: in `eval_checkpoint`, the append line was
|
||
|
||
```bash
|
||
printf '%s\n' "$(cat "$(ma_hist_file "$opp")")" "$wr" | tail -n 5 | tr '\n' ' ' > "$(ma_hist_file "$opp")"
|
||
```
|
||
|
||
Bash sets up the `>` redirect **before** running the command substitution, so `$(cat f)` always read the freshly truncated file ⇒ every cycle wiped history to `" $wr "`. Reproduced standalone on bash 5.3 before touching the script; no other writer exists (grep). Fix hoists the read into its own statement and word-splits it so `tail -n 5` keeps exactly the last MA_WINDOW values. Verified live in attempt-3: two eval cycles produced `[0 0 ]` per opponent (previously impossible).
|
||
|
||
## Attempt-3 launch (2026-08-23 07:01:28 CEST, epoch 1787461288)
|
||
|
||
- Attempt-2 stopped cleanly (tmux kill; no stray procs); artifacts (76 items, 336 MB incl. `.part` corpses + MA/best/latest/counter) → `/tmp/v2_attempt2_backup/`; final round_counter **8452**; `weights/` verified empty.
|
||
- Twin regenerated via documented transient-seed method: throwaway Nim script (`initSACTrainer(35,4)` + `saveWeights`, hidden=256 env, alpha=1.0) seeded as transient `weights/sac_best.zip` → `./make_twin.sh` → twin zips byte-identical to seed (`cmp` OK), twin counter=0 → transient removed.
|
||
- Launch line = locked config verbatim plus the two live deltas:
|
||
|
||
```bash
|
||
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:2,Target:1,SacTwin:1' \
|
||
SAC_EVAL_OPPONENTS='Corners,Crazy,Target' \
|
||
SAC_EVAL_INTERVAL=2 SAC_EVAL_ROUNDS=10 \
|
||
SAC_TOTAL_ROUNDS=25000 SAC_CHUNK_SIZE=10 SAC_MAX_CRASHES=5 \
|
||
SACLSTM_HIDDEN_SIZE=256 SACLSTM_BATCH_SIZE=16 SACLSTM_SAVE_INTERVAL=1 \
|
||
SACLSTM_LR_CRITIC=1e-4 \
|
||
./sac_train.sh 2>&1 | tee -a campaign_v2_stdout.log
|
||
```
|
||
|
||
- Safety net: existing transient unit `sac-ceiling-net-v2.service` re-verified active; fire epoch start+384587 s = **1787810380 = Thu 2026-08-27 07:59:40 CEST exactly** (delta 0 vs mandate) — kept, not duplicated.
|
||
|
||
### Health baseline (t+15 m)
|
||
|
||
| Check | Evidence |
|
||
|-------|----------|
|
||
| Counter | 33 (t+8m) → **125** (t+15m) |
|
||
| Metrics JSONL | flowing (6 lines); baseline line: `critic_loss=23.09, actor_loss=-3.228, alpha_loss=0.0, alpha=0.99970` (random-init magnitudes — attempt-3's divergence curve starts here); t+15m line: critic 84.5, actor −9.99, alpha 0.9937 |
|
||
| Saves | `SAVE_INTERVAL=1`: `sac_latest.zip` written from first grad-step onward |
|
||
| Eval rotation | full cycle landed: Corners/Crazy/Target ×10 games each in `eval_log.jsonl` |
|
||
| MA files | `[0 0 ]` per opponent after two cycles — fix confirmed in vivo |
|
||
| Best gate | `best_score.txt=0.0000` written on first composite (0 > −1 default) |
|
||
| Crash banners | 0 |
|
||
|
||
|
||
|
||
### Progress graphs
|
||
|
||
One live dashboard: `docs/campaign_dashboard.svg` (current run only — test wins, real-fight wins, losses, alpha, throughput; auto-reloads every 60 s when open in Chrome). Keep it fresh with `tools/watch_dashboard.sh` (regenerates every 60 s), or one-shot `python3 tools/plot_progress.py` (pure stdlib; paths overridable via argv, `--selftest` for sanity check).
|
||
|
||
## ~10:47 — the dashboard's axes were upside-down since creation
|
||
|
||
The progress graphs have been lying since they were made: 0% was drawn at the TOP of every panel and the newest games appeared on the LEFT. The cause is a one-line formula bug in `tools/plot_progress.py`: `map_fn` interpolated as `p1 - t*(p1-p0)` instead of `p0 + t*(p1-p0)`, so all five panels plotted `100 - value` on y and reversed time on x. Tick labels were computed by separate (correct) code, which is why the numbers on the axes never matched the ink.
|
||
|
||
Fix + guard: formula corrected; every call site audited (panels 1/2/5 use `map_fn` for both axes and are fixed by the same line; panels 3/4 already used a correct local x-lambda; no other consumer of `map_fn` exists in the repo). `--selftest` now renders a known rising series through the full build path and fails loudly unless higher value = smaller SVG y and newer data = further right — proven to catch this exact bug when the old formula is re-injected. Dashboard regenerated from live logs.
|
||
|
||
Honest status while reading the now-correct charts: run 3 is only hours old and winning ~0% — recent evals are 0/10 vs Corners, Crazy and Target alike, and real-fight buckets sit at 0–1 wins per 100 games. Expected for a fresh brain. Night-1's gains were real but were intentionally reset by the stability restart that began attempt-3; the curve starts from zero again here.
|