From 781e41595e870b78e995d7e56f5e82a0bf11d672 Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Sun, 23 Aug 2026 08:48:37 +0200 Subject: [PATCH] docs(SAC_LSTM_Bot): simple-words story + dictionary for readability --- SAC_LSTM_Bot/docs/campaign_notebook.md | 45 ++++++++++++++++++++++++++ 1 file changed, 45 insertions(+) diff --git a/SAC_LSTM_Bot/docs/campaign_notebook.md b/SAC_LSTM_Bot/docs/campaign_notebook.md index 81f7a82..c5effa0 100644 --- a/SAC_LSTM_Bot/docs/campaign_notebook.md +++ b/SAC_LSTM_Bot/docs/campaign_notebook.md @@ -1,3 +1,46 @@ +## The story so far, in simple words + +This project trains a robot tank. It plays many fights against other tanks. +After each fight it changes itself a little. It keeps the changes that helped it win. + +Night 1 (run 1) finished without problems. It ran for 14 hours alone. It never crashed. +It beat an old copy of itself most of the time. It beat Crazy about half the time. +Three things went badly. First, learning was not stable. Good skill appeared, then disappeared again. +Second, the saved "best" version came from one lucky perfect score. It was not really its best. +Third, the bot learned to hide and survive. It almost never shot back. + +The human approved five fixes. All five were put into the code. +Run 2 used these fixes. Its error numbers grew far too big. Learning broke. +We made one speed number smaller. This number sets how fast one part learns. +Then we dropped the broken progress and started clean. This is run 3. It is running now. + +Next we watch run 3. One of three doors will open. +Door 1: it stays steady. We let it run to the end. +Door 2: the numbers grow too big again. We turn the next speed number down. +Door 3: it stays steady but still fights badly. We teach aiming as a separate, direct lesson. + +Updated: 2026-08-23 — this section is refreshed at every major step. + +## Small dictionary + +- **training**: the time when the bot plays fights and changes itself to improve. It learns only during training. +- **battle**: one group of fights against one opponent. The bot restarts between groups. +- **round**: one single fight. Win it by destroying the enemy tank or outliving it. +- **chunk**: one work block: a battle of up to 10 rounds, then some learning from it. +- **eval (test match)**: a test match. The bot does not learn during these. We use them only to measure. +- **win rate**: how many test matches were won, as a percent. 8 wins in 10 matches = 80%. +- **checkpoint**: a saved copy of the bot's brain (a zip file). Written every few learning steps. +- **"best" checkpoint**: the saved copy we currently call best. Run 1 picked one from a lucky score, hence the quotes. +- **replay buffer**: the bot's memory of past moments: what it saw, did, and received. Learning picks old moments from it. +- **loss (critic/actor)**: a number saying how wrong the bot's inner guesses are. Lower usually means better. Losses growing huge mean trouble. +- **alpha**: a dial setting how much the bot tries new moves instead of repeating known good ones. +- **MA / composite score**: MA is the average of the last few win rates; it smooths luck. Composite is the average of MAs across all test opponents. +- **twin (SacTwin)**: a frozen copy of our own bot, used as a practice partner. Beating it proves real improvement. +- **lever**: one numbered change we prepared, waiting for approval. There are levers 1 to 5. +- **watchman**: a helper who checks the running training at set times and stops it if something breaks. + +--- + # Campaign Notebook — campaign-v1 (SAC_LSTM_Bot) Overnight training campaign on branch `research/goto-controller`. @@ -44,6 +87,8 @@ SACLSTM_SAVE_INTERVAL=5 \ ## Phase log +Every entry below starts with a plain-language first sentence. Technical detail follows for those who want it. + - **2026-08-21 23:22** — Claim posted on #56 (comment 471). Config locked, notebook committed (`2f49cb2`). - **2026-08-21 23:27** — Smoke weights archived; release build; bootstrap + probe battles vs Corners established a hidden-256 baseline checkpoint (`sac_latest.zip`, 10.5 MB) and twin seed. - **2026-08-21 23:32** — Twin regenerated (md5-verified). **Launch attempt 1** (SAVE_INTERVAL=20, TOTAL_ROUNDS=2000): ran 16+ chunks, evals every 2 chunks — but **zero checkpoints persisted** (see incident). Killed 23:48.