docs(SAC_LSTM_Bot): simple-words story + dictionary for readability

This commit is contained in:
2026-08-23 08:48:37 +02:00
parent 40e074e5a3
commit 781e41595e
+45
View File
@@ -1,3 +1,46 @@
## The story so far, in simple words
This project trains a robot tank. It plays many fights against other tanks.
After each fight it changes itself a little. It keeps the changes that helped it win.
Night 1 (run 1) finished without problems. It ran for 14 hours alone. It never crashed.
It beat an old copy of itself most of the time. It beat Crazy about half the time.
Three things went badly. First, learning was not stable. Good skill appeared, then disappeared again.
Second, the saved "best" version came from one lucky perfect score. It was not really its best.
Third, the bot learned to hide and survive. It almost never shot back.
The human approved five fixes. All five were put into the code.
Run 2 used these fixes. Its error numbers grew far too big. Learning broke.
We made one speed number smaller. This number sets how fast one part learns.
Then we dropped the broken progress and started clean. This is run 3. It is running now.
Next we watch run 3. One of three doors will open.
Door 1: it stays steady. We let it run to the end.
Door 2: the numbers grow too big again. We turn the next speed number down.
Door 3: it stays steady but still fights badly. We teach aiming as a separate, direct lesson.
Updated: 2026-08-23 — this section is refreshed at every major step.
## Small dictionary
- **training**: the time when the bot plays fights and changes itself to improve. It learns only during training.
- **battle**: one group of fights against one opponent. The bot restarts between groups.
- **round**: one single fight. Win it by destroying the enemy tank or outliving it.
- **chunk**: one work block: a battle of up to 10 rounds, then some learning from it.
- **eval (test match)**: a test match. The bot does not learn during these. We use them only to measure.
- **win rate**: how many test matches were won, as a percent. 8 wins in 10 matches = 80%.
- **checkpoint**: a saved copy of the bot's brain (a zip file). Written every few learning steps.
- **"best" checkpoint**: the saved copy we currently call best. Run 1 picked one from a lucky score, hence the quotes.
- **replay buffer**: the bot's memory of past moments: what it saw, did, and received. Learning picks old moments from it.
- **loss (critic/actor)**: a number saying how wrong the bot's inner guesses are. Lower usually means better. Losses growing huge mean trouble.
- **alpha**: a dial setting how much the bot tries new moves instead of repeating known good ones.
- **MA / composite score**: MA is the average of the last few win rates; it smooths luck. Composite is the average of MAs across all test opponents.
- **twin (SacTwin)**: a frozen copy of our own bot, used as a practice partner. Beating it proves real improvement.
- **lever**: one numbered change we prepared, waiting for approval. There are levers 1 to 5.
- **watchman**: a helper who checks the running training at set times and stops it if something breaks.
---
# Campaign Notebook — campaign-v1 (SAC_LSTM_Bot)
Overnight training campaign on branch `research/goto-controller`.
@@ -44,6 +87,8 @@ SACLSTM_SAVE_INTERVAL=5 \
## Phase log
Every entry below starts with a plain-language first sentence. Technical detail follows for those who want it.
- **2026-08-21 23:22** — Claim posted on #56 (comment 471). Config locked, notebook committed (`2f49cb2`).
- **2026-08-21 23:27** — Smoke weights archived; release build; bootstrap + probe battles vs Corners established a hidden-256 baseline checkpoint (`sac_latest.zip`, 10.5 MB) and twin seed.
- **2026-08-21 23:32** — Twin regenerated (md5-verified). **Launch attempt 1** (SAVE_INTERVAL=20, TOTAL_ROUNDS=2000): ran 16+ chunks, evals every 2 chunks — but **zero checkpoints persisted** (see incident). Killed 23:48.