docs(SAC_LSTM_Bot): simple-words story + dictionary for readability
This commit is contained in:
@@ -1,3 +1,46 @@
|
||||
## The story so far, in simple words
|
||||
|
||||
This project trains a robot tank. It plays many fights against other tanks.
|
||||
After each fight it changes itself a little. It keeps the changes that helped it win.
|
||||
|
||||
Night 1 (run 1) finished without problems. It ran for 14 hours alone. It never crashed.
|
||||
It beat an old copy of itself most of the time. It beat Crazy about half the time.
|
||||
Three things went badly. First, learning was not stable. Good skill appeared, then disappeared again.
|
||||
Second, the saved "best" version came from one lucky perfect score. It was not really its best.
|
||||
Third, the bot learned to hide and survive. It almost never shot back.
|
||||
|
||||
The human approved five fixes. All five were put into the code.
|
||||
Run 2 used these fixes. Its error numbers grew far too big. Learning broke.
|
||||
We made one speed number smaller. This number sets how fast one part learns.
|
||||
Then we dropped the broken progress and started clean. This is run 3. It is running now.
|
||||
|
||||
Next we watch run 3. One of three doors will open.
|
||||
Door 1: it stays steady. We let it run to the end.
|
||||
Door 2: the numbers grow too big again. We turn the next speed number down.
|
||||
Door 3: it stays steady but still fights badly. We teach aiming as a separate, direct lesson.
|
||||
|
||||
Updated: 2026-08-23 — this section is refreshed at every major step.
|
||||
|
||||
## Small dictionary
|
||||
|
||||
- **training**: the time when the bot plays fights and changes itself to improve. It learns only during training.
|
||||
- **battle**: one group of fights against one opponent. The bot restarts between groups.
|
||||
- **round**: one single fight. Win it by destroying the enemy tank or outliving it.
|
||||
- **chunk**: one work block: a battle of up to 10 rounds, then some learning from it.
|
||||
- **eval (test match)**: a test match. The bot does not learn during these. We use them only to measure.
|
||||
- **win rate**: how many test matches were won, as a percent. 8 wins in 10 matches = 80%.
|
||||
- **checkpoint**: a saved copy of the bot's brain (a zip file). Written every few learning steps.
|
||||
- **"best" checkpoint**: the saved copy we currently call best. Run 1 picked one from a lucky score, hence the quotes.
|
||||
- **replay buffer**: the bot's memory of past moments: what it saw, did, and received. Learning picks old moments from it.
|
||||
- **loss (critic/actor)**: a number saying how wrong the bot's inner guesses are. Lower usually means better. Losses growing huge mean trouble.
|
||||
- **alpha**: a dial setting how much the bot tries new moves instead of repeating known good ones.
|
||||
- **MA / composite score**: MA is the average of the last few win rates; it smooths luck. Composite is the average of MAs across all test opponents.
|
||||
- **twin (SacTwin)**: a frozen copy of our own bot, used as a practice partner. Beating it proves real improvement.
|
||||
- **lever**: one numbered change we prepared, waiting for approval. There are levers 1 to 5.
|
||||
- **watchman**: a helper who checks the running training at set times and stops it if something breaks.
|
||||
|
||||
---
|
||||
|
||||
# Campaign Notebook — campaign-v1 (SAC_LSTM_Bot)
|
||||
|
||||
Overnight training campaign on branch `research/goto-controller`.
|
||||
@@ -44,6 +87,8 @@ SACLSTM_SAVE_INTERVAL=5 \
|
||||
|
||||
## Phase log
|
||||
|
||||
Every entry below starts with a plain-language first sentence. Technical detail follows for those who want it.
|
||||
|
||||
- **2026-08-21 23:22** — Claim posted on #56 (comment 471). Config locked, notebook committed (`2f49cb2`).
|
||||
- **2026-08-21 23:27** — Smoke weights archived; release build; bootstrap + probe battles vs Corners established a hidden-256 baseline checkpoint (`sac_latest.zip`, 10.5 MB) and twin seed.
|
||||
- **2026-08-21 23:32** — Twin regenerated (md5-verified). **Launch attempt 1** (SAVE_INTERVAL=20, TOTAL_ROUNDS=2000): ran 16+ chunks, evals every 2 chunks — but **zero checkpoints persisted** (see incident). Killed 23:48.
|
||||
|
||||
Reference in New Issue
Block a user