Main bot integration #48
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Parent
#37
What to build
Wire all modules into the working SAC_LSTM_Bot. The skeleton bot (which already connects and has radar lock + colors) gains the full RL pipeline:
The bot should fight and learn — observable through decreasing critic loss and improving match outcomes over training rounds.
Acceptance criteria
Blocked by
Settled architecture decisions (Q1–Q14) — recorded from the #48 wayfinding session (2026-08-21). This thread/channel architecture is the agreed spec for implementation.
SACLSTM_UTD_RATIO. If round produces 200 ticks, do 200 gradient steps.Tensor[float32]— lives entirely on training thread, no cross-thread issue.Tensor[float32]on bot thread. Zeros at battle start, persists across rounds. Never crosses to training thread.ChannelwithtrySend. Bot/training thread never blocks on disk I/O.canSample()already encodes minimum — no separate warm-up.trySend, drain-then-train loop on training thread.state: array[35, float32],action: array[4, float32],reward: float32,nextState: array[35, float32],done: bool. No logProb (SAC recomputes). No LSTM state (burn-in reconstructs).scannedBotId: intfrom first scan. Protocol sends no names. If ID changes across battles → buffer clears (safe default).TrainingMsg variant:
Transition(state, action, reward, nextState, done) | NewBattle(enemyId: int) | ShutdownChannel/shared-state map:
transitionChan: Channel[TrainingMsg](cap 256, bot→training) ·weightLock: Lock+ latest actor snapshot (training writes, bot reads) ·saveChan(cap 1, training→I/O) · existinggTickChan/gIntentChan/gEventChanuntouched.ORC tensor safety: No Arraymancer tensor crosses a thread boundary (PPO_Bot SIGSEGV lesson). Each thread builds its own tensors from plain arrays/seqs/floats.
Claimed for implementation (self-assigned as
SirStone— the MCP issue-edit surface here exposes no assignee field, so this comment is the assignment record). Working on branchresearch/goto-controller.Resolution — main bot integration (commit
32b71d9)All modules (#38–#47) wired into a working bot per the Q1–Q14 decisions.
What was built
src/SAC_LSTM_Bot/integration.nim(new): thread plumbing —TrainingMsgobject-variant (Transition(state/action/reward/nextState/done)|NewBattle(enemyId)|Shutdown) carried as plain fixed arrays; flat weight-snapshot pack/unpack (actor + 4 critics + alpha, layout asserts against net shapes); shared actor snapshot (Lock+ version, allocated once, written in place, copied out element-wise — zero cross-thread refcount traffic); testableTrainState/handleTrainingMsg/trainPass; training + I/O thread entries; lifecycleinitIntegration/shutdownIntegration.src/SAC_LSTM_Bot.nim(rewritten): bot thread inference loop — battle-change detection → weight pull → bullet tracking → 35-dim state → LSTM actor forward → action mapping → intents →go(); transition finalization each tick; event-driven rewards through the Welford normalizer (onBulletHit,onHitByBullet,onHitWall,onBulletHitWall); radar lock preserved; LSTM h/c persist across rounds as plain arrays on the bot object, zeroed at battle start.config.nims: added--path:src(binary now imports submodules; tests already had their own path).replay_buffer.nim: folds in the pre-existing uncommittedclear()used by the NewBattle/opponent-changed rule.Thread/channel inventory (5 threads total)
Channel[TrainingMsg]cap-256 (trySend, drops on overflow), UTD ratio viaSACLSTM_UTD_RATIO, owns SACTrainer + ReplayBuffer, publishes actor to the locked snapshot, NewBattle clears buffer only whenenemyIdchangedChannel[FullSnap], atomic zip saves via weights.nim everySACLSTM_SAVE_INTERVALsteps + one final save on ShutdownNo Arraymancer tensor crosses a thread boundary; channels/snapshot carry plain arrays/seqs only.
Tests:
nimble test— 8/8 suites green (7 existing + newtest_integration: channel round-trip incl. closed-channel semantics, NewBattle clear-only-on-opponent-change, trainPass no-op below canSample, flat snapshot pack/unpack round-trip). Additional end-to-end smoke (no server): threads spawn → real sacUpdate on hidden=32 → weight publish v2 → zip saved → clean join. Binary builds clean and starts/exits cleanly without a server.New env vars:
SACLSTM_UTD_RATIO(1),SACLSTM_BATCH_SIZE(16),SACLSTM_SAVE_INTERVAL(500),SACLSTM_WEIGHTS_PATH(/weights/sac_latest.zip).Deferred: opponent identity is the numeric
scannedBotIdfor now — name-based identification deferred to #49 (protocol exposes no names). Adam momentum is not carried in the periodic flat save (optimizer restarts fresh per process); #49's checkpoint management can switch to saveCheckpoint/loadCheckpoint end-to-end if resume quality matters.