spec: SAC_LSTM_Bot — Recurrent SAC with LSTM for Robocode TankRoyale #37
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem Statement
The current RL bot (PPO_Bot) has accumulated architectural debt and design flaws. Rather than patching it further, a new bot is needed — built from scratch with a cleaner architecture, proper nimble project structure, and a more sample-efficient off-policy algorithm (SAC) combined with recurrent memory (LSTM) for long-term enemy tactic recognition in partially observable 1v1 combat.
Solution
Build SAC_LSTM_Bot: a native Nim bot for Robocode TankRoyale using Recurrent Soft Actor-Critic (SAC-v2) with LSTM cells. The bot connects directly to the TankRoyale WebSocket server, runs inference in under 2ms per tick, and trains in a background thread from a sequential replay buffer. The LSTM enables the bot to build a persistent model of the enemy's tactics across rounds within a battle.
User Stories
src/convention, so that the codebase follows community standards and is easy to navigate.tankroyale_botapiinlibs/, so that no middleware or extra socket layers are needed.radar_lockmodule inlibs/that any bot can import, so that radar lock-on logic is reusable without bot-specific coupling.onScannedBotin the same tick (not the tick after), so that scan data is never stale by one turn.maxTurn = 10 - 0.75 * abs(speed)), so that the network's full output range always maps to achievable turn rates..npyfiles viastd/zipfiles, so that all tensors are in a single file but individually inspectable.RunTraining.javafor match orchestration with a new SAC-specific shell script, so that server lifecycle and opponent management are not reimplemented.Implementation Decisions
Algorithm: SAC-v2 with LSTM
State Vector (~35 dimensions)
Action Space (4 continuous dimensions)
a[0]: Body turn rate — mapped viatanh * (10 - 0.75 * abs(currentSpeed))(speed-aware)a[1]: Acceleration — mapped to[-2, +1](asymmetric: decel faster than accel, matching physics)a[2]: Gun turn rate — mapped to[-20, +20]degrees/ticka[3]: Fire — negative = don't fire, positive = fire power mapped to[0.1, 3.0]Network Architecture
LSTM Hidden State Management
Replay Buffer
Reward Function
+(4p + 2(p-1))where p = fire power-(4p_enemy + 2(p_enemy - 1))-5.0per tick of impact-0.1 × p+20.0/-10.0Radar
radar_lockmodule inlibs/radar_lock/— reusable across botsradarTurnRateto keep the enemy in scan arcinit()for defaults,doRadar(radarHeading, enemyBearing, ...) → floatonScannedBothandler (same-tick reactivity)adjustRadarForBodyTurn=true,adjustRadarForGunTurn=trueProject Structure (nimble standard)
Weight Persistence
.npytensors inside a single.zipfile viastd/zipfilesweights/latest/for current, periodic snapshotsTraining Infrastructure
RunTraining.javafor match orchestration (server lifecycle, opponent management, liveness detection)sac_train.shshell script for SAC-specific training flowConfiguration (environment variables)
SACLSTM_HIDDEN_SIZE(default: 256)SACLSTM_BUFFER_CAPACITY(default: 500000)SACLSTM_BURN_IN(default: 8)SACLSTM_TRAIN_WINDOW(default: 16)SACLSTM_TARGET_ENTROPY(default: -4.0)SACLSTM_TAU(default: 0.005)SACLSTM_LR_ACTOR,SACLSTM_LR_CRITIC,SACLSTM_LR_ALPHASACLSTM_EVAL_MODE(0/1)Testing Decisions
What makes a good test
Tests verify external behavior through the module's public API — given inputs, assert outputs. No testing of internal data structures, private helpers, or implementation details. If refactoring internals doesn't break the public contract, no test should fail.
Module test seams (8 total, one per module)
test_state.nim—buildState()produces correct tensor shape, values within normalization ranges, handles missing scan data gracefullytest_actions.nim—mapActions()applies speed-aware turn mapping correctly, clamps acceleration to [-2, +1], fire threshold behavior at boundary valuestest_network.nim— actor/critic forward pass produces correct output shapes, hidden state dimensions propagate correctly, deterministic mode suppresses sampling noisetest_training.nim— SAC update produces finite losses, soft target update blends weights correctly at configured tau, alpha stays positivetest_replay_buffer.nim— sequences never cross battle boundaries, burn-in/train split is correct, buffer wraps correctly at capacity, empty buffer doesn't crashtest_rewards.nim— reward values match the documented formula for each event type, running normalization converges to zero-mean unit-variancetest_radar_lock.nim— lock produces correct turn rate for cardinal bearings, handles angle wrapping (359° → 1°), clamps to ±45°test_weights.nim— save then load round-trips all tensors exactly, file is valid zip containing.npyentriesTest tooling
unittestonly — no frameworksnimble testdiscovers and runs alltests/test_*.nimtests/config.nimssets--path:"../src"and--path:"../../libs"Out of Scope
Further Notes
libs/radar_lock/module is the first shared library in this repo. It sets a precedent for extracting reusable bot components.Delivered. All child tickets #38–#49 implemented and closed:
df256b4)Also: vendored
botapiupdated to v1.0.1 content (commit7104645), enabling name-based opponent identification.End-to-end training loop verified live via a
sac_train.shsmoke run. Closing as delivered.