spec: SAC_LSTM_Bot — Recurrent SAC with LSTM for Robocode TankRoyale #37

Closed
opened 2026-08-20 23:05:47 +02:00 by SirStone · 1 comment
Owner

Problem Statement

The current RL bot (PPO_Bot) has accumulated architectural debt and design flaws. Rather than patching it further, a new bot is needed — built from scratch with a cleaner architecture, proper nimble project structure, and a more sample-efficient off-policy algorithm (SAC) combined with recurrent memory (LSTM) for long-term enemy tactic recognition in partially observable 1v1 combat.

Solution

Build SAC_LSTM_Bot: a native Nim bot for Robocode TankRoyale using Recurrent Soft Actor-Critic (SAC-v2) with LSTM cells. The bot connects directly to the TankRoyale WebSocket server, runs inference in under 2ms per tick, and trains in a background thread from a sequential replay buffer. The LSTM enables the bot to build a persistent model of the enemy's tactics across rounds within a battle.

User Stories

  1. As a developer, I want a proper nimble project structure with src/ convention, so that the codebase follows community standards and is easy to navigate.
  2. As a developer, I want each module (state, actions, network, training, replay buffer, rewards, radar lock, weights) testable in isolation, so that bugs are caught early and localized.
  3. As a developer, I want the bot to connect to TankRoyale via the existing tankroyale_botapi in libs/, so that no middleware or extra socket layers are needed.
  4. As a developer, I want LSTM hidden state to persist across rounds within a battle and reset only at battle start, so that the bot accumulates tactical knowledge about the enemy over multiple rounds.
  5. As a developer, I want a standalone radar_lock module in libs/ that any bot can import, so that radar lock-on logic is reusable without bot-specific coupling.
  6. As a developer, I want the radar to react to onScannedBot in the same tick (not the tick after), so that scan data is never stale by one turn.
  7. As a developer, I want 4-dimensional raw continuous actions (body turn, acceleration, gun turn, fire), so that the bot can learn advanced movement patterns like wave surfing that controller-based actions prevent.
  8. As a developer, I want body turn rate mapping to be speed-aware (maxTurn = 10 - 0.75 * abs(speed)), so that the network's full output range always maps to achievable turn rates.
  9. As a developer, I want acceleration control (not target speed) mapped to [-2, +1], so that the bot has per-tick control over speed changes for precise evasive maneuvers.
  10. As a developer, I want a single fire action dim where negative means don't fire and positive maps to fire power [0.1, 3.0], so that the action space stays minimal.
  11. As a developer, I want the state vector (~35 dims) to include bullet tracking and scan staleness but exclude explicit enemy history, so that the LSTM learns temporal patterns rather than receiving them pre-computed.
  12. As a developer, I want dual Q-critics with independent LSTM cells (standard SAC-v2 twin critics), so that Q-value overestimation is suppressed.
  13. As a developer, I want automatic entropy temperature (alpha) tuning with a configurable target entropy (default: -4), so that exploration adapts during training without manual tuning.
  14. As a developer, I want a sequential replay buffer with configurable capacity (default: 500K transitions) that samples contiguous sequences and respects battle boundaries as episode terminators, so that LSTM training gets coherent temporal data.
  15. As a developer, I want configurable burn-in (default: 8 steps) and training window (default: 16 steps) within sampled sequences, so that LSTM hidden states are warmed up before computing loss.
  16. As a developer, I want running reward normalization, so that reward scale stays compatible with SAC's entropy term (O(1)).
  17. As a developer, I want weights saved as zipped .npy files via std/zipfiles, so that all tensors are in a single file but individually inspectable.
  18. As a developer, I want the training harness to reuse RunTraining.java for match orchestration with a new SAC-specific shell script, so that server lifecycle and opponent management are not reimplemented.
  19. As a developer, I want all hyperparameters (hidden size, buffer capacity, burn-in length, training window, target entropy, learning rates, tau) exposed as environment variables, so that tuning doesn't require recompilation.
  20. As a developer, I want a deterministic evaluation mode (no sampling noise), so that training progress can be measured reliably.
  21. As a developer, I want checkpoint management (periodic saves, best-of tracking by evaluation score), so that training progress is never lost.

Implementation Decisions

Algorithm: SAC-v2 with LSTM

  • Soft Actor-Critic v2 with automatic alpha tuning (no separate V-network)
  • LSTM cells (not GRU) — chosen for the separate cell state that protects long-term memory from short-term interference
  • 3 independent LSTMs: actor, critic1, critic2 — no shared encoder, no gradient leakage
  • Target networks with soft update (tau=0.005, configurable). CrossQ (BatchNorm critics, no target nets) documented as a known upgrade path but not implemented initially
  • Target entropy default: -4 (for 4-dim action space, configurable)

State Vector (~35 dimensions)

  • Own bot: position, direction, speed, energy, gun direction, gun heat
  • Enemy current: position, direction, speed, energy, fired flag, last fire power
  • Derived: enemy acceleration, enemy turn rate, relative bearing, distance
  • Wall distances: 4 cardinal directions
  • Bullet tracking: up to 3 in-flight bullets × 4 features (relative position, speed, time to impact)
  • Scan staleness: ticks since last scan, clamped
  • No explicit enemy history window — the LSTM learns temporal patterns

Action Space (4 continuous dimensions)

  • a[0]: Body turn rate — mapped via tanh * (10 - 0.75 * abs(currentSpeed)) (speed-aware)
  • a[1]: Acceleration — mapped to [-2, +1] (asymmetric: decel faster than accel, matching physics)
  • a[2]: Gun turn rate — mapped to [-20, +20] degrees/tick
  • a[3]: Fire — negative = don't fire, positive = fire power mapped to [0.1, 3.0]

Network Architecture

  • Hidden size: 256 (configurable)
  • Actor: Linear(state_dim → 256) + ReLU → LSTM(256, 256) → Linear(256 → 128) + ReLU → two heads: μ (4) and log_σ (4, clamped [-5, 2])
  • Critic (×2): Linear(state_dim + action_dim → 256) + ReLU → LSTM(256, 256) → Linear(256 → 128) + ReLU → Linear(128 → 1)
  • Squashed Gaussian policy with reparameterization trick and tanh correction
  • Weight init: He-style for hidden layers, smaller scale for output layers

LSTM Hidden State Management

  • Reset h and c to zeros at battle start only (not per-round)
  • Hidden state persists across rounds within a battle — enables cross-round tactic recognition
  • Replay buffer marks battle boundaries as episode terminators; round boundaries are not episode breaks

Replay Buffer

  • Ring buffer in memory, configurable capacity (default 500K transitions)
  • Stores: state, action, reward, next_state, done flag, plus LSTM hidden states at sequence start
  • Sequence sampling: contiguous windows of length burn_in + train_window (default 24)
  • Sequences never cross battle boundaries
  • Burn-in (default 8 steps): forward pass only, updates hidden state, no loss computation
  • Train window (default 16 steps): full loss computation and backpropagation

Reward Function

  • Damage inflicted: +(4p + 2(p-1)) where p = fire power
  • Damage received: -(4p_enemy + 2(p_enemy - 1))
  • Wall hit: -5.0 per tick of impact
  • Wasted shot: -0.1 × p
  • Win/loss: +20.0 / -10.0
  • Running mean/std normalization applied before feeding to SAC

Radar

  • Standalone radar_lock module in libs/radar_lock/ — reusable across bots
  • 1v1 narrow lock-on: computes radarTurnRate to keep the enemy in scan arc
  • API: init() for defaults, doRadar(radarHeading, enemyBearing, ...) → float
  • Activated inside onScannedBot handler (same-tick reactivity)
  • adjustRadarForBodyTurn=true, adjustRadarForGunTurn=true

Project Structure (nimble standard)

SAC_LSTM_Bot/
  SAC_LSTM_Bot.nimble
  src/
    SAC_LSTM_Bot.nim              # main: WebSocket, event handlers, run loop
    SAC_LSTM_Bot/
      state.nim                   # state vector construction
      actions.nim                 # action mapping (speed-aware, accel, fire)
      network.nim                 # LSTM actor + dual critics
      training.nim                # SAC loss, gradient updates, soft target sync
      replay_buffer.nim           # sequential ring buffer
      rewards.nim                 # reward computation + running normalization
      weights.nim                 # zipped .npy save/load
  tests/
    config.nims
    test_state.nim
    test_actions.nim
    test_network.nim
    test_training.nim
    test_replay_buffer.nim
    test_rewards.nim
    test_weights.nim

libs/
  radar_lock/
    radar_lock.nimble
    radar_lock.nim                # standalone, bot-agnostic
    tests/
      test_radar_lock.nim

Weight Persistence

  • All network tensors (actor, 2 critics, 2 targets, alpha, Adam optimizer states) saved as individual .npy tensors inside a single .zip file via std/zipfiles
  • Atomic save: write to temp file, then rename
  • Checkpoint directory: weights/latest/ for current, periodic snapshots

Training Infrastructure

  • Reuse RunTraining.java for match orchestration (server lifecycle, opponent management, liveness detection)
  • New sac_train.sh shell script for SAC-specific training flow
  • Two-thread architecture: WebSocket client thread (inference) + background trainer thread (SAC updates)
  • All hyperparameters via environment variables

Configuration (environment variables)

  • SACLSTM_HIDDEN_SIZE (default: 256)
  • SACLSTM_BUFFER_CAPACITY (default: 500000)
  • SACLSTM_BURN_IN (default: 8)
  • SACLSTM_TRAIN_WINDOW (default: 16)
  • SACLSTM_TARGET_ENTROPY (default: -4.0)
  • SACLSTM_TAU (default: 0.005)
  • SACLSTM_LR_ACTOR, SACLSTM_LR_CRITIC, SACLSTM_LR_ALPHA
  • SACLSTM_EVAL_MODE (0/1)

Testing Decisions

What makes a good test

Tests verify external behavior through the module's public API — given inputs, assert outputs. No testing of internal data structures, private helpers, or implementation details. If refactoring internals doesn't break the public contract, no test should fail.

Module test seams (8 total, one per module)

  1. test_state.nim — buildState() produces correct tensor shape, values within normalization ranges, handles missing scan data gracefully
  2. test_actions.nim — mapActions() applies speed-aware turn mapping correctly, clamps acceleration to [-2, +1], fire threshold behavior at boundary values
  3. test_network.nim — actor/critic forward pass produces correct output shapes, hidden state dimensions propagate correctly, deterministic mode suppresses sampling noise
  4. test_training.nim — SAC update produces finite losses, soft target update blends weights correctly at configured tau, alpha stays positive
  5. test_replay_buffer.nim — sequences never cross battle boundaries, burn-in/train split is correct, buffer wraps correctly at capacity, empty buffer doesn't crash
  6. test_rewards.nim — reward values match the documented formula for each event type, running normalization converges to zero-mean unit-variance
  7. test_radar_lock.nim — lock produces correct turn rate for cardinal bearings, handles angle wrapping (359° → 1°), clamps to ±45°
  8. test_weights.nim — save then load round-trips all tensors exactly, file is valid zip containing .npy entries

Test tooling

  • stdlib unittest only — no frameworks
  • nimble test discovers and runs all tests/test_*.nim
  • tests/config.nims sets --path:"../src" and --path:"../../libs"

Out of Scope

  • Frame stacking / attention alternatives to LSTM — future experiments, not this bot
  • CrossQ / RedQ / DroQ — documented as upgrade paths, not in initial implementation
  • Multi-bot support (melee) — this bot is 1v1 only; radar lock assumes single enemy
  • Radar as a learned action — radar is a fixed lock-on module, not a network output
  • goto/aimTo controllers — deliberately excluded to enable raw movement learning (wave surfing)
  • GPU acceleration — Arraymancer CPU only for now; GPU is a future optimization if training is too slow
  • Python/Rust implementation — Nim only
  • PPO_Bot modifications — PPO_Bot is paused, this spec does not touch it

Further Notes

  • Known upgrade path: CrossQ — if sample efficiency becomes a bottleneck, add BatchNorm to critic layers and remove target networks. Zero inference cost, only affects training. Reference: CrossQ (ICLR 2024).
  • Known upgrade path: larger hidden size — 256 is the baseline. 512 is feasible (~2ms inference) if 256 proves insufficient for complex tactics.
  • Reward scale sensitivity — SAC is notoriously sensitive to reward scale. Running normalization is the first defense; if training is unstable, revisit the raw reward magnitudes.
  • The libs/radar_lock/ module is the first shared library in this repo. It sets a precedent for extracting reusable bot components.
  • LSTM vs GRU was decided: LSTM's separate cell state protects long-term tactical memory from short-term interference. The extra compute cost (~0.3ms) is acceptable within the 2ms inference budget.
## Problem Statement The current RL bot (PPO_Bot) has accumulated architectural debt and design flaws. Rather than patching it further, a new bot is needed — built from scratch with a cleaner architecture, proper nimble project structure, and a more sample-efficient off-policy algorithm (SAC) combined with recurrent memory (LSTM) for long-term enemy tactic recognition in partially observable 1v1 combat. ## Solution Build **SAC_LSTM_Bot**: a native Nim bot for Robocode TankRoyale using Recurrent Soft Actor-Critic (SAC-v2) with LSTM cells. The bot connects directly to the TankRoyale WebSocket server, runs inference in under 2ms per tick, and trains in a background thread from a sequential replay buffer. The LSTM enables the bot to build a persistent model of the enemy's tactics across rounds within a battle. ## User Stories 1. As a developer, I want a proper nimble project structure with `src/` convention, so that the codebase follows community standards and is easy to navigate. 2. As a developer, I want each module (state, actions, network, training, replay buffer, rewards, radar lock, weights) testable in isolation, so that bugs are caught early and localized. 3. As a developer, I want the bot to connect to TankRoyale via the existing `tankroyale_botapi` in `libs/`, so that no middleware or extra socket layers are needed. 4. As a developer, I want LSTM hidden state to persist across rounds within a battle and reset only at battle start, so that the bot accumulates tactical knowledge about the enemy over multiple rounds. 5. As a developer, I want a standalone `radar_lock` module in `libs/` that any bot can import, so that radar lock-on logic is reusable without bot-specific coupling. 6. As a developer, I want the radar to react to `onScannedBot` in the same tick (not the tick after), so that scan data is never stale by one turn. 7. As a developer, I want 4-dimensional raw continuous actions (body turn, acceleration, gun turn, fire), so that the bot can learn advanced movement patterns like wave surfing that controller-based actions prevent. 8. As a developer, I want body turn rate mapping to be speed-aware (`maxTurn = 10 - 0.75 * abs(speed)`), so that the network's full output range always maps to achievable turn rates. 9. As a developer, I want acceleration control (not target speed) mapped to [-2, +1], so that the bot has per-tick control over speed changes for precise evasive maneuvers. 10. As a developer, I want a single fire action dim where negative means don't fire and positive maps to fire power [0.1, 3.0], so that the action space stays minimal. 11. As a developer, I want the state vector (~35 dims) to include bullet tracking and scan staleness but exclude explicit enemy history, so that the LSTM learns temporal patterns rather than receiving them pre-computed. 12. As a developer, I want dual Q-critics with independent LSTM cells (standard SAC-v2 twin critics), so that Q-value overestimation is suppressed. 13. As a developer, I want automatic entropy temperature (alpha) tuning with a configurable target entropy (default: -4), so that exploration adapts during training without manual tuning. 14. As a developer, I want a sequential replay buffer with configurable capacity (default: 500K transitions) that samples contiguous sequences and respects battle boundaries as episode terminators, so that LSTM training gets coherent temporal data. 15. As a developer, I want configurable burn-in (default: 8 steps) and training window (default: 16 steps) within sampled sequences, so that LSTM hidden states are warmed up before computing loss. 16. As a developer, I want running reward normalization, so that reward scale stays compatible with SAC's entropy term (O(1)). 17. As a developer, I want weights saved as zipped `.npy` files via `std/zipfiles`, so that all tensors are in a single file but individually inspectable. 18. As a developer, I want the training harness to reuse `RunTraining.java` for match orchestration with a new SAC-specific shell script, so that server lifecycle and opponent management are not reimplemented. 19. As a developer, I want all hyperparameters (hidden size, buffer capacity, burn-in length, training window, target entropy, learning rates, tau) exposed as environment variables, so that tuning doesn't require recompilation. 20. As a developer, I want a deterministic evaluation mode (no sampling noise), so that training progress can be measured reliably. 21. As a developer, I want checkpoint management (periodic saves, best-of tracking by evaluation score), so that training progress is never lost. ## Implementation Decisions ### Algorithm: SAC-v2 with LSTM - Soft Actor-Critic v2 with automatic alpha tuning (no separate V-network) - LSTM cells (not GRU) — chosen for the separate cell state that protects long-term memory from short-term interference - 3 independent LSTMs: actor, critic1, critic2 — no shared encoder, no gradient leakage - Target networks with soft update (tau=0.005, configurable). CrossQ (BatchNorm critics, no target nets) documented as a known upgrade path but not implemented initially - Target entropy default: -4 (for 4-dim action space, configurable) ### State Vector (~35 dimensions) - Own bot: position, direction, speed, energy, gun direction, gun heat - Enemy current: position, direction, speed, energy, fired flag, last fire power - Derived: enemy acceleration, enemy turn rate, relative bearing, distance - Wall distances: 4 cardinal directions - Bullet tracking: up to 3 in-flight bullets × 4 features (relative position, speed, time to impact) - Scan staleness: ticks since last scan, clamped - No explicit enemy history window — the LSTM learns temporal patterns ### Action Space (4 continuous dimensions) - `a[0]`: Body turn rate — mapped via `tanh * (10 - 0.75 * abs(currentSpeed))` (speed-aware) - `a[1]`: Acceleration — mapped to `[-2, +1]` (asymmetric: decel faster than accel, matching physics) - `a[2]`: Gun turn rate — mapped to `[-20, +20]` degrees/tick - `a[3]`: Fire — negative = don't fire, positive = fire power mapped to `[0.1, 3.0]` ### Network Architecture - Hidden size: 256 (configurable) - Actor: Linear(state_dim → 256) + ReLU → LSTM(256, 256) → Linear(256 → 128) + ReLU → two heads: μ (4) and log_σ (4, clamped [-5, 2]) - Critic (×2): Linear(state_dim + action_dim → 256) + ReLU → LSTM(256, 256) → Linear(256 → 128) + ReLU → Linear(128 → 1) - Squashed Gaussian policy with reparameterization trick and tanh correction - Weight init: He-style for hidden layers, smaller scale for output layers ### LSTM Hidden State Management - Reset h and c to zeros at **battle start** only (not per-round) - Hidden state persists across rounds within a battle — enables cross-round tactic recognition - Replay buffer marks **battle boundaries** as episode terminators; round boundaries are not episode breaks ### Replay Buffer - Ring buffer in memory, configurable capacity (default 500K transitions) - Stores: state, action, reward, next_state, done flag, plus LSTM hidden states at sequence start - Sequence sampling: contiguous windows of length burn_in + train_window (default 24) - Sequences never cross battle boundaries - Burn-in (default 8 steps): forward pass only, updates hidden state, no loss computation - Train window (default 16 steps): full loss computation and backpropagation ### Reward Function - Damage inflicted: `+(4p + 2(p-1))` where p = fire power - Damage received: `-(4p_enemy + 2(p_enemy - 1))` - Wall hit: `-5.0` per tick of impact - Wasted shot: `-0.1 × p` - Win/loss: `+20.0` / `-10.0` - Running mean/std normalization applied before feeding to SAC ### Radar - Standalone `radar_lock` module in `libs/radar_lock/` — reusable across bots - 1v1 narrow lock-on: computes `radarTurnRate` to keep the enemy in scan arc - API: `init()` for defaults, `doRadar(radarHeading, enemyBearing, ...) → float` - Activated inside `onScannedBot` handler (same-tick reactivity) - `adjustRadarForBodyTurn=true`, `adjustRadarForGunTurn=true` ### Project Structure (nimble standard) ``` SAC_LSTM_Bot/ SAC_LSTM_Bot.nimble src/ SAC_LSTM_Bot.nim # main: WebSocket, event handlers, run loop SAC_LSTM_Bot/ state.nim # state vector construction actions.nim # action mapping (speed-aware, accel, fire) network.nim # LSTM actor + dual critics training.nim # SAC loss, gradient updates, soft target sync replay_buffer.nim # sequential ring buffer rewards.nim # reward computation + running normalization weights.nim # zipped .npy save/load tests/ config.nims test_state.nim test_actions.nim test_network.nim test_training.nim test_replay_buffer.nim test_rewards.nim test_weights.nim libs/ radar_lock/ radar_lock.nimble radar_lock.nim # standalone, bot-agnostic tests/ test_radar_lock.nim ``` ### Weight Persistence - All network tensors (actor, 2 critics, 2 targets, alpha, Adam optimizer states) saved as individual `.npy` tensors inside a single `.zip` file via `std/zipfiles` - Atomic save: write to temp file, then rename - Checkpoint directory: `weights/latest/` for current, periodic snapshots ### Training Infrastructure - Reuse `RunTraining.java` for match orchestration (server lifecycle, opponent management, liveness detection) - New `sac_train.sh` shell script for SAC-specific training flow - Two-thread architecture: WebSocket client thread (inference) + background trainer thread (SAC updates) - All hyperparameters via environment variables ### Configuration (environment variables) - `SACLSTM_HIDDEN_SIZE` (default: 256) - `SACLSTM_BUFFER_CAPACITY` (default: 500000) - `SACLSTM_BURN_IN` (default: 8) - `SACLSTM_TRAIN_WINDOW` (default: 16) - `SACLSTM_TARGET_ENTROPY` (default: -4.0) - `SACLSTM_TAU` (default: 0.005) - `SACLSTM_LR_ACTOR`, `SACLSTM_LR_CRITIC`, `SACLSTM_LR_ALPHA` - `SACLSTM_EVAL_MODE` (0/1) ## Testing Decisions ### What makes a good test Tests verify **external behavior through the module's public API** — given inputs, assert outputs. No testing of internal data structures, private helpers, or implementation details. If refactoring internals doesn't break the public contract, no test should fail. ### Module test seams (8 total, one per module) 1. **`test_state.nim`** — `buildState()` produces correct tensor shape, values within normalization ranges, handles missing scan data gracefully 2. **`test_actions.nim`** — `mapActions()` applies speed-aware turn mapping correctly, clamps acceleration to [-2, +1], fire threshold behavior at boundary values 3. **`test_network.nim`** — actor/critic forward pass produces correct output shapes, hidden state dimensions propagate correctly, deterministic mode suppresses sampling noise 4. **`test_training.nim`** — SAC update produces finite losses, soft target update blends weights correctly at configured tau, alpha stays positive 5. **`test_replay_buffer.nim`** — sequences never cross battle boundaries, burn-in/train split is correct, buffer wraps correctly at capacity, empty buffer doesn't crash 6. **`test_rewards.nim`** — reward values match the documented formula for each event type, running normalization converges to zero-mean unit-variance 7. **`test_radar_lock.nim`** — lock produces correct turn rate for cardinal bearings, handles angle wrapping (359° → 1°), clamps to ±45° 8. **`test_weights.nim`** — save then load round-trips all tensors exactly, file is valid zip containing `.npy` entries ### Test tooling - stdlib `unittest` only — no frameworks - `nimble test` discovers and runs all `tests/test_*.nim` - `tests/config.nims` sets `--path:"../src"` and `--path:"../../libs"` ## Out of Scope - **Frame stacking / attention alternatives** to LSTM — future experiments, not this bot - **CrossQ / RedQ / DroQ** — documented as upgrade paths, not in initial implementation - **Multi-bot support (melee)** — this bot is 1v1 only; radar lock assumes single enemy - **Radar as a learned action** — radar is a fixed lock-on module, not a network output - **goto/aimTo controllers** — deliberately excluded to enable raw movement learning (wave surfing) - **GPU acceleration** — Arraymancer CPU only for now; GPU is a future optimization if training is too slow - **Python/Rust implementation** — Nim only - **PPO_Bot modifications** — PPO_Bot is paused, this spec does not touch it ## Further Notes - **Known upgrade path: CrossQ** — if sample efficiency becomes a bottleneck, add BatchNorm to critic layers and remove target networks. Zero inference cost, only affects training. Reference: CrossQ (ICLR 2024). - **Known upgrade path: larger hidden size** — 256 is the baseline. 512 is feasible (~2ms inference) if 256 proves insufficient for complex tactics. - **Reward scale sensitivity** — SAC is notoriously sensitive to reward scale. Running normalization is the first defense; if training is unstable, revisit the raw reward magnitudes. - **The `libs/radar_lock/` module** is the first shared library in this repo. It sets a precedent for extracting reusable bot components. - **LSTM vs GRU was decided**: LSTM's separate cell state protects long-term tactical memory from short-term interference. The extra compute cost (~0.3ms) is acceptable within the 2ms inference budget.
SirStone added the ready-for-agent label 2026-08-20 23:05:47 +02:00
Author
Owner

Delivered. All child tickets #38–#49 implemented and closed:

  • Modules: #38–#47
  • Main bot integration: #48 (commit 32b71d9; architecture decisions in comment id=437 on #48)
  • Training harness: #49 (commit df256b4)

Also: vendored botapi updated to v1.0.1 content (commit 7104645), enabling name-based opponent identification.

End-to-end training loop verified live via a sac_train.sh smoke run. Closing as delivered.

Delivered. All child tickets #38–#49 implemented and closed: - Modules: #38–#47 - Main bot integration: #48 (commit 32b71d9; architecture decisions in comment id=437 on #48) - Training harness: #49 (commit df256b4) Also: vendored `botapi` updated to v1.0.1 content (commit 7104645), enabling name-based opponent identification. End-to-end training loop verified live via a `sac_train.sh` smoke run. Closing as delivered.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#37