Spec: EvoBot round-end checkpointing — best-so-far per-enemy weights #92

Closed
opened 2026-08-26 07:28:56 +02:00 by SirStone · 1 comment
Owner

Problem Statement

Training matches run for hundreds or thousands of rounds, and they get interrupted — Ctrl-C, server shutdown, dropped connections. Today the bot writes weights only when a game ends cleanly, so an interrupt at round 600 throws away every generation since the match started. On top of that, the save was last-write-wins: a worse champion could silently clobber a better one on disk. The trainer loses hours of evolution and may not even notice the regression.

Solution

Weights are checkpointed at every round end, gated by best-so-far comparison, so an interrupted match never loses more than the round in flight and the file on disk always holds the best champion seen for that enemy. Files live beside the bot binary and are never corrupted by a mid-write kill.

User Stories

  1. As a trainer running 1000-round matches, I want weights saved at each round end, so that interrupting the match loses at most one round of training.
  2. As a trainer whose bot dies early in a round, I want that round's results still saved, so that death does not skip a checkpoint.
  3. As a trainer, I want only champions better than the stored one written to disk, so that the file always holds the best-so-far, never the most recent.
  4. As a trainer, I want the stored champion's fitness kept alongside its weights, so that "is the candidate better?" survives process restarts.
  5. As a trainer with pre-existing weight files, I want legacy headerless files to load normally, so that old training data stays usable.
  6. As a trainer with legacy files, I want them treated as never-better, so that the first improved candidate upgrades them without manual migration.
  7. As a trainer hit by a kill during a save, I want the previous file intact, so that a partial write cannot corrupt my best weights.
  8. As a maintainer, I want exactly one code path that saves weights, so that persistence behavior is trivial to reason about and audit.
  9. As a trainer who moves or copies the bot folder, I want weight files resolved next to the running binary, so that every run reads and writes the right files regardless of working directory.
  10. As a trainer, I want load order to stay per-opponent → global → random init, so that existing champion-selection behavior is preserved.
  11. As a trainer fighting many different enemies, I want each enemy's best weights tracked independently, so that training one matchup never contaminates another.
  12. As a trainer, I want the global fallback weights to follow the same best-so-far rule, so that the cold-start path also only ever improves.
  13. As a trainer watching the stderr log, I want a log line whenever a new best is stored, so that I can follow training progress live.
  14. As a trainer, I want the snapshot written to be the champion already delivered between threads, so that saving never touches or stalls the evolving thread.
  15. As a maintainer, I want the checkpoint path covered by tests driving the real event handlers, so that wiring regressions fail loudly.
  16. As a trainer, I want the loss window on hard kills documented, so that I know exactly what the guarantee does not cover.

Implementation Decisions

  • Single trigger: the round-ended event handler is the ONLY place that persists weights. Death, game-end, disconnect, and OS signals are explicitly not save points; dead bots still receive the round-ended event, which covers case 2.
  • Loss ceiling: a hard kill loses at most the rounds since the last round-end boundary — accepted, no signal handlers.
  • Snapshot: the champion last drained from the evo thread's champion channel (bot-thread-visible state). The evolving thread is never inspected or blocked. Staleness bounded by the drain cadence within a round.
  • Champion transport carries fitness: the weights→bot delivery pairing is extended so each champion arrives with the GA fitness that crowned it (per-generation internal fitness, per wayfinder decision).
  • Best-gating: the candidate replaces the stored file iff its fitness strictly exceeds the stored fitness. Ties keep the incumbent.
  • File format: one metadata header line above the 745 float lines, carrying a format marker and the stored fitness. Loaders accept both formats; headerless (legacy) files parse as before and rank as never-better (−inf), upgraded by the first improved candidate.
  • Atomic write: write to a temporary file in the same directory, then rename over the target. No fsync requirement.
  • File location: the weights directory resolves at runtime to the directory of the running binary, replacing the compile-time source anchoring. The resolver is injectable so tests can redirect it.
  • Load order unchanged: per-opponent best → global best → random init, honoring ADR 0001; the per-opponent slot simply gains best-so-far meaning.
  • Modules touched: persistence module (format, gating, atomic write, location resolution), gun module (champion+fitness transport, drained-state access), main bot module (event wiring swap: save moves from game-end to round-end), test harnesses.
  • Replacement events emit one stderr log line with enemy key and new fitness.

Testing Decisions

  • Tests assert external behavior only: files on disk (name, header, float count), what a fresh load returns, and log output — never internal thread or buffer state.
  • Primary seam (the only new-capability seam needed, already existing): the compile-time-gated test export on the main bot module that lets a harness drive REAL event handlers with fabricated events. Full lifecycles are simulated there: multi-round games, bot death mid-game, name-lag → heal, kill-during-write injection (failure injected at the write call), legacy-file upgrade.
  • Secondary seam: the persistence module's existing self-check (its main-module block), extended for header round-trip, gating arithmetic (better/tie/worse/−inf legacy), and atomic-rename behavior.
  • Prior art: the handler-seam regression harness and the key-lag repro harness created during the weight-save bugfix; per-lib main-module self-checks across the project.
  • All existing harnesses must stay green after the change; the old game-end-save assertion set is updated to expect round-end semantics.

Out of Scope

  • Checkpoint history/versioning beyond the single best slot per enemy (+ global).
  • Migrating or renaming legacy files beyond read-compatibility (poisoned keys stay stale until overwritten naturally).
  • Deploy/build hygiene and the evo-thread lifecycle leak across games.
  • Validating virtual-gun simulation fidelity (recorded assumption: GA fitness is only trustworthy if vgun simulation is correct).
  • Signal-handler-based flush on Ctrl-C/SIGKILL.

Further Notes

Provenance: every implementation decision above is a locked wayfinder decision — see map "EvoBot weight checkpointing: never lose the best per-enemy weights" and its closed children (what "better" means; which moments the platform actually delivers; final trigger/write-rules lock). Platform facts backing the trigger choice (round-ended delivered per-round to every participant including dead bots; aborts deliver nothing) were verified against Tank Royale protocol schemas and server source during research.

## Problem Statement Training matches run for hundreds or thousands of rounds, and they get interrupted — Ctrl-C, server shutdown, dropped connections. Today the bot writes weights only when a game ends cleanly, so an interrupt at round 600 throws away every generation since the match started. On top of that, the save was last-write-wins: a worse champion could silently clobber a better one on disk. The trainer loses hours of evolution and may not even notice the regression. ## Solution Weights are checkpointed at every round end, gated by best-so-far comparison, so an interrupted match never loses more than the round in flight and the file on disk always holds the best champion seen for that enemy. Files live beside the bot binary and are never corrupted by a mid-write kill. ## User Stories 1. As a trainer running 1000-round matches, I want weights saved at each round end, so that interrupting the match loses at most one round of training. 2. As a trainer whose bot dies early in a round, I want that round's results still saved, so that death does not skip a checkpoint. 3. As a trainer, I want only champions better than the stored one written to disk, so that the file always holds the best-so-far, never the most recent. 4. As a trainer, I want the stored champion's fitness kept alongside its weights, so that "is the candidate better?" survives process restarts. 5. As a trainer with pre-existing weight files, I want legacy headerless files to load normally, so that old training data stays usable. 6. As a trainer with legacy files, I want them treated as never-better, so that the first improved candidate upgrades them without manual migration. 7. As a trainer hit by a kill during a save, I want the previous file intact, so that a partial write cannot corrupt my best weights. 8. As a maintainer, I want exactly one code path that saves weights, so that persistence behavior is trivial to reason about and audit. 9. As a trainer who moves or copies the bot folder, I want weight files resolved next to the running binary, so that every run reads and writes the right files regardless of working directory. 10. As a trainer, I want load order to stay per-opponent → global → random init, so that existing champion-selection behavior is preserved. 11. As a trainer fighting many different enemies, I want each enemy's best weights tracked independently, so that training one matchup never contaminates another. 12. As a trainer, I want the global fallback weights to follow the same best-so-far rule, so that the cold-start path also only ever improves. 13. As a trainer watching the stderr log, I want a log line whenever a new best is stored, so that I can follow training progress live. 14. As a trainer, I want the snapshot written to be the champion already delivered between threads, so that saving never touches or stalls the evolving thread. 15. As a maintainer, I want the checkpoint path covered by tests driving the real event handlers, so that wiring regressions fail loudly. 16. As a trainer, I want the loss window on hard kills documented, so that I know exactly what the guarantee does not cover. ## Implementation Decisions - Single trigger: the round-ended event handler is the ONLY place that persists weights. Death, game-end, disconnect, and OS signals are explicitly not save points; dead bots still receive the round-ended event, which covers case 2. - Loss ceiling: a hard kill loses at most the rounds since the last round-end boundary — accepted, no signal handlers. - Snapshot: the champion last drained from the evo thread's champion channel (bot-thread-visible state). The evolving thread is never inspected or blocked. Staleness bounded by the drain cadence within a round. - Champion transport carries fitness: the weights→bot delivery pairing is extended so each champion arrives with the GA fitness that crowned it (per-generation internal fitness, per wayfinder decision). - Best-gating: the candidate replaces the stored file iff its fitness strictly exceeds the stored fitness. Ties keep the incumbent. - File format: one metadata header line above the 745 float lines, carrying a format marker and the stored fitness. Loaders accept both formats; headerless (legacy) files parse as before and rank as never-better (−inf), upgraded by the first improved candidate. - Atomic write: write to a temporary file in the same directory, then rename over the target. No fsync requirement. - File location: the weights directory resolves at runtime to the directory of the running binary, replacing the compile-time source anchoring. The resolver is injectable so tests can redirect it. - Load order unchanged: per-opponent best → global best → random init, honoring ADR 0001; the per-opponent slot simply gains best-so-far meaning. - Modules touched: persistence module (format, gating, atomic write, location resolution), gun module (champion+fitness transport, drained-state access), main bot module (event wiring swap: save moves from game-end to round-end), test harnesses. - Replacement events emit one stderr log line with enemy key and new fitness. ## Testing Decisions - Tests assert external behavior only: files on disk (name, header, float count), what a fresh load returns, and log output — never internal thread or buffer state. - Primary seam (the only new-capability seam needed, already existing): the compile-time-gated test export on the main bot module that lets a harness drive REAL event handlers with fabricated events. Full lifecycles are simulated there: multi-round games, bot death mid-game, name-lag → heal, kill-during-write injection (failure injected at the write call), legacy-file upgrade. - Secondary seam: the persistence module's existing self-check (its main-module block), extended for header round-trip, gating arithmetic (better/tie/worse/−inf legacy), and atomic-rename behavior. - Prior art: the handler-seam regression harness and the key-lag repro harness created during the weight-save bugfix; per-lib main-module self-checks across the project. - All existing harnesses must stay green after the change; the old game-end-save assertion set is updated to expect round-end semantics. ## Out of Scope - Checkpoint history/versioning beyond the single best slot per enemy (+ global). - Migrating or renaming legacy files beyond read-compatibility (poisoned keys stay stale until overwritten naturally). - Deploy/build hygiene and the evo-thread lifecycle leak across games. - Validating virtual-gun simulation fidelity (recorded assumption: GA fitness is only trustworthy if vgun simulation is correct). - Signal-handler-based flush on Ctrl-C/SIGKILL. ## Further Notes Provenance: every implementation decision above is a locked wayfinder decision — see map "EvoBot weight checkpointing: never lose the best per-enemy weights" and its closed children (what "better" means; which moments the platform actually delivers; final trigger/write-rules lock). Platform facts backing the trigger choice (round-ended delivered per-round to every participant including dead bots; aborts deliver nothing) were verified against Tank Royale protocol schemas and server source during research.
SirStone added the ready-for-agent label 2026-08-26 07:28:56 +02:00
Author
Owner

Superseded — dropping per-enemy weights, pivoting to global-only weights + improved GA/network.

Superseded — dropping per-enemy weights, pivoting to global-only weights + improved GA/network.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#92