tm-learning-tracks.md covers three things, all marked [FACT]/[INFERENCE]/ [UNKNOWN]: - Section A: the Tsetlin gun's label is measured against the wrong baseline. predX = linearX + cx, so rx = actual - predX = delta - cx, and inside tmLearnOne error = residual - predicted = (delta - cx) - cx = delta - 2cx. The fixed point is cx = delta/2 -- HALF the correction needed, even with perfect Granmo feedback. Fix: store linearX/linearY in TmTrace and train on delta. Also: hits zero the label instead of carrying their true residual, and the per-clause step is magnitude-blind. - Section B: what a TM is actually good at (AND-clauses over binary literals, readable output) and why this repo suits it -- the gun already builds an 83-bit x 10-frame Gray-coded window (870 bits). Includes a falsifiable known-rule benchmark proposal. - Section C: delayed-reward learning belongs to the MOVEMENT layer, not the gun. The gun's outcome is delayed but exactly pairable via (fireTick, powerBin), so its effective lambda is 1 and discounting would only destroy information. Also records that docs/papers/tm-deb-paper.pdf was deleted by the user as AI-generated and unverifiable, while the Granmo-based feedback diff in tm-deb-assessment.md stands on its own.
32 KiB
Tsetlin Machine learning tracks — target shape, high-level patterns, delayed reward
Date: 2026-09-20
Scope: Design notes only. No .nim file was edited, nothing was built, no battle was run, no /tmp artifact was touched. Every claim is tagged by its evidence class.
Author's method: Read common_libs/guns/tsetlin.nim, common_libs/gun_harness/gun_interface.nim, common_libs/gun_harness/virtual_bullets.nim, the common_libs/movements/* modules, ModularBot_garage/src/ModularBot.nim, BNNBot_garage/src/{binary_encoding,tsetlin_predictor,wisard_predictor}.nim, SAC_LSTM_Bot_garage/src/SAC_LSTM_Bot/rewards.nim, docs/rl-algorithm-choice.md, docs/research/tm-deb-assessment.md, and the installed robocode_tankroyale_botapi schemas.nim / bot.nim.
Evidence legend
- [FACT] — verified by reading the cited code or paper at the cited location.
- [INFERENCE] — reasoned from facts; not measured.
- [UNKNOWN] — needs an experiment; no claim is made.
Provenance caveat.
docs/papers/tm-deb-paper.pdf(TM-DEB) was deleted by the user as AI-generated and unverifiable (see the editorial note indocs/research/tm-deb-assessment.md). Nothing in this document uses it as evidence, and the Section C delayed-reward design does not adopt any TM-DEB mechanism.
SECTION A — The Tsetlin gun's target is the wrong shape (do this first)
A.0 Correction to the brief, before anything else
The brief states that "the gun currently learns a real-valued (cx,cy) pixel correction from a binary hit/miss outcome." That does not match HEAD. The current training site already computes a signed residual. [FACT] common_libs/guns/tsetlin.nim, proc onResult*:
# Directional residual: actual enemy pos minus our prediction
# On hit residual is 0 (we were right); on miss we push toward actual position.
let rx = if e.hit: 0.0 else: clamp(e.actualX - t.predX, -TM_RESID_MAX, TM_RESID_MAX)
let ry = if e.hit: 0.0 else: clamp(e.actualY - t.predY, -TM_RESID_MAX, TM_RESID_MAX)
let lits = tmMakeLiterals(t.input)
g.net.tmLearnOne(0, lits, t.cache, rx)
g.net.tmLearnOne(1, lits, t.cache, ry)
The previous revision (git show e536900~1:common_libs/guns/tsetlin.nim) did the same. So the binary hit flag is not the label anywhere in the gun's history; it is only the fitness signal in virtual_bullets.nim. The brief's concern is therefore a non-issue here — but a different and sharper target-shape defect is present, derived below from the code.
A.1 Why a hit/miss-only label would be ill-posed (stated for completeness)
If the label were only hit ∈ {0,1}:
- A hit says the total 2D prediction was good; a miss says it was bad. Neither says which axis was off, nor by how much, nor which direction. [FACT — by construction of a scalar label vs a 2D continuous target.]
- Fitting a signed 2D correction (two real outputs, each with sign and magnitude) to a scalar binary outcome would need the direction to be invented from something else (e.g. the predicted-vs-actual geometry). It would not be a well-posed supervised problem.
This is not what the gun does, per A.0.
A.2 The harness hands the gun an exact signed residual
[FACT] common_libs/gun_harness/gun_interface.nim defines FeedbackEvent with actualX, actualY, prediction, plus fireTick, powerBin, missDistance, hit. At resolution, common_libs/gun_harness/virtual_bullets.nim (tickBullets) reads the target's actual position from the enemies table and builds the event:
let missDist = hypot(bx - ex, by - ey)
let fe = FeedbackEvent(
prediction: GunPrediction(x: b.aimX, y: b.aimY),
actualX: ex, actualY: ey, ... fireTick: b.fireTick, powerBin: b.powerBin, ...)
So residual = actual − prediction is known exactly at resolution time, per axis, with no ambiguity. [FACT]
[INFERENCE] Therefore the gun's learning task is supervised regression with an exact continuous label, not reinforcement learning. There is no hidden credit to assign, so no eligibility discounting is needed — the gun's effective λ = 1.
[FACT] The gun already exploits that: onResult retrieves the exact trace by (fireTick, powerBin) via tmTraceSlot(e.fireTick, binIdx) (ring of TM_TRACE_SLOTS = 1024), checking t.fireTick == e.fireTick and t.powerBin == binIdx. trainedShots counts successful pairings, traceMisses counts failures. This is the repo-native form of an eligibility buffer, and the pairing is exact. [FACT] docs/research/tm-deb-assessment.md §5 reaches the same conclusion: effective λ = 1.
A.3 The real target-shape defect: the label is measured from the wrong baseline
This is the code-grounded finding that replaces the brief's premise.
Let, in common_libs/guns/tsetlin.nim:
L= the linear extrapolation baseline, computed inpredict()aslinearX = state.enemyX + cos(headingRad) * state.enemySpeed * ticksToArrive. [FACT]c= the TM's output correction,tmForwardWithCachereturns(vx / TM_T * TM_RESID_MAX, …). [FACT]P= stored prediction,predX = clamp(linearX + cx, …), andTmTrace.predX = predX. [FACT]A=e.actualXat resolution.
At training time rx = A − P = A − L − c. [FACT]
Inside tmLearnOne, predicted is the TM's own output recomputed from the cached clause outputs:
for c in 0..<TM_N_CLAUSES: vote += tmPolarity(c) * float(cache[outIdx * TM_N_CLAUSES + c])
vote = clamp(vote, -TM_T, TM_T)
let predicted = vote / TM_T * TM_RESID_MAX # == c
let error = residual - predicted # == rx - c
So error = (A − L − c) − c = (A − L) − 2c. [FACT — direct substitution.]
[INFERENCE] The update minimizes |error|, whose fixed point is c* = (A − L)/2 — half the correction that makes P hit A. The code passes a label that already contains one copy of the TM's output, then subtracts a second copy. Even with a perfect feedback rule (Granmo Table 2/3), the TM is being asked for a systematically shrunken correction. This is a well-posedness bug in the target definition, not a learning-rate smell.
Corroborating detail: the same baseline pattern exists in the historical BNNBot_garage/src/tsetlin_predictor.nim, proc learn ("residualX/Y: pixel correction needed (actual_target - aimed_point)") and in BNNBot_garage/src/TsetlinBot.nim (residualX = bot.lastEnemyX - bulletX, where bulletX already includes the correction). [FACT] So this is not a one-off typo introduced by the current gun; it is a repeated design slip.
Three further target-shape defects at the same site:
- Hits are trained toward zero, not toward their true (small) residual. [FACT]
rx = if e.hit: 0.0 else: …. A hit means the 2DmissDistance < BotRadius = 18.0[FACT], but each axis can still be off by up to 18 px. Zeroing both axes discards that information and biases the TM toward under-correction. A hit is not "residual exactly 0"; it is "|(rx, ry)| < 18". - The update is magnitude-blind. [FACT]
pFeedback = min(1.0, abs(error) / (2.0 * TM_RESID_MAX))only scales the probability that a clause receives feedback; the per-clause state step is always ±1 (st = min(st+1, …)/max(st-1, …)). [INFERENCE] Direction is carried bysign(error)and clause polarity, but magnitude must be inferred statistically across many shots — slow, high-variance, and unable to express "this shot needed 40 px". - Range/scale. [FACT]
residualis clamped to±TM_RESID_MAX = 80and the output vote to±TM_T = 25→±80. Because of the factor-2 defect above, the correction the TM converges to is half the baseline error: to correct a baseline errorδit must drive its output toδ/2, so it saturates at a baseline error of±160rather than the±80the constant name suggests. [INFERENCE]
A.4 Consequence: the right target is the residual — measured from the baseline
[INFERENCE] The correct supervised target for the TM's correction is δ = A − L (the leftover error of the linear baseline), and onResult should pass δ, not A − P. Equivalently: store linearX, linearY in TmTrace alongside predX, predY and pass e.actualX − t.linearX. The current TmTrace stores predX, predY but not the baseline [FACT], so this requires a small struct change (but still no change to the harness — the baseline is computable at predict time and the actual position already arrives in the event).
A.5 Learning a continuous residual with binary, gradient-free machinery
Tsetlin Machines are binary and gradient-free (a hard constraint for this project). [FACT — the automata in tsetlin.nim are discrete integer states updated by the Tables 2/3-style reward/penalty; there is no gradient.] Two honest ways to learn a continuous residual:
Option A — quantise the residual and classify (recommended).
- Split each axis's residual range
[−R, +R]intoKbins. Run two independent multi-class TMs (one per axis), each withKclasses; each class gets its own clause team and vote, predictargmax, map the winning bin back to its centre pixel value. This is Granmo's standard multi-class formulation. [INFERENCE — standard TM construction; not present in this repo's code.] - Tradeoffs:
- Resolution vs data:
Kclasses needKteams and enough per-class samples. The residual distribution is heavily concentrated near 0, so most classes are rare. [INFERENCE] Finer bins = better magnitude resolution = more data per bin = slower convergence. - Magnitude vs direction: a binned classifier captures both; a pure sign classifier (2 classes/axis) captures only direction and would still need a magnitude estimate (a fixed step, or a second head). [INFERENCE]
- The no-correction class: a hit should map to an explicit zero bin (e.g.
|residual| < BotRadius). Modelling "no correction" as its own class lets the TM emit exactly 0 when it is right, instead of always nudging. Without it, hits trained toward zero (A.3 defect 1) will drag the whole output distribution toward 0.
- Resolution vs data:
Option B — predict the SIGN per axis as an independent binary task.
- 2 classes/axis → 2 clauses-team pairs; the smallest-data option and the most robust head. [INFERENCE] Magnitude must come from elsewhere (fixed step size, or a separate regression/binned head), so this is a partial solution but a good first milestone: prove the TM can learn which way the correction goes before asking it to learn how far.
[INFERENCE] Recommend B as the first experiment, A as the destination: the sign task is cheap and isolates the clause-learning question from the magnitude question.
A.6 Recommended change, in order
- [HIGH] Fix the feedback rule first. Implement Granmo Table 2/3 faithfully in
tmLearnOne(the Type IcOutconditioning and the Type II direction). [FACT] The current Type I ignorescOut, and Type II is unsatisfiable dead code — seedocs/research/tm-deb-assessment.md§6.3 and its §6.5 fixes 1–2. Changing the target shape while the update direction is broken is unmeasurable. - [HIGH] Fix the label baseline. Store
linearX, linearYinTmTrace; passδ = actual − lineartotmLearnOne. Remove the factor-2 shrink and stop zeroing residuals on hits (map hits into the zero bin / keep their true smallδ). - [MEDIUM] Choose the head. Start with per-axis sign classification (Option B) to validate clause learning, then move to binned multi-class (Option A) with an explicit zero bin.
- [LOW] Keep the exact
(fireTick, powerBin)pairing. [FACT] It already works (trainedShots/traceMisses). Do not add discounting.
Measurement that would prove it worked [INFERENCE]:
- Build a deterministic oracle: for a constant-velocity enemy,
δ = 0by construction, so a correct TM must output ≈ 0 (not oscillate). For a constant-acceleration / fixed-perpendicular-drift enemy,δ = actual − linearis analytically known; check the TM's output converges towardδ, notδ/2. - Log per shot
(linear, actual, prediction, correction); metric = mean absolutecorrection − δon held-out ticks, plus hit rate vs theLineargun (the same adversary, same seed). Success = the factor-2 bias disappears and hit rate beatsLinear. [UNKNOWN] the exact accuracy the TM can reach in the repo's data budget.
SECTION B — Can a Tsetlin Machine learn high-level patterns?
B.1 What a TM is actually good at
[FACT] A Tsetlin Machine learns a DNF over binary literals: each clause is a conjunction (AND) of a subset of the 2 × n_input literals (a bit and its complement), and the output vote combines positive- and negative-polarity clauses. In this repo, common_libs/guns/tsetlin.nim tmEvalClause returns 1 iff at least one literal is included and every included literal's value is 1; tmForwardWithCache sums tmPolarity(c) * clause_output into vx, vy. [FACT]
- Learning is by discrete Tsetlin automata (finite-state machines with
Reward/Penalty), updated by Granmo's Tables 2/3; there is no gradient. [FACT] - The learned model is human-readable: a clause is a list of included literals, which can be decoded back to
(frame index, field, bit). [FACT] The clause state array is indexed bytmStateIdx(outIdx, clause, lit);st > 0= included. [FACT] - [INFERENCE] This is the right tool when the target is genuinely a conjunction of discrete conditions, and the wrong tool when the target is a smooth signed magnitude that must be fitted precisely (Section A).
B.2 The repo is already built for a temporal TM — [FACT]s
- Frame-stacked binary encoding already exists in the gun. [FACT]
common_libs/guns/tsetlin.nim:TM_FRAME_BITS = 83,TM_SELF_BITS = 40,TM_WINDOW_SIZE = 10,TM_TOTAL_BITS = 83*10 + 40 = 870.- Gray coding via
tmToGray(value) = value xor (value shr 1)andtmToBits. - The window is LIFO (index 0 = newest) and is shifted once per tick (
if state.tick != g.frameTick), with a stale comment documenting a past over-shift bug. - [INFERENCE] Because every one of the 10 frames is laid out in the same fixed 83-bit block, a single clause can AND literals from different frames, i.e. it can express a temporal conjunction across the last 10 ticks.
- The per-frame field layout (from
tmEncodeFrame): bearing sin 8 + bearing cos 8 + distance 7 + velocity 5 + heading sin 8 + heading cos 8 + 4 wall distances × 7 + energy 11 = 83 bits. [FACT] - A clause over 10 frames is a high-level pattern. [INFERENCE] "closing distance" = a distance-field bit pattern that changes across frames; "decelerating" = consecutive velocity-field codes ordered by Gray code; "low energy across the last N frames" = an energy-field bit repeated across frames. A conjunction of those literals is exactly the kind of rule a TM clause represents natively. [INFERENCE]
Mismatch to record. [FACT] The gun's
tmEncodeFramecall passesstate.selfEnergyfor the frame's energy field, with the comment# use self energy as proxy (enemy energy not in WorldState). That comment is stale:common_libs/gun_harness/gun_interface.nimline 25 definesWorldState.enemyEnergy*. So the current gun encodes self energy twice and never sees enemy energy — which means a rule like "the enemy turns hard when its energy drops below X" is not even representable by the current gun, independent of its clause learner.
B.3 What is reusable from BNNBot_garage today — proc-level
BNNBot_garage/src/binary_encoding.nim — encoder is fully reusable
- Types:
BinaryVector = array[TOTAL_BITS, uint8],EnemyScanFrame,SelfState,OutputVector. - Constants:
FRAME_BITS = 83,SELF_BITS = 40,WINDOW_SIZE = 10,TOTAL_BITS = 870,OUTPUT_BITS = 7,AIM_MIN = -60.0,AIM_MAX = 60.0. - Reusable procs:
encodeFrame(frame: EnemyScanFrame): array[FRAME_BITS, uint8],encodeSelf(self: SelfState): array[SELF_BITS, uint8],encodeFullVector(window, self): BinaryVector, and the utility setformatBinary,popcount,hammingDistance,bitwiseAnd,formatVectorBinary,formatVectorDecimal,formatFrameBinary,encodeOutput(angle): OutputVector,decodeOutput(vec): float. [FACT] - NOT reusable (self-documented placeholders that return hardcoded values):
formatFrameDecimal(returns"199,199,99,15,199,199,99,99,150") anddecodeFrameFields(returns a fixed[0,…,50,50,100]array). [FACT] These are stubs, not decoders — do not build on them. - [FACT] The encoder is already inlined into
common_libs/guns/tsetlin.nimastmEncodeFrame/tmEncodeSelf/tmEncodeFullVector(the gun's header says "adapted from BNNBot_garage/src/binary_encoding.nim"), so there is nothing to import; BNNBot is the reference copy. One real divergence: BNNBot'sEnemyScanFramecarriesenemyEnergy, while the gun's inlined copy takesstate.selfEnergy(B.2 mismatch above).
BNNBot_garage/src/tsetlin_predictor.nim — reusable as scaffolding, but it is a regression TM with the same feedback defects
- Types:
TsetlinNet,ClauseCache,EligibilityTrace,VirtualBullet. Constants:N_IN = 870,N_OUT = 2,N_CLAUSES = 64(32 pos + 32 neg),N_STATES = 15,T = 32.0,S = 4.0,RESID_MAX = 80.0,TRACE_MAX_AGE = 40. - Procs:
stateIdx,clausePolarity,evalClause,makeLiterals,computeVote,initTsetlinNet,forward,forwardWithCache,learnOne,learn. [FACT] - Honest assessment of the learning core [FACT]:
- It is a regression TM (continuous
(cx, cy)scaled to±RESID_MAX), not Granmo's multi-class classifier; the output isvote / T * RESID_MAX. - Its "Type Ia" (
cOut == 1) and "Type Ib" (else) branches execute identical bodies (both grow toward the current input), so the split is cosmetic. The counteracting(c=0, literal=1) → toward-Excludeforce is absent. - Its "Bug 3 fix" Type II decrements included false literals (
if literals[lit] == 0 and st > 0) rather than incrementing excluded false literals; whencOut == 1that condition is unsatisfiable, so the branch is dead code (same defect as the current gun — seetm-deb-assessment.md§6.3). - Verdict: reusable as a prototype scaffold (types, vote, clause cache, eligibility trace), but not as a correct learning core. The fixed feedback rule from Section A.6 step 1 is a prerequisite.
- It is a regression TM (continuous
BNNBot_garage/src/wisard_predictor.nim — not a TM
- Types
WiSARDNet,WaveTrace,VirtualBullet; procsinitWiSARD(seed),computeAddresses,predictCorrection,learnCorrection; constantsK = 14,N_NEURONS = 50,LUT_SIZE = 16384. [FACT] - It is a WiSARD/n-tuple LUT regressor averaging
(sumX/count, sumY/count). [FACT] Reusable only as a non-TM baseline; it regresses(dx, dy)directly, so it bypasses the binary-machinery question entirely. It has no interpretable clauses.
Summary: the encoder is reusable as-is (minus the two stub decoders). The TM code is reusable as reference/scaffolding but is a regression TM with broken feedback rules, not a correct classifier. Both already live inside the gun, so the practical value of BNNBot_garage today is historical reference, not an importable library.
B.4 Concrete, falsifiable experiment
Goal: prove a TM can recover a known high-level temporal rule from the frame-stacked encoding — and measure whether it can do so sparsely.
Setup [INFERENCE unless noted]:
- Define a synthetic enemy with a rule that is (a) deterministic, (b) temporal, (c) expressible as a conjunction over the encoded frame fields. Two candidates:
- Rule E (energy-gated turn): the enemy turns hard left/right on tick
tiff its energy has dropped below a fixed threshold within the lastkframes. [FACT] energy is an 11-bit Gray-coded field per frame; [FACT] the gun currently feeds self energy, so this rule needs theenemyEnergyfix (B.2). - Rule D (fixed perpendicular drift): the enemy holds a constant heading offset relative to the bearing for
kframes → the velocity/heading fields repeat across the window.
- Rule E (energy-gated turn): the enemy turns hard left/right on tick
- Feed the 10-frame stacked 870-bit vector to a standalone classifier TM (not the gun's 2-output regression head): multi-class over
{turn-left, turn-right, no-turn}or{drift, no-drift}. - Train/test split, fixed seed, and a majority-class baseline reported alongside.
Accuracy / sparsity bar that would prove the TM found the rule [INFERENCE — proposed thresholds, not measured]:
- Accuracy: on held-out noiseless frames, a correct rule should be recovered near-ceiling (propose ≥ 95 % on the deterministic rule, vs 50 % chance for a balanced binary task, and vs the majority-class rate). Anything at the majority-class rate means no rule was learned.
- Sparsity: the clauses that win should each include a small single-digit number of literals, and the count should be stable across runs. [FACT] The current gun's clauses saturate at ~131 literals/clause when the feedback rule is broken (assessment §6.4); that is the failure signature to watch for.
- Interpretability check: decode the highest-weight clause's literals via
tmStateIdxback to(frame, field, bit)and verify they land on the energy/velocity/heading fields and the expected frames. This is the test that it found a high-level rule, not a coincidence.
Can the current gun represent such a rule, once its feedback tables are fixed?
- Clause width: [FACT] there is no width cap; a clause may include any of
TM_N_LITERALS = 1740literals (st > 0). Representability is not blocked by clause width. - Feature count: [FACT] 870 input bits / 1740 literals, with the energy field spanning 11 bits × 10 frames. [INFERENCE] Gray-coding means a monotone threshold ("energy < X") is not one literal — it is a sub-pattern of the 11-bit code; a single AND clause captures a specific bit pattern, and the class team as a whole is what learns the threshold. So the rule is representable by the team, not by a single clause.
- Data volume: [FACT]
docs/rl-algorithm-choice.mdstates episodes are ~30–180 ticks; the gun only creates a trace on ticks with a valid power bin, andtrainedShots = 2141(cited as observed intm-deb-assessment.md§4) is the accumulated sample count. [UNKNOWN] whether that is enough for convergence — this is exactly what the experiment settles. - Head shape: [FACT] the gun's head is
TM_N_OUT = 2regression outputs (cx, cy), not discrete rule classes. So the gun as-is cannot be pointed at the Rule E/D classification task — the experiment must use a standalone classifier TM. The gun can only benefit from these findings indirectly.
SECTION C — Delayed reward: the movement case
C.1 Delayed supervision vs delayed reward — a hard distinction
- Delayed supervision (the gun): the outcome (a shot resolving) arrives 4–128 ticks after the fire tick. [FACT]
bulletSpeed(power) = 20 − 3·power(gun_interface.nim) andPowerBins = [1.0, 1.5, 2.0, 3.0]→ speeds17, 15.5, 14, 11px/tick (virtual_bullets.nim); at typical ≈400 px engagement that is ≈23–36 ticks, and the gun's own comment notes a power-3 long shot can take ~128 ticks. But the outcome is exactly attributable per shot:FeedbackEventcarriesfireTickandpowerBin, andonResultindexestmTraceSlot(e.fireTick, binIdx)to the exact stored trace. [FACT] So the label is delayed, not confounded. Effectiveλ = 1; no discounting needed. [FACT, matchingtm-deb-assessment.md§5] - Delayed reward (movement): one sparse outcome per round (win/loss/rank/score) with no action-to-outcome mapping. You cannot say which of the hundreds of per-tick turn/thrust decisions caused a win. [INFERENCE] This is the case where credit assignment — discounting an eligibility buffer, or a TD/REINFORCE estimator — is the correct tool. [FACT]
tm-deb-assessment.md§7 already identifies a learned movement module as the repo's genuine delayed-credit-assignment case.
[INFERENCE] Discounting is only justified for the second case. Applying it to the gun would delete exactly-attributable long-range samples (assessment §4: γ^Δt = 4.4e-7 at Δt = 90).
C.2 What the movement layer actually does today
[FACT] The movement modules in common_libs/movements/ are: the_floor_is_lava.nim, phantom_meteor.nim, wave_surfer.nim, minimum_risk.nim, oscillator.nim, random_oscillator.nim, rammer.nim.
- [FACT] The live mover is TFIL.
ModularBot_garage/src/ModularBot.nimdeclaresmover: TFILModule, initialisesmover: TFILModule(debugGraphics: true), and importsmovements/the_floor_is_lava. - [FACT] TFIL is hand-tuned with no learnable parameters. Every coefficient is a
const:BulletCore = 10.0,BulletAura = 5.0,EnemyCore = 40.0,EnemyAura = 10.0,CorridorHeat = 20.0,WallHotness = 30.0,WallRadiance = 10.0,PillarHotness = 30.0,PillarRadiance = 10.0,CommitTicks = 15,MinCommitTicks = 5,DangerReplanThreshold = 25.0,CoolestLevels = 2,MaxTrackedBullets = 20,GridSize = 36.0,MaxSpeed = 8.0.computeMovebuilds a lava grid, scores reachable tiles, commits forCommitTicks, and picks among safe tiles with a random tie-break (rand(candidates.high)). No reward ever reaches it;resetRoundonly clears transients. - [FACT] Other modules:
minimum_risk.nim— all scoring weights areconst(KEnemy = 1.0,KWall = 0.5,KCorner = 0.8,KTravel = 0.003).oscillator.nim/random_oscillator.nim/rammer.nim— fixed period/threshold constants.wave_surfer.nimhas an observed-danger histogrambins: array[31, float64]incremented as enemy waves pass, andresetRounddoes not clearbins(it persists across rounds).phantom_meteor.nimlikewise keeps itsDangerHistogramacross rounds. [FACT] - [INFERENCE] These histograms are counters, not reward-optimized parameters: no reward is ever propagated into them; they are incremented purely from observed bullet geometry. There is no existing learnable movement parameter trained by any outcome signal.
Reward signal available from the bot API [FACT]:
ResultsForBot(robocode_tankroyale_botapi/schemas.nim) exposes per round:rank,survival,lastSurvivorBonus,bulletDamage,bulletKillBonus,ramDamage,ramKillBonus,totalScore, plus season countersfirstPlaces/secondPlaces/thirdPlaces. Delivered inonRoundEndedviaRoundEndedEventForBot. [FACT]ModularBot.onRoundEndedalready readse.results.totalScoreande.results.bulletDamage(it logs them to/tmp/gun_stats.jsonl), so the plumbing for a per-round reward already exists. [FACT]- Dense per-tick proxies also exist:
getEnergy()(buildStatereads it),onHitByBullet(damage taken),onBulletHit(damage dealt),onBotDeath. [FACT]SAC_LSTM_Bot_garage/src/SAC_LSTM_Bot/rewards.nimshows a hand-designedcomputeReward(...)combining win/loss (+20/−10), bullet damage, ram penalty, charge penalty, and wall ticks. [FACT]
C.3 Minimal delayed-reward learning design
[INFERENCE] A deliberately small first design, one live parameter at a time:
- One live knob. Pick a single low-dimensional movement decision with a small discrete action set — e.g. a bias toward/away from a tile class in TFIL's scoring, or an "engagement distance / aggression" scalar (mirroring
wave_surfer'sWS_PrefDist), or the reversal threshold in the oscillators. Keep it to a few discrete actions so a per-action clause team is trainable. - Bounded experience buffer. Each tick, store
(encodedStateBits, actionIdx)in a fixed-size ring (a few hundred ticks). Encode the state with the existing 870-bit encoder frombinary_encoding.nim(or a movement-relevant subset); this is the same eligibility-buffer shape the gun already uses, just keyed by tick instead of(fireTick, powerBin). [INFERENCE] - Dispatch at round end. In
onRoundEnded, compute a scalar rewardrfromResultsForBot(totalScore/rank/survival; optionally plus a dense term as inrewards.nim), then replay the buffer: for each buffered(state, action)apply Type I/II feedback for the chosen action's clause team, weighted byγ^(T − t). [INFERENCE] - Credit-assignment honesty. One reward per round spread over ~30–180 ticks is extremely sparse and high-variance. [FACT for the tick range from
docs/rl-algorithm-choice.md; INFERENCE for the consequence.] Theγ^(T − t)decay is exactly a heuristic for "earlier actions are less likely to be responsible", and it does not make the attribution correct — it only hedges. [INFERENCE]
How many rounds before a signal is detectable [INFERENCE — standard two-proportion power calculation, computed by hand; no runtime measurement]:
For a win-rate reward, two-proportion sample size per arm
n = (z_{α/2} + z_β)² · [p₁(1−p₁) + p₂(1−p₂)] / (p₁ − p₂)², with α = 0.05, power 0.80 ⇒ (1.96 + 0.8416)² ≈ 7.85.
- Detect
0.50 → 0.55(5 pts):n ≈ 7.85 · (0.25 + 0.2475) / 0.0025 ≈ 1560rounds per arm. - Detect
0.50 → 0.60(10 pts):n ≈ 7.85 · (0.25 + 0.24) / 0.01 ≈ 385rounds per arm.
These are lower bounds for a between-condition comparison; a single adaptive learner is worse (nonstationarity, exploration cost, interference), so treat them as optimistic. [INFERENCE]
Feasibility. Given the brief's figure of ~1 round per few seconds, 385 rounds ≈ 20–30 min, and 1560 rounds ≈ 1.5–2 h per arm. [FACT] docs/rl-algorithm-choice.md also states 10–35 rounds per match and a training window of "~few hundred milliseconds" between rounds. [INFERENCE] So an offline, between-battle experiment is feasible within hours; online learning inside a single match is not (a match has ~10–35 rounds). [UNKNOWN] the actual wall-clock duration of one round in this repo's harness — no file I read records it, so the minutes/hours above are conditional arithmetic, not measured.
C.4 What is UNPROVEN, and the experiment that would settle it
- [UNKNOWN] Does the per-round reward carry enough signal to be learnable at all? It may be dominated by opponent identity, arena randomness, and starting positions rather than by the movement knob.
- [UNKNOWN] The wall-clock cost per round, hence total experiment time and whether the round budget is realistic.
- [UNKNOWN] Whether γ-discounted per-tick feedback beats a fixed hand-tuned policy (TFIL) at all, and by how much — no measurement exists.
- [UNKNOWN] Whether one live knob is enough to move
totalScoremeasurably, or whether the effect is buried in variance.
Experiment that would settle it [INFERENCE]:
- Offline counterfactual first (cheap gate). Replay a fixed corpus of recorded enemy movement tapes. For each candidate action bias, simulate the resulting movement and compute the reward proxy. If no single-knob policy beats TFIL's fixed policy offline, do not spend the online round budget — the signal is not there. If one does, use its effect size to size the online run.
- Online A/B, seeded. Fixed opponent set and seeds; arm A = current TFIL, arm B = TM-biased TFIL. Run
Nrounds per arm from the power calculation above (start at ≈ 385/arm for a 10-pt effect, expand for smaller effects). Pre-registertotalScore/rank/survivalas the metrics and report confidence intervals, not point estimates. - Instrument variance. Log per-round reward and its spread before learning; if the round-to-round standard deviation is large relative to any plausible effect, report that directly as the blocker rather than running an under-powered test.
Explicit non-import. None of the above relies on
docs/papers/tm-deb-paper.pdf. The only external claim used is the standard power calculation (C.3). The TM-DEB idea of discounting a replayed buffer is conceptually the right shape for this case (astm-deb-assessment.md§7 already noted), but no result, equation, or number from that deleted document is treated as evidence here.