cc332138b36acc755f27409feea884c84ea357e0
79 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9cd6e9b8ce |
Ablation: the radial TM is replaceable by a CONSTANT, and its avenue is dead on the
shipped metric The radial TM beats Linear on bmPoint, but its head never beat the majority baseline after the label bias was fixed - suggesting the win is a constant lean rather than learning. So: sweep a stateless constant short-range offset (new `common_libs/guns/radial_offset.nim`, no learning at all) against the learned TM. VERDICT (measured, offline range, seeds=3, 18 paired runs, 231 TM rounds): 1. **bmPoint - REPLACE the TM with a constant.** `RO_s0.95` (aim distance x0.95) TIES it early (9/9, p=1.0) and BEATS it overall (15/3, p=0.0075; 7.47% vs 6.89% per-run mean). A fixed -20px does the same. The head never beats its majority baseline (56.2% vs 57.2%). 2. **bmPath (the SHIPPED metric) - the radial avenue is a DEAD END.** TMRadial is a systematic LOSS there (2/16, p=0.0013); every constant is within +-0.2pp; the only real bmPath effect is the BotRadius clamp. So the radial shift cannot help the shipped configuration. 3. The per-adversary optimum DOES vary (fixed -10 for crazy, -30 for tr_crazy, scale 0.95 for three others) - but ONE GLOBAL CONSTANT still beats the adaptively-trained head, so the "fragility justifies learning" argument FAILS. THE REAL FINDING UNDERNEATH, and it generalises beyond this gun: the base linear prediction systematically OVERSHOOTS. Measured raw per-tick base radial error has mean -71 to -100 px; the enemy is NEARER than the prediction in 63-81% of shots and farther in only 4-14%, CONSISTENT ACROSS ALL SIX CAPTURES. Radial label histogram [415166,126461,120325,44694,19351] = 57.2% majority class, mean label -82.3 px, mean applied shift -37.6 px. So the net-short bias is a GENUINE property of these range-holders against a constant-velocity extrapolation (they decelerate and turn, so the true position is closer than the straight-line guess) - NOT a fixture artefact. That is worth chasing for the guns that actually ship. Caveat: bmPoint is not the shipped metric (bmPath won the real-hit-rate A/B for SELECTION), so a bmPoint win is not yet evidence of a real win. That needs a live test - and the natural target is Pattern, which is now the default and best gun. Adds radial_offset.nim + sweep_radial_offset.nim; tm_pattern.nim gains additive instrumentation only (radial label mean and applied-shift mean; no behaviour change, and test_tm_pattern_registration still passes all 20 checks). |
||
|
|
589a230106 |
TM radial gun: registered (default OFF) + label-bias fix that removes the bias but
retracts its own earlier learning claim === TASK 1: REGISTERED AS GUN 14, DEFAULT `off` === The radial TM gun is now a first-class rack member (`TMPATTERN`, id 14), forceable alone with `TR_RACK_TMPATTERN=both` plus every other `TR_RACK_*=off`. DEFAULT IS `off`, and the justification matters: `both` would let it compete for selection AND (because the shared VirtualTracker ring is order-sensitive) shift every other gun's learning order, so it CANNOT leave the default path unchanged. With `off` its predict and spawnBullets are additionally GATED on rack admission (the only gun wired that way), so the shipped default never spawns it at all: zero cost, zero ring perturbation. Live proof: 1-round battle with only TMPATTERN racked -> `gun 14 (TMPattern): vShots=400 selected=104 other-gun selections=0`. Default-path-unchanged proof: parity checks that the 15-gun default bestGun/ selectGun equals the old 14-gun rack RNG-draw-for-RNG-draw, that gun 14 is never selected by default, and acceptance 12/12. Cost: 0.36 ms/tick (predict 0.30 + onResult 0.05) ~= 3% of the 13.16 ms budget. Tsetlin in the same harness is 1.62 ms/tick, so the new gun is ~4.5x cheaper. === TASK 2: THE LABEL-BIAS FIX - AND A RETRACTION === Root cause confirmed: under bmPoint a SHORT radial correction resolves the virtual bullet BEFORE the base arrival tick, so the label was dropped (labelMisses). Fix: defer the label in a pending queue and flush it once the arrival tick is recorded; labels still come from the BASE arrival tick. labelMisses 4,281,695 -> 0 training samples 1,071,824 -> 5,345,847 (x5) radial head acc 48.8% -> 57.0% (shuffled control 20.0%) bmPoint hit rate 9.4/5.8% -> 9.1/5.7% (unchanged, within noise) So the fix IMPROVES LEARNING but NOT the metric. **RETRACTION OF THE PREVIOUS JOB'S CLAIM.** It reported the radial head's 48.8% against a 36.7% majority baseline and concluded "conditional learning, not a constant bias". With the bias removed, the correctly-measured majority baseline is **58.2%** - so the head at 57.0% is AT/BELOW majority. The earlier apparent conditional learning was PARTLY AN ARTEFACT OF THE BIASED SAMPLE. The bmPoint metric win is real (TMRadial > Linear early 16/2 p=0.0013, overall 18/0 p<0.0001; > shuffled 18/0 p<0.0001) but it comes from a NET-POSITIVE AVERAGE RADIAL SHIFT, not from beating a majority classifier. Recorded plainly rather than left standing. Guards: test_tm_pattern_registration 20 (new), test_tm_pattern_rack_live 4 (new), test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3 (the SIGSEGV is gone - the knn_gun rewrite is now committed), test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28, test_rack_membership 38, test_selector_tiebreak 19, test_tm_pattern_learning 3, acceptance_offline_vs_online 12/12. ModularBot compiles (release). Note: `common_libs/tests/range_guns.nim` still builds 14 offline drivers (the offline sweep constructs TmPatternGun directly and acceptance only inspects ids 0..13), so nothing breaks - but a future job wanting it in the offline rack must add a 15th driver and mirror the live admission gating. gun_stats.jsonl now emits 15 rows; downstream tooling should ignore id 14. |
||
|
|
4657fe715e |
wave pairing: 36-58% of GF/DecayGF/KNN learning samples were MISLABELLED
The audit inferred (from code) that GF/DecayGF/KNN pop the OLDEST wave on resolution, while under bmPath bullets leave the arena in NON-FIFO order - so an outcome could be attached to the wrong wave. It also noted that `starved=0` does NOT rule this out. Both halves are now MEASURED. MISPAIRING RATE (10 DrussGT fixtures, real VirtualTracker, 344k resolutions/gun): gun bmPath mispair label err bmPoint mispair label err GuessFactor 36.48% 19.39% 18.24% 7.62% DecayGF 36.85% 19.52% 20.57% 8.64% KNN 57.91% 27.63% 29.75% 11.58% (starved = 0 everywhere, exactly as the audit predicted) So ~1 in 5 GF/DecayGF learning samples and ~1 in 4 KNN samples carried a WRONG guess-factor bin. This is a material corruption of the learning signal. FIX: the same fireTick-keyed ring scheme `tsetlin.nim`/`tm_selector.nim` already use - `slot = (fireTick*4 + bin) mod 1024` (period 256 ticks, longer than the ~91-tick max flight), looked up by exact key. Public interfaces unchanged; added `waveResolved`/`waveMispaired` integrity counters. AFTER: mispaired = 0 and starved = 0, both metrics, all three guns. EFFECT ON HIT RATE: SMALL AND NOT SIGNIFICANT. bmPath 4000 samples/gun: GuessFactor 23.20% -> 23.02% (-0.18pp, per-run sign-flip p=0.750) DecayGF 23.80% -> 24.25% (+0.45pp, p=0.625) KNN 18.27% -> 18.80% (+0.53pp, p=0.547) bmPoint: +0.05 / +0.33 / -0.15pp, p = 1.00 / 0.50 / 0.50. Per-run ranges overlap almost completely. A bullet-level z-test is anti-conservative (bullets within a fixture share a trajectory) and its KNN p=1.9e-16 cannot be trusted given ~10 effective independent runs. PLAIN READING: this is a CORRECTNESS fix, not a measurable hit-rate win. It removes a 36-58% mislabelling of the learning signal; the point estimates move by at most ~0.5pp, within run-to-run noise. Stated plainly rather than oversold. A REGRESSION IT CAUGHT IN ITSELF (and this explains the SIGSEGV another job saw and correctly attributed to a concurrent knn_gun.nim rewrite): the first implementation put an inline `array[1024, KNNWave]` (~100KB) inside each gun, which overflowed the default 8MB stack and made `test_power_selection` SIGSEGV. Causation was proven by stashing only the three gun files (test passed), then fixed by making the rings heap-backed `seq`. Verified: `test_power_selection` 3 PASS on the default stack, and zero inline `array[1024]` remain. Guards: test_wave_pairing 17 (new, pure), test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28. ModularBot compiles. Adds audit_wave_pairing.nim and compare_pairing.nim. |
||
|
|
1ea72c7f14 |
TM gun round 2: base was never behind; RADIAL target beats Linear on bmPoint
=== TASK 1: MY PREMISE WAS REFUTED ===
I instructed the job to "fix the baseline" because an earlier measurement said the
TM gun's base did not iterate flight time like `LinearGun`. MEASURED: the new gun's
base is BYTE-FOR-BYTE `LinearGun` - 18/18 runs tie exactly, p=1.000, every per-run
row byte-identical. The "non-iterating baseline" belonged to the OLD `tsetlin.nim`,
not this gun. So no fix was needed, and the earlier inference should not have been
generalised to the new gun. (It did still align the zero-correction clamp to
LinearGun's exact [0, arena] range, and reports the old BotRadius-inset base was a
wash/marginally better at 34.2%/24.7%.)
=== TASK 2: THE RADIAL TARGET - A CONTROL-VALIDATED WIN, BUT ONLY ON bmPoint ===
Instead of the lateral (GF-bucket) component - which the linear lead already
captures - the TM now predicts the RADIAL component: will the enemy be nearer or
farther than the base prediction when our bullet arrives? A 5-class radial head
sharing the same 40-bit context and TM core; the readout advances/retards the aim
distance along the base bearing.
under bmPath (the SHIPPED metric): STRUCTURAL NO-OP
synthetic 8/8 exact ties, p=1.0; real 33.9%/24.1% vs Linear 34.0%/24.3%
under bmPoint: A WIN, control-validated
TMRadial 9.4% (6013/63785) / 5.8% (42079/726652)
Linear 7.2% / 4.7% overall 17/1, p=0.0001
Tsetlin 7.0% / 4.8% overall 15/3, p=0.0075
shuffled 7.0% / 3.6% early 17/1 p=0.0001; overall 18/0, p<0.0001
radial head online accuracy 48.8% vs 19.9% shuffled chance and 36.7% majority
-> it is CONDITIONAL learning, not a constant short-range bias.
Best config: TM_RADIAL_RANGE=60, TM_RAD_MARGIN=0.25, 5 classes.
CAVEAT THAT MATTERS: a win on `bmPoint` is NOT yet evidence of a real win. `bmPath`
is the shipped SELECTION metric precisely because it beat `bmPoint` on real hit
rate (7.43% vs 4.70%). But that A/B was about which gun to PICK, not about gun
QUALITY - a gun can be better in reality while scoring worse on the selection
metric. So this needs a LIVE test, and it is the decisive one.
=== TASK 3: REVERSAL TARGET - CLEAN NEGATIVE ===
The label positive rate is only 9.7% (rev=[24772,2673]) and the head's 86.8%
accuracy is BELOW the 90.3% majority baseline: it does not learn the positive
class at all. Hit-rate effect neutral (bmPath 19.5%/18.4% vs shuffled 19.1%/17.8%,
p=0.24/0.82). Dropped.
=== OVERALL ===
Not competitive on the shipped bmPath metric (gated GF 28.3%/22.2% vs Linear
34.0%/24.3%, p=0.0075). Better than Linear on bmPoint via TMRadial (+2.2pp early,
+1.1pp overall). Per-enemy reset exists; a fresh gun per round; NO cross-battle
persistence (the user's non-negotiable).
MEASURED LIMITATION: radial mode has a high labelMiss because aiming short
resolves BEFORE the base arrival tick, biasing training toward resolvable samples.
The metric win is label-independent. A deferred-label fix is the next refinement.
INFERRED: the mechanism is surfers being NEARER than the base prediction
(range-holding); a constant-short-offset ablation would separate a learned
short-range bias from genuine per-tick conditional prediction.
|
||
|
|
78975a35c4 |
cost: parallel per-enemy virtual bullets are NOT affordable as proposed
Benchmark driving the real rack and the real VirtualTracker over 7 recorded DrussGT fixtures synthesised into an N-enemy melee. Answers "the virtual bullets are cheap, why not keep fitness for every enemy in parallel?" (the user's idea, motivated by making kill-stealing target switches free). BASELINE: exactly 4.0 predict calls per gun per tick (one per power bin) - 52/tick for the shipped 13-gun rack (TMSelect is compiled out). The task's 56/tick was the 14-gun figure. VERDICT: NOT AFFORDABLE. Budget is 13.16 ms/tick (76 ticks/s measured live). N=1 46% of budget N=2 94% <- already at the edge N=4 189% N=6 274% Marginal cost ~= 5.9 ms per extra target, linear. TWO FINDINGS THE PROPOSAL MISSED: 1. `onResult` TRAINING dominates, not predict. Tsetlin's onResult alone is 3.40 ms/tick - ~99.5% of all 13-gun onResult cost - doing ~174k rand() calls per resolved bullet. Every spawned bullet that resolves triggers it, so it scales 1:1 with targets. The per-target cost is the Tsetlin training pass. 2. `MaxBullets=8192` is a HARD BLOCKER, not just CPU. Spawn rate is 56*N/tick and path-metric bullets live until they hit a wall (40-90 ticks). Measured dropped bullets/tick: N=1 -> 0, N=2 -> ~3, N=4 -> ~180, N=6 -> ~300. At N=6 the ring wraps every ~24 ticks, so most bullets are silently clobbered and never scored. A working N=6 pipeline needs MaxBullets ~30k-50k (~4-6 MB, cheap RAM). ALSO MEASURED: Tsetlin and KNN do NOT cache per tick - they redo the full TM forward pass / full KNN scan for EACH of the 4 power bins (Tsetlin 1.94 ms/tick of predict, KNN 0.45). The earlier "tick-only cache" fix never touched the two most expensive predicts. Pattern and TMSelect do cache fully. MITIGATIONS (measured predict+spawn at N=6 vs 13.59 ms baseline): nearest-K=1 only 45% budget nearest-K=2 95% rotate every 3 ticks 95% drop Tsetlin for extras ~68% (INFERRED from Tsetlin's measured 90% share) Tsetlin is ~90% of the per-target cost, so excluding it from non-primary targets makes N=6 fit. "Resolve less often" is not a separate lever - resolution IS when training happens. ARCHITECTURAL CAVEAT (correctness, not cost - and not priced into the proposal): the shared-rack topology is broken for this. The guns are global singletons, so predicting for enemy B ADVANCES/OVERWRITES enemy A's velocity tracker, KNN feature history and Tsetlin frame window in the SAME instance. Per-enemy fitness with correct histories therefore requires PER-ENEMY GUN INSTANCES, which is what this benchmark measured. That multiplies the (already dominant) Tsetlin cost. CONSEQUENCE FOR THE PLAN: combined with the measured finding that the selector is negative value and the rack should shrink to a few good guns, this work is much less valuable than assumed - with a small rack (Pattern's predict is 40us and fully cached) the cost falls proportionally. Priority lowered accordingly. Caveat: the host was heavily loaded (load 15/16), so absolute ms carry ~30-50% noise; min-of-2 and two independent runs agree on the trend, the Tsetlin dominance, and the ring overflow. No melee fixture exists in the repo, so the 7 enemies are 7 distinct recorded trajectories (stated in the file header). |
||
|
|
0ede6d12ec |
selector: arrival-accuracy tie-break measured NEGATIVE; randomness is load-bearing
Hypothesis under test (from the gun audit, which named the tie-band as "the lever that matters most"): `bmPath` is deliberately generous (2.3-3.6x `bmPoint`), so a gun can sit in the tied band on a ray that sweeps the target's path while its bullets ARRIVE badly. So: keep the `path`-ranked band (path beat point on real hit rate 7.43% vs 4.70%, z=5.56), but narrow the random draw inside it using a parallel `point` (arrival-accuracy) window. RESULT: NO EFFECT. Real DrussGT, ONE frozen binary (/tmp/ModularBot_tieband, md5 2c0c56e6...), env knobs only, 7 runs x 7 rounds per arm, server-side events sidecar, exact two-sided permutation test on per-run rates. arm runs shots real % dmg/run d p tbbase (shipped) 7 4128 7.17 175 -- -- tbpt path-rank + point-narrow 7 3938 7.08 165 +0.14 0.88 tbpc =commit control 7 3759 4.44 98 +2.74 0.0012 tbpt25 point margin 0.25 7 3683 5.59 119 +1.65 0.20 tbtie05 / tbtie40 (band width) 7 3937/3917 5.84/6.28 133/144 1.49/1.00 0.11/0.25 tbwin50 (SelectorWindow=50) 7 3983 6.05 139 +1.20 0.11 tbfloor10 (FloorPeakFrac=0.10) 7 3829 5.33 118 +2.12 0.11 tbpt vs base: fully overlapping ranges, p=0.88. This is a REAL null, not a dead arm - the mechanism was live, and it visibly changed the selected-gun mix (Pattern 24%->16%, Accel 6%->16%, Tsetlin ~0%->13%). CONTROL VALIDATED, AND THIS IS THE THIRD TIME: removing the random draw inside the band is SIGNIFICANTLY WORSE (4.44%, p=0.0012). Combined with the earlier hysteresis A/B (7.02% -> 5.10% for commitment) and the light-hysteresis result, the selector's per-tick randomness is now load-bearing on three independent measurements. Narrowing the band on ANY second virtual statistic has not helped. Every knob swept (band width, floor, window) is nominally worse than shipped at n=7; that is "no credible win" rather than "proven harm" (sd ~1.8pp, ~1pp resolution, underpowered). Shipped default stays `GUN_SELECTOR_TIEBREAK=off`; the feature is opt-in, fully guarded, and costs zero extra work on the default path (point windows are scored only when the mode is on). Guards: test_selector_tiebreak 19 (new, pure), test_gun_harness 39, test_vbullet_metric 11, test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28, test_rack_membership 38, acceptance_offline_vs_online 12/12 PASS (offline path calls neither chooseFromFit nor the tie-break). STRATEGIC CONCLUSION: three selection-side attempts have now failed (hysteresis, commitment, point tie-break). The selector is at a local optimum and the remaining lever is the QUALITY OF THE GUNS, not the selection among them. |
||
|
|
ca82053a11 |
TM gun: the discrete-target diagnosis was RIGHT - it learns now. Still loses to Linear.
The user's goal: a TM gun that is the best 1v1 gun, starting from scratch every
battle but quickly overfitting the current enemy. The previous attempt (knob
tuning) failed: NO configuration beat its own shuffled-feedback control, and the
TM-off ablation scored the same as TM-on, i.e. the TM's correction was
near-zero-mean noise. Diagnosis then: a Tsetlin Machine is a CLASSIFIER, and we
were asking it for an absolute aim point - a regression target. So this attempt
gave it a DISCRETE target (multi-class over guess-factor buckets) with 40
binary/bucketed motion features, and measured it against Linear, the default
Tsetlin gun, and a MANDATORY shuffled control.
THE DIAGNOSIS IS CONFIRMED - THE TM LEARNS, DECISIVELY:
online class accuracy 46.0% vs shuffled control 20.0% (2.3x chance)
raw ungated argmax 21.2%/18.6% vs shuffled 15.2%/8.3% (18/18, p<0.0001)
TMPattern > its shuffled control, overall 17/1 runs, p=0.0001
Compare the previous attempt, which could not beat shuffled feedback at all.
TMPattern also beats the default Tsetlin gun early (17/1, p=0.0001), so it is a
strictly better TM gun than the one in the rack.
BUT IT IS NOT COMPETITIVE WITH LINEAR ON REAL SURFERS:
real DrussGT, bmPath (the shipped metric), 3 seeds, pooled early/overall
Linear 34.0% (6358/18715) 24.3% (58297/239943)
TMPattern (gated) 27.9% (15514/55535) 22.0% (158658/719681)
TMPatternShuf 28.7% 19.4%
Linear > TMPattern: 15/18 early p=0.0075, 15/18 overall p=0.0075
bmPoint: neutral (7.2%/4.6% vs Linear 7.2%/4.7%)
synthetic controlled motion: matches/edges Linear (66.8%/60.6% vs 66.4%/59.6%,
shuffled 55.7%/50.1%) - the mechanism works when motion is predictable.
So: the representation fix moved this from "learns nothing" to "learns strongly
but applies its knowledge badly". INFERRED reason for the residual loss: the
linear lead is already the modal GF bucket (the label histogram is centred), so
corrective excursions away from it are net-negative. The measured deficit lives
in the BASELINE and in RANGE, not in the TM knobs - which is why further knob
tuning was never going to work.
Best config: gated hard K=5, TM_CONF_MARGIN=0.25, TM_SHRINK=0.5.
NOT TRIED (time-boxed): the binary-reversal target, and a RADIAL (range-holding)
target - the latter is the top next step.
Adds `common_libs/guns/tm_pattern.nim` (NOT registered in the rack),
`common_libs/tests/sweep_tm_pattern.nim`, and a durable writeup at
`common_libs/tests/tm_pattern_sweep_results.md`.
|
||
|
|
a73de13458 |
racks: separate melee and 1v1 gun racks, plus per-mode real hit-rate data
The user's plan: "separate racks for melee and 1v1, so the bot switches from those based on the situation, and we can put the guns we want in one or both racks." MECHANISM - `RackMode` (rm1v1/rmMelee) derived from SERVER TRUTH: `rackMode(enemyCount)` = 1v1 when the count is 1, melee otherwise. This is the SAME `getEnemyCount()` value the radar already uses, so there is now ONE definition of the mode. (Using the tracker's known-enemy count was a previous bug in the radar: it read 1 before the second enemy was scanned.) - `RackMembership` per gun: both (default) | 1v1 | melee | off. - The selector ranks only admitted guns - including the floor path and the incumbent-hysteresis path. - Empty filtered set FALLS BACK to the full rack, so the bot can never end up with no gun. - Env-overridable at process start, no rebuild: `TR_RACK_<GUN>` for all 14 guns (TR_RACK_HEADON, TR_RACK_LINEAR, ... TR_RACK_TMSELECT), values both|1v1|melee|off. Empty/unknown -> both + a stderr warning, never fatal. - `[rack] mode=<1v1|melee> active=<guns> overrides=<...>` logged once per mode change, never per tick. DEFAULT IS UNCHANGED: every gun ships `rmBoth`, so behaviour is byte-identical until the user re-racks anything. Verified by the unit test's default-config selection parity (RNG draw for RNG draw) and by `test_gun_harness` 39 and acceptance 12/12. `chooseFromFit` iterates the admitted list in ascending id order, so the random tie-break draws are unchanged. NO TUNING DONE, deliberately: we had no per-gun melee hit-rate data, and an earlier 15-paired-run experiment found pruning neutral-to-negative on hit rate (p=0.57/0.21). So all guns stay `both` and the membership pass waits for data. PER-MODE DATA PLUMBING (this is what unblocks that pass): per-gun real shot accounting is now split by the rack in force at fire time, adding to gun_stats.jsonl: realShots1v1, realHits1v1, realHitRate1v1, realShotsMelee, realHitsMelee, realHitRateMelee. Verification: test_rack_membership 38/38 (new, pure, no battle); test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28; acceptance_offline_vs_online 12/12 VERDICT PASS; ModularBot compiles. The live `[rack]` line was observed switching 1v1 -> melee when the enemy died. The offline range never calls the selector (only spawnBullets/tickBullets/ reportFor), so mode filtering cannot change the offline result and no offline mode parameter was needed - confirmed by reasoning over the source and by 12/12. |
||
|
|
ab86c0481f |
gun audit: the virtual system is sound; the rack is redundant, not broken
Audited all 14 guns offline over the committed DrussGT fixtures (~150k resolved bullets/gun) plus 123 rounds of live gun_stats. Prompted by a GUI observation that selected guns "fire dozens of pixels away" and a suspicion of reverse selection. MY HYPOTHESIS WAS WRONG. I expected guns to be ignoring `bulletSpeed`, which would make their 4 power bins identical and the per-bin fitness pure noise. MEASURED: only `HeadOn` is speed-blind (100% identical bins) and that is its correct definition. Every other gun emits 91-94% DISTINCT per-bin predictions (mean intra-tick bin spread 62-103px). The earlier "tick-only cache collapsed all bins onto bin 0" fix is complete across the whole rack. LEAD/SIGN/UNITS ARE CORRECT: replaying synthetic ground truth, all 14 guns score 100% on a stationary target (which also proves predictions are ABSOLUTE - a relative or angle return would score 0), ~100% on constant-velocity for every leaded gun, Circular 100% / Accel 99.8% on a 3deg/tick circle, WallBounce 98.2% on a bounce. No missing lead, no sign inversion. Resolution is right (BotRadius 18, hit credited to the owning gun). FEEDBACK IS INTACT: offline pushes=151260/starved=0, Tsetlin trained=149205/ traceMisses=0; live `vStarved=0` and `vDropped=0` across all 123 rounds. THE "43% FLAT / 4x OPTIMISTIC" EVIDENCE I CITED IS NOT REPRODUCIBLE on current code/data. Live virtual/real ratios against DrussGT are 0.8-2.0 for most guns (Linear 10.9 virt / 13.4 real; KNN 8.1/7.1; DecayGF 10.5/9.8). The 43%-flat session matches an older config or a weak opponent (SittingDuck), not DrussGT. `bmPath` IS 2.3-3.6x `bmPoint` - but by design and documented: it asks "does the ray eventually sweep the target's path", a deliberately generous relative signal. So the flat tie is a RANKING artefact: many guns share the same base forecast and, with a near-zero learned correction, collapse onto the same ray; RelTieMargin=0.20 then treats the top ~half of the rack as tied. DUTY AND OVERLAP (>=50% of ticks within 20px = redundant): Tsetlin ~ StopShot 87% -> duplicate pair DecayGF ~ GuessFactor 92% -> duplicate pair Accel ~ Circular 65% -> partial duplicate WallBounce~ Linear 58% AvgLead = the MEAN of Linear+Circular+WallBounce (constructed redundancy) Displace worst point% (6.6) AND worst real% (2.4); wins no bucket TMSelect DEAD - never spawned (EnableTmSelector=false), 0 shots in every log Pattern the ONLY gun competitive in every distance/speed bucket HeadOn/Linear/Tsetlin/StopShot are identical copies of each other on a real surfer (v<1 ~50.8%, everything else ~2%) RECOMMENDED LEAN RACK (8): HeadOn, Linear, Circular, Accel, Pattern, GuessFactor, KNN, WallBounce. DROP (6): TMSelect (dead), AvgLead (constructed mean), Displace (worst), DecayGF (92% GF), StopShot (87% Tsetlin), Tsetlin (the repo's own sweep already showed it learns nothing on DrussGT). HONEST HEADLINE: pruning is NOT expected to raise hit rate - an earlier 15-paired-run experiment found it neutral-to-negative (p=0.57/0.21). The mechanism by which it could help is a SELECTOR effect (shrinking the tied band), not a gun effect, and that is UNVERIFIED until A/B'd. The virtual system and the rack are basically sound; the lever that matters most is the selector's metric/tie-band, not deleting guns. DESIGN SMELL FOUND (INFERRED, not measured): GF/DecayGF/KNN `onResult` pops the OLDEST wave, but under bmPath bullets leave the arena in non-FIFO order, so a resolution can be paired with a neighbouring tick's wave. starved=0 does not rule this out. Candidate fix: key waves by fireTick, as Tsetlin/TMSelect do. Adds common_libs/tests/audit_virtual_guns.nim (offline, no shipped file touched). |
||
|
|
994f88d7a7 |
ramming: make the decision PROACTIVE, with a bullet-rain abort
The user watched 1v1 and melee runs and saw ram opportunities arise that the bot
declined: "there were moments where the bot could jump over the enemy and shred
it but shot it down instead."
DIAGNOSIS - a chicken-and-egg loop. `ramOpportunity` required dist < 50px, but
an offline measurement over 15 rounds vs DrussGT found the closest approach was
118.7px and the <50px trigger had NEVER fired: the mover has no reason to close,
so the trigger waited for a proximity nothing created. The MECHANISM to close
already existed (the ring mover expresses a ram as band=(0,50)); what was
missing was a decision that fires at a range the bot can actually close from.
Changes:
- opportunity gate relaxed: dist 50 -> TR_RAM_OPP_DIST (200), energy margin
+30 -> TR_RAM_OPP_MARGIN (15). Both env-tunable, no rebuild needed.
- New pure module `common_libs/movements/ram_decision.nim` holding the trigger
and abort logic (no battle/API deps), so it is unit-testable.
- BULLET-RAIN ABORT (the user asked for this earlier): `onHitByBullet` now
accumulates `e.bullet.power` into a 15-turn ring; damageRatePerTurn = sum/15;
an in-progress ram aborts when rate > TR_RAM_ABORT_DMG (0.5/turn). On abort:
isRamming=false, cooldown 30, TARGET KEPT, and the mover returns to the normal
range band. It never stops the bot.
- Opt-in, DEFAULT-OFF `plan` trigger for "change of plan when the gun duel is
failing" (dist<250, margin+20, selected gun's pooled virtual rate < 0.05).
Left off because a cold gun reads 0.0 and would qualify - speculative.
- `TR_RAM_LOG=1` change-gated line: `[ram] ON reason=opportunity dist=143
selfE=78 enemyE=41 cap=3.0 band=[0,50]` / `[ram] OFF reason=bulletRain`.
- Existing cooldown/duration/stuck machinery untouched (stuck>10 or duration>60
-> abort + cooldown 30). A refactor bug that briefly DROPPED the
`ramCooldownTicks == 0` gate was caught and fixed.
Trigger set (first match wins): finisher (dist<300, enemy<20, we are healthier);
opportunity (dist<200, we lead by 15+); desperation (both <5, dist<150); plan
(off). All require enemy>0, a valid target, and no cooldown.
Proof the intent now fires at a closable distance (28/28 unit checks):
PASS: opportunity fires at dist 143 with a 37-energy lead <- the exact case
PASS: old gate (dist<50, margin+30) does NOT fire at 143 <- the old bug
PASS: fires at 199px / does NOT fire at 201px
+ margin, finisher priority, desperation, plan on/off, window mean, abort
threshold checks.
HONEST FRAMING: ram damage is 0.6 per CONTACT EVENT, one-shot (collision
resolution rewinds positions so contacts do not stream) - small next to a p=3.0
bullet hit (16). The payoff is that point-blank forces hit probability toward 1,
so heavy bullets stop missing and E[dE]=p(3P-1) turns positive above P=1/3; ram
damage also scores 2.0/point (highest in the game) and a ram kill carries a 0.30
bonus vs 0.20. So this is "force the fight to point-blank", not "the ram shreds
them". Base rate is rare (2 collisions in the whole fixture corpus).
Guards: test_ram_decision 28 (new), test_power_policy 26, test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24. Compiles (release).
UNVERIFIED: the live effect. No A/B has run, and whether the mover actually
reaches contact is unproven.
|
||
|
|
c9825dfb0b |
power policy: cap power by range and energy, gate 3.0 on above-average chances
Implements the user's energy management request: "firing from more than 200px should be a 'not good chances zone' so faster bullets and more chances to hit matters more than single hit damage with low chances. When we are lower than 50 health, same thing. I would like to use 3.0 power only when the chances of hitting are higher than average." Design: a CAP on top of the existing `bestPower`, not a rewrite. `bestPower` still answers "which bin does this gun's own data prefer"; the policy caps it: ramming -> 3.0 (reason ram, exempt) dist > TR_POWER_FAR_DIST (200) -> 1.0 (far) elif selfEnergy < TR_POWER_LOW_ENERGY (50) -> 1.0 (lowEnergy) elif pEst <= pRef -> 2.0 (belowAvg) else -> 3.0 (full) power = min(gunPreferredBinPower, cap) # can only LOWER power p=1.0 is the right "low" tier on the measured mechanics: bullet speed 20-3p so p=1.0 gives speed 17 vs 11 at p=3.0 (55% faster = less lead error), fire interval 10+2p so 12 ticks vs 16 (33% more shots), and drain 0.083/turn vs 0.1875 (2.25x slower). All three things the user asked for at long range. pEst = the chosen bin's virtual rate (gun aggregate when the bin is empty); pRef = the gun's aggregate mean unless TR_POWER_REF > 0. No-data guns are vacuously below-average -> cap 2.0 (conservative, documented). Control arm: TR_POWER_POLICY=0 = uncapped = today's behaviour exactly. Knobs: TR_POWER_POLICY, TR_POWER_FAR_DIST, TR_POWER_LOW_ENERGY, TR_POWER_FAR_CAP, TR_POWER_MID_CAP, TR_POWER_REF, TR_POWER_LOG. TR_POWER_MID_CAP exists because the user did not specify the middle case (close + healthy + not-above-average); 2.0 is the default, flippable to 1.0. Seam: the cap lives in a pure `applyPowerPolicy` and is applied only in `selectShot` (the single place real shots are chosen), so the logic is testable without a battle. Ram is wired from `shouldRam` - the same value the movement dispatch uses for the (0,50) band. CORRECTION TO AN ASSUMPTION IN THE TASK: `offline_range.nim` does NOT call `bestPower`/`selectShot` - it only replays virtual-bullet spawn/resolve across all power bins, independent of the real shot's power. So there is no offline power-selection path that could diverge from the live one, and the acceptance test guards the metric, not the policy. Policy coverage therefore comes from the new unit test. Verification: test_power_policy 26/26 in BOTH modes (default and TR_POWER_POLICY=0 control arm); test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24; acceptance_offline_vs_online 12/12 VERDICT PASS (live battle). ModularBot compiles. UNVERIFIED: the live effect on damage/survival/score. No A/B has run. |
||
|
|
9caf1d3728 |
movement: range-weighted TFIL variant + tamed heat field (opt-in, default unchanged)
New mover `the_floor_is_lava_ring.nim`, a COPY of `the_floor_is_lava.nim` (which stays byte-identical - the user explicitly wants the current TFIL preserved). Selected only via `TR_MOVEMENT=tfil_ring`; the default stays `tfil`. WHY: our measured real hit rate vs DrussGT is strongly range-dependent - 21.6% at 0-100px, 27.1% at 100-200px, 19.3% at 200-300, 10.9% at 300-400, 6.8% at 400-600, 5.4% at 600-800 - but we shoot from ~450px on average. Plain TFIL has no range preference at all. THE ONE CHANGE: the final tile draw is re-weighted toward a target band. rangeW(d) = 1.0 if lo<=d<=hi; exp(-((lo-d)/K)^2) if d<lo; exp(-((d-hi)/K)^2) if d>hi w_i = rangeW(d_i)^(1/T); chosen ~ Categorical(w) FLAT TOP on purpose: a Gaussian centred on the band midpoint would collapse the band to a point and destroy the within-band hedge. `T` is the only knob; `TR_TFIL_RANGE_TEMP=0` gives plain `rand(candidates.high)` - the exact control arm. Safety stays a HARD constraint: the weighting only reorders the draw among the pool the old code already accepted, so it can never pick a tile the old code rejected (monotone refinement). Small pools (<4) stay uniform. Randomness is deliberately KEPT: a measured A/B showed committing to the "best" tile made real hit rate WORSE (7.02% -> 5.10%), so the distribution is tilted, never removed. HEAT TAMING (ring copy only; env-overridable): TR_TFIL_CORRIDOR_HEAT 20.0 -> 5.0 TR_TFIL_WALL_HOTNESS 30.0 -> 10.0 Rationale, measured: `CorridorHeat=20` is TWICE `PathDangerThreshold=10`, so a single corridor could poison a path by itself; `WallHotness=30` with `WallRadiance=10` put the outer two tile rings over threshold on their own. Per-source shares of total lava: wall 60.6%, corridor 25.4%, pillar 7.4%, everything else <3%. MEASURED EFFECT (primary fixture, 20,026 ticks / 15 rounds, field identity verified max diff 0.000e+00): metric original(20/30) ring(5/10) band-weightable ticks 10.79% 26.45% mean safeTiles/tick 15.19 85.04 ticks with 0 safe (pre-fallback) 58.5% 7.5% safePool >= 4 39.05% 92.47% >=1 safe tile in 100-200px 11.84% 26.64% MEAN CLOSEST-SAFE-TILE DISTANCE 397.78px 284.84px tiles > 10 threshold 0.61 0.15 The 397.78px figure is why the bot stayed far away: the safety filter left nothing safe near the target, and 397px is our WORST range. Control: setting corridor=20 wall=30 reproduces the original baseline exactly. CEILING, honestly: even at corridor 0 / wall 0 only ~40% of ticks are band-weightable, so no constant tweak fully unlocks the range weighting. Ram unification: the ring mover takes a `band` field; ramming becomes just `band=(0,50)`, so there is one movement engine. The `tfil` path is unchanged. Observability: magenta annulus at the band edges, candidates tinted by weight, chosen tile marked; one `[tfil_ring]` log line on change (now including corridorHeat/wallHotness). Guards: test_tfil_ring_weights 24/24 (new, pure, no battle), test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41. UNVERIFIED: the mover's live effect. It has not been run in a battle yet. |
||
|
|
2daa519e15 |
TFIL heat field: the safety model keeps us ~400px away, which is our worst range
Offline diagnostic driving the REAL TFILModule.computeMove over the committed
DrussGT fixtures (re-derived field matched the module's own m.lava bit-for-bit,
max diff 0.000e+00). Answers "is the heat map too hot, and are the corridors to
blame?" - the user's suspicion after watching a GUI run stay far away.
PRIMARY FIXTURE (tr_drussgt_vs_modularbot, 20,026 ticks / 15 rounds):
Field saturation
tiles == 0 22%
tiles > 0 78%
tiles > PathDangerThreshold(10) 61% (worst tick 91%)
median / p90 / max lava 20.06 / 44.53 / 77.51
> 10 with NO bullets at all 44% <- wall radiance + pillar alone
early/mid/late frac > 10 0.61 / 0.63 / 0.60 (saturated from tick 0, not degrading)
Safe pool - THIS IS THE KEY NUMBER
inside-hull tiles/tick 173.7
safeTiles/tick 15.19
ticks with ZERO tile passing the filter 58.5% (2-tile promote fallback used 59.0%)
ticks where the ring weighting is enabled (pool >= MinRingPool=4) 39.1%
ticks with >=1 safe tile in the 100-200px band 11.84%
MEAN DISTANCE TO THE CLOSEST SAFE TILE 397.8 px
ticks both pool>=4 AND band present ("band-weightable") 10.79%
So the safety filter leaves nothing safe near the target: the closest safe tile
averages 398px away. Our measured hit rate is 27.1% at 100-200px and ~5% at
450px, so TFIL's danger model structurally parks us at our worst range. This -
not only the env-var issue - is why the bot stays far away.
Per-source attribution (share of total lava / of the over-10 set)
wall 60.56% / 56.39% <- saturates the RAW field
corridor 25.43% / 21.35% <- blocks the BAND
pillar 7.44% / 5.06%
bullet_aura 2.22% / 1.47%
enemy_core 1.97% / 0.62%
bullet_core 1.17% / 0.33%
enemy_aura 1.21% / 0.95%
Note CorridorHeat=20 is TWICE PathDangerThreshold=10, so a single corridor can
poison a path on its own; WallHotness=30 with WallRadiance=10 puts the outer two
tile rings at/over threshold by themselves (38.6% of all tiles).
Counterfactuals (shipped constants NOT changed) - band-weightable ticks
corridor 20 (shipped) 10.79% band-safe 11.84% pool 15.19
corridor 10 16.47% pool 30.54
corridor 5 23.80% band-safe 24.21% pool 43.93
corridor 0 38.74% pool 62.14
wall 30->10 only pool 15.19 -> 24.80, band unchanged (12.31%)
corridor 5 + wall 10 26.45% band-safe 26.64% pool 85.04
Reachability - NOT the blocker
band inside the 50-tick reachable hull 64.94% of ticks
0-300px inside hull 90.34%
So the band is reachable 65% of the time but SAFE only 12%: the 53-point gap is
heat, not hull geometry.
VERDICT: heat saturation is the real blocker; the WALL is the largest raw-heat
source but the CORRIDORS are the band blocker (removing them multiplies
band-weightable ticks 3.6x, while taming walls leaves the band unchanged).
Even at corridor=0/wall=0 the band is weightable only 40% of ticks, so no
constant tweak fully unlocks the range weighting - the safe set against DrussGT
rarely reaches 100-200px at all. Recommended (NOT applied): CorridorHeat 20->5
and WallHotness 30->10, to be validated by a live A/B.
Adds common_libs/tests/measure_tfil_heat_field.nim (offline, no shipped file
touched; both movers byte-identical).
|
||
|
|
391318a7bd |
cornering/ramming premise REFUTED on three independent measurements
Hypothesis (user's): pushing an enemy toward a wall makes it predictable, which
both enables a ram and raises our gun hit rate. Measured offline over the
committed DrussGT fixtures using REAL server event attribution (an events
sidecar survived: tools/robocode_shim/evidence/tr_drussgt_vs_modularbot.events.json,
2534 fires / 229 hits; per-round event tick joins the fixture global tick at
global = round.startTick + tick - 2, verified exact over all 2534 fires).
M1 - cornering does NOT raise hit rate.
REAL attribution, all 1134 ModularBot shots, bucketed by DrussGT's distance
to the nearest wall at fire time:
<=30px 69 shots 5 hits 7.25%
30-60 336 18 5.36%
60-120 555 26 4.68%
120-250 172 11 6.40%
>250 2 0 0.00%
TOTAL 1134 60 5.29%
Adjacent(<=60) 5.68% vs Open(>60) 5.08%, z=+0.435 -> NOT significant.
Per-round ranges fully overlap (adjacent 0-18.2%, open 0-9.6%).
Virtual per-gun within-gun check: most guns neutral-to-negative; only DecayGF
favours it. Across all 10 fixtures every one of 12 guns scores LOWER adjacent
(range-confounded, directional only).
M2 - a wall-adjacent enemy is LESS predictable, not more.
30-degree tolerance, uniform chance 16.7%, adjacent vs open:
keep-direction (1 tick) 93.35% vs 95.88% z=-13.82
turn-persistence 87.6% vs 90.9%
constant-velocity err H=10 45.9% vs 25.7% (1.8x MORE deviation)
"move away from nearest wall" 1.1% vs 8.4%
"move toward centre" 0.4% vs 3.0%
Wall-adjacent DrussGT reverses more and deviates from constant-velocity ~1.8x
more. It does NOT flee the wall - it surfs perpendicular. Base rate of
wall-adjacency: 20.2% of moving ticks.
M3 - the ram is a near-zero-frequency opportunity against DrussGT.
Strict contact (<=36px): ZERO ticks in all 10 fixtures. Closest global
approach 39.1px. In the primary fixture (ModularBot vs DrussGT) the closest
approach was 118.7px - 0 ticks <=80px, 0 near-contact episodes, and ZERO ram
collisions in 15 rounds. Real ram collisions anywhere in the corpus: 2 total
(drussgt_vs_ramfire 1/20 rounds, tr_drussgt_vs_crazy 1/10), each a ONE-SHOT
0.6 energy to both bots, no sustained multi-tick stream.
CORRECTION TO AN EARLIER CLAIM: ram damage is 0.6 per CONTACT EVENT, not
0.6/turn sustained. The efficiency ratio (0.6 damage for 0.6 energy taken,
scored 2.0/pt) still beats firing, but the magnitude is 0.6 vs 16 for a p=3
bullet hit, and against DrussGT the frequency is zero.
CAVEAT (from the analysis): the fixtures capture DrussGT's NATURAL wall
behaviour, not an enemy being actively pushed into a corner by a rammer, so the
exact scenario is not directly represented. But M3 shows we never get close
enough to push in the first place - ModularBot's closest approach in 15 rounds
was 118.7px, so the <50px ram trigger has never fired against this adversary.
Adds two reusable offline instruments:
- measure_cornering_guns.nim (replays a fixture through the real VirtualTracker,
attributing each resolved virtual bullet to its fire-tick wall bucket)
- measure_cornering_ram.py (real-event join, predictability, ram base rate)
Neither edits offline_range.nim; the 12/12 deterministic-gun contract is
untouched and was not re-run (it requires a live battle).
|
||
|
|
07f6f3af3f |
Tsetlin gun: NO configuration adapts faster than random feedback
The user's goal was "a TM gun that can learn fast and generalize better". Swept offline over the real DrussGT fixtures (no live battles) by coordinate descent, one lever at a time, with a SHUFFLED-FEEDBACK CONTROL - a TM trained on randomised targets. That control is what settles the question. Final confirmation, 4 seeds each (~74,600 first-100-tick bullets per config): config EARLY(first 100) OVERALL Shuf_w3 (RANDOM feedback) 23.9% 20.0% win3_s1.1 (best real TM found) 23.7% 20.2% Shuf_w10 (RANDOM feedback) 23.1% 20.0% win3_st100 (prior job's edit) 23.0% 20.1% win3_off (TM correction ~= 0) 22.7% 20.2% def_w10 (shipped default) 22.1% 20.3% Linear (deterministic reference) 34.0% 24.3% The best real config beats the default early (23.7% vs 22.1%, non-overlapping per-seed ranges, z=+7.34, p=2e-13) - but its own SHUFFLED control scores 23.9%, i.e. HIGHER, z=-0.91, p=0.37. Random targets do at least as well. So the early gain is not learning. Per-lever screens were flat: TM_N_CLAUSES 25/50/100/200 all 23.0% early, completely flat; TM_N_STATES 4/32/100 all ~22-23% (unstable across seeds); TM_S mildly monotonic (lower better early); TM_T flat; TM_WINDOW_SIZE 2/3/10 all within noise of each other and of the shuffled control. Two further findings: - The TM-off ablation (correction ~= 0) scores 22.7%/20.2%, essentially the same as TM-on. The TM's correction is near-zero-mean noise; the gun's one-shot internal linear baseline accounts for its accuracy. - The TM gun is 10.3 pp behind Linear early and 4.1 pp behind overall. That deficit is in the BASELINE MODEL (LinearGun iterates flight time; this gun does not), not in the TM hyper-parameters. Tuning knobs cannot close it. Conclusion: do not tune TM hyper-parameters further. Either the input representation or the prediction target is what needs to change - the shuffled control shows the TM is not extracting target information beyond its baseline. Defaults left UNCHANGED (window=10/states=32/S=1.5/T=25/clauses=50); an uncommitted prior edit (window=3/states=100) was reverted as unsupported. Hyper-parameters are now compile-time overridable (-d:TM_WINDOW_SIZE=3 etc.) so future sweeps need no gun edit. NOT MEASURED: real hit rate vs DrussGT (offline only by design). The repo's own docs/gun_rack_analysis.md 2 reports the virtual metric is a sign-unstable ranker of real hit rate, so the comparison against "Linear 10.7% real" is not direct - whether the TM is competitive live is INFERRED-unknown, not measured. Guards: test_gun_harness 39/39, test_vbullet_metric, test_power_selection, test_tsetlin_gun, test_tm_pattern_learning all green. |
||
|
|
fb36a0a685 |
tracker: corpses do not exist - revert the fix and retire the workaround
The belief "BotDeathEvent never reaches ModularBot, so enemyTracker keeps dead
enemies alive forever" was written into a code comment and then believed twice.
It is FALSE. Measured in a 7-bot melee with a per-tick probe comparing
enemyTracker's alive count against the server's getEnemyCount():
metric 1.3.1 (20 rd) 0.35.5 (15 rd)
observed enemy deaths 83 68
...non-round-ending 83 (100%) 66 (97%)
ekBotDeath events DROPPED 0 0
max dispatch lag (turns behind) 1 1
phantom ticks 1 / 16,820 1 / 12,596
MAX CORPSE LIFETIME 0 ticks 0 ticks
victims still alive at round end 0 0
onBotDeath fires for every death, including non-round-ending ones. The
API-level event-drop mechanism IS real (test_event_drop_mechanism.nim proves
it: ekBotDeath is not in isCritical and MAX_EVENTS_AGE=2) - the bot simply
never falls far enough behind for it to trigger (max lag 1 turn).
Removed:
- reconcileWithServer + ReconcilePersistTicks/mismatchTicks/sawServerAlive
(uncommitted, and ON BY DEFAULT despite the premise being false). Its own
comment admitted a shorter window once KILLED A LIVE ENEMY ("it fired three
more times after the tracker marked it dead") - a latent mis-prune path
defending against a bug that does not exist.
- The radar's CorpseTicks=40 filter and the same-class age>60 filter in
recordRadarStats, both carrying the false comment. Removal changes no real
behaviour: buildState feeds the radar enemyTracker.allAlive(), so a dead
enemy never reaches computeScan.
Kept:
- The TR_TRACKER_PROBE instrument (default OFF), which produced the table above.
- test_event_drop_mechanism.nim - the drop mechanism is a genuine library
behaviour worth guarding.
- isAlive/aliveCount on the tracker.
Added: docs/tracker_death_events.md (the durable negative, so this is not
re-invented a third time) and test_enemy_tracker_death.nim (13 checks) in place
of the test for the deleted feature.
Guards: test_gun_harness 39/39, test_vbullet_metric 11, test_power_selection 3,
test_adaptive_radar 41/41, test_event_drop_mechanism 6, test_enemy_tracker_death
13, acceptance 12/12, ModularBot compiles.
|
||
|
|
c091bf3c34 |
harness: upgrade to server 1.3.1, keep 0.35.5 selectable, re-baseline
All prior measurements ran on server 0.35.5. The default is now the current 1.3.1 jar, with the legacy jar kept and switchable via TR_SERVER_JAR (no code edit). test_gauntlet_5bots.nim no longer clobbers a caller's TR_SERVER_JAR - it used to putEnv() unconditionally, so an override was silently ignored. RE-BASELINE (controlled RulesProbe battle, stationary bot, powers 0.1/0.5/1/2/3): dimension 1.3.1 0.35.5 verdict bullet damage per hit 0.4/2/4/10/16 identical SAME bullet speed (20-3p) within noise within noise SAME post-fire gun heat (1+p/5) identical identical SAME cooling 0.1/tick 0.1/tick SAME bulletDamage SCORE exactly 100/round 104..113/round DIFFERENT bulletKillBonus (20%) 20/round 20..23/round DIFFERENT LOUD FINDING - a SCORING rule changed, physics did not: 0.35.5 credits OVERKILL to bulletDamage (the killing bullet's full damage even past 0 energy); 1.3.1 caps it at the energy actually removed. Every 0.35.5 score is therefore inflated ~5-6%, and bulletKillBonus inherits the inflation. Gauntlet totals shift accordingly (SittingDuck 1936 -> 1800, WaveSurfer 1886 -> 1669). Consequence: score-based numbers recorded on 0.35.5 are NOT comparable to 1.3.1. Our gun A/Bs used real HIT RATE, not score, so those conclusions stand. Runner 1.0.2 (unchanged, no newer one on the box) is measured compatible with the 1.3.1 server. Note TrBattleCapture uses the runner's EMBEDDED server, which is 1.0.2 - so the capture path still runs an older engine than the gauntlet. Also re-ran acceptance_offline_vs_online on the new default: 12/12. |
||
|
|
18f778056b |
gun selector: hysteresis measured NEGATIVE, shipped at the lightest setting
Hypothesis under test: the selector chatters (~54 switches/100 ticks) and that chatter suppresses firing, so committing to the virtual-best gun should raise real hit rate. MEASURED AGAINST THE REAL DRUSSGT: it does not. setting switches/100t real hit % dmg/run shots/run no hysteresis 0/0 54.29 7.02% 217 244.8 light 10/0.05 1.95 6.22% 191 235.3 moderate 30/0.15 1.33 5.10% 156 233.9 aggressive 60/0.30 - 5.72% 175 242.5 (16 runs x 7 rounds per config except aggressive = 8; server-side events sidecar; permutation test baseline-vs-moderate p=0.002, baseline-vs-light p=0.18.) Hysteresis cuts chatter 28-54x but every variant fires slightly FEWER shots and deals LESS damage than baseline. Mechanism [INFERRED, consistent with docs/gun_rack_analysis.md 2/4]: the per-tick random tie-break among the tied band is a hedge, and hysteresis destroys it by committing to the virtual-best gun - which is not the real-best, because the virtual metric is a weak, sign-unstable ranker. The chattering was load-bearing. Shipped: GunDwellTicks=10, GunSwitchMargin=0.05 (GUN_SELECTOR_DWELL / GUN_SELECTOR_MARGIN) - the only setting within the baseline's run-to-run spread. GUN_SELECTOR_DWELL=0 GUN_SELECTOR_MARGIN=0 reproduces the pre-change selector exactly. Seam: VirtualTracker, which already owns the other selection state (fitness, the relative floor's peakRateRef), so the bot needs no new fields. bestGun and chooseFromFit stay pure/memoryless, which is why the existing random-tiebreak test needed no change. Guards: test_gun_harness 39/39 (33 original + 6 new hysteresis checks), test_vbullet_metric, test_power_selection, acceptance_offline_vs_online 12/12. |
||
|
|
68e0375be2 |
feat(radar): adaptive melee radar sweeps only the arc containing all enemies
Replaces melee_scan in the rack. melee_scan spun the radar at the 45 deg/tick cap unconditionally, so a full 360 deg revolution took 8 ticks and every enemy was scanned roughly every 8 ticks. The new module starts with the same full spin, and once it is SURE it has covered every enemy it sweeps back and forth over only the minimal covering arc of all enemy bearings. MEASURED, real melees via the bridge, per-enemy onScannedBot counts: 3-bot melee (2 enemies): 29.6 -> 61.1 scans/100 melee-ticks (2.06x) 4-bot melee (3 enemies): 36.0 -> 75.7 scans/100 melee-ticks (2.10x) Covering-arc widths observed: mostly <90 deg in the 2-enemy case, up to 240 deg in the 3-enemy case, so the gain shrinks as the arc widens - and at the ExitTrackWidthDeg=300 fallback it degenerates to exactly the old full spin, so there is no loss when narrowing would not help. TRADEOFF, recorded rather than hidden: a wider arc legitimately takes longer to traverse, so the freshness window costs 5-7 points (fresh<=16: 93-95% vs 98-100%) and more at fresh<=8 (75-76% vs 97-100%). More scans per enemy, at slightly staler individual fixes. DESIGN: acquisition spins 360 until every live known enemy was seen within FreshnessTicks=16 (two revolutions of slack), no new id appeared, and the live count matches getEnemyCount(); that must hold FreshStreakTicks=3 consecutive ticks. Tracking then bang-bang sweeps the wraparound-aware covering arc (350+10 -> 20 through 0) widened by MarginDeg=20 each end, at up to 45 deg/tick. Fallbacks return to acquisition: any stale enemy, any new id, or an arc >= 300 deg. Enter 270 / Exit 300 gives 30 deg of hysteresis so it cannot flap. Adds EnemyInfo.lastSeenTick (additive) so coverage is judged on staleness, not mere knowledge - without it an enemy that slipped behind the sweep would keep contributing its own stale bearing, which is self-confirming. The offline range now round-trips that field from the fixture 'lst'. COMPANION FIX, and it matters: the radar-mode switch used the TRACKER's known enemy count, so in melee the bot saw one enemy before scanning the second, locked to 1v1, and the melee radar never ran at all. Now uses getEnemyCount() (server truth), so melee mode persists until one enemy is genuinely left. 41 new unit checks (wraparound arcs, straddle at 0/360, single/empty enemies, the 45 deg/tick cap, every phase transition and fallback). melee_scan is kept but marked DEPRECATED; nothing in the rack imports it. Non-regression: 33 gun-harness checks, vbullet metric, power selection, and 12/12 offline==online acceptance all pass. |
||
|
|
c214abcfa8 |
fix(adversaries): repair four of the five sparring bots
The user suspected these were bugged. They were, and the verdicts are not uniform - three genuinely broken, one merely sloppy, one fine: - WaveSurfer: GENUINELY BUGGED, worst of the five. (a) The enemy velocity decomposition was sin/cos SWAPPED - enemyVx used sin and enemyVy used cos, while Tank Royale is 0 deg = East, CCW+, so it must be cos for X and sin for Y. Its linear-prediction gun was aiming at a reflected position. (b) The wall escape flipped strafeDir on EVERY tick the bot was inside the wall margin, so instead of turning away it flip-flopped in place: measured standing still (speed < 0.5) for 96.2% of ticks with a longest continuous standstill of 1398 ticks. Fixed with a hysteretic wall-escape selection plus a corner escape, dead enemyLastDir removed, and per-round state reset. AFTER, measured through the booter: rest 12.3%, mean speed 6.69, full speed 79.7%, longest zero run 18, perpendicular 0.845 / radial 0.012 - it now actually strafes. Gun sanity: lead error 1.0 px vs 106 px for head-on on a constant-velocity target; lead gun 45.8% hits vs 29.3% for head-on. - PatternMover: GENUINELY BUGGED. Real deadlock - it decremented its step counter by the REQUESTED amount while issuing setTargetSpeed(8), so against a wall the counter never reached 0, advanceStep never ran and it was stuck forever (309-tick standstill). Now counts down by ACTUAL distance/turn with a STALL_LIMIT watchdog and steers toward the arena centre. Standstill 309 -> 19 ticks; full-speed ticks 10.0% -> 28.4%. - OscillatorBot: GENUINELY BUGGED, milder. No wall handling at all, so it ground along walls 53.4% of ticks and could pin in a corner. Added wall steering that preserves the fixed 25-tick reversal cadence. Wall-band 53.4% -> 18.6%, mean wall distance 72 -> 119. - RandomMover: merely sloppy, not broken. Its turn intent saturated against the speed-dependent limit (18.4% of moving ticks clamped) and the fire gate was a very loose 10 deg. Now clamps to calcMaxTurnRate and fires within 3 deg. Saturation 18.4% -> 3.9%. - SittingDuck: FINE. Speed 0 for 100% of ticks, zero shots. Left untouched - it is a duck by design. Adds test_wavesurfer_velocity.nim, a direct assertion that the decomposition is cos/sin and explicitly NOT the swapped form (7 cases). KNOWN ISSUE, not fixed: RandomMover/PatternMover/WaveSurfer import tankroyale_botapi 1.0.1 and intermittently SIGSEGV in tankroyale_botapi/event_queue.nim:89 addEvent, freezing the bot for the rest of the battle. It reproduces on old and new code and never occurs for SittingDuck/ OscillatorBot, which import robocode_tankroyale_botapi 1.0.7. Migrating the three to 1.0.7 would likely fix it and is worth doing - it is a real reliability risk for these as sparring partners. |
||
|
|
4cd5618435 |
fix(test): repair the flaky offline==online acceptance; make the tie-break truly random
TASK 2 - THE FLAKY ACCEPTANCE TEST, root-caused. It was NOT a live/offline boundary race as suspected. The replay spawned gun 13 (TMSelect) while live has EnableTmSelector = false and never does. The shared VirtualTracker ring is ORDER-SENSITIVE, so gun 13's extra 4 bullets/tick shift the ring head and permute the per-tick RESOLUTION ORDER of every other gun. The learning guns append observations in resolution order, so their predictions shifted and produced small hit deltas that moved between runs. Evidence: the first KNN divergence was at rtick=174 with the SAME resolution set merely reordered (live ft133,138,139,142,148,150,151 vs offline ft150,151,133,138,139,142,148); after closing gun 13's ready gate offline the live and offline KNN traces became BYTE-IDENTICAL (diff empty, 904/904 lines). Fix: mirror the live rack in the replay. No tick exclusion, no tolerance loosening. Stability: 5/5 consecutive runs now report 12/12 exact, each with enemyDied=true - the death boundary is included, not excluded. The proof is now real rather than a lucky run. TASK 1 - the tie-break was not random. randomize() was only reached incidentally through initTsetlinGun(), so a rack without Tsetlin had a fixed rand() stream and ties always resolved the same way across process restarts. Added seedSelectorRng() after gun construction, honouring GUN_SELECTOR_SEED. Evidence: unseeded, 6 separate processes gave different pick sequences; with GUN_SELECTOR_SEED=42, 3 processes gave identical sequences. TASK 3 - PRUNING DOES NOT HELP; keep the full rack. 15 PAIRED runs per variant vs DrussGT, 8 rounds, identical seeds: baseline 3238 shots 6.18% (events 6.16%) 200 dmg/run Tsetlin disabled 3522 shots 5.76% (events 5.71%) 197 dmg/run Tsetlin+Displace 3478 shots 5.46% (events 5.37%) 183 dmg/run Paired permutation tests: -0.34pp p=0.57 and -0.70pp p=0.21. Per-run distributions completely overlap (baseline range [2.68, 10.00]; 15/15 and 14/15 runs inside it). A Crazy control showed no separation either. So removing the measured-worst real performers is neutral-to-slightly-negative, and with sd ~1.8pp a definitive claim either way would need far more runs. CORRECTION TO A CLAIM I MADE: the 'virtual metric is INVERTED' finding does NOT reproduce. Job-24 measured Spearman -0.374; this job measures +0.335 over the same 13 guns with a different but equally defensible aggregation. Two opposite signs means the correlation is NOT robustly negative - it is WEAK AND SIGN-UNSTABLE. The honest statement is that virtual hit rate is a poor ranker, not an inverted one. The docs assert the inversion and need correcting. Also adds per-process GUN_STATS_PATH/GUN_SHOTLOG_PATH so concurrent A/B runs do not clobber each other, and an env-gated GUN_RACK_DISABLE for rack A/Bs. All default behaviour is unchanged when the env vars are unset. |
||
|
|
57b2ac3849 |
feat(guns): scale-aware power selection (+52% damage); TM classifier gun built, measured, DISABLED
TASK 2 - power selection, a clear win. bestPower used an ABSOLUTE MinHitRate = 0.40 bar. Measured per-bin virtual rates (rolling-100 fraction) show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0) even where higher bins were comparable: Linear p1.0 44% p1.5 39% p2.0 30% p3.0 29% old bin 0 -> new bin 3 Accel p1.0 44% p1.5 40% p2.0 26% p3.0 29% old bin 1 -> new bin 3 Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12% old bin 1 -> new bin 2 Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless FRACTION of the gun's own best bin rate). 13 of 14 selections now pick heavier bullets. Real effect vs DrussGT (8 rounds x 3 runs): hit rate unchanged (7.56% -> 7.47%) but damage dealt +52% (157 -> 239 per run) and rounds end faster. Same accuracy, half the shots, half again more damage. TASK 1 - the TM pattern-classifier gun does NOT earn its slot. It was built as a mixture of experts with a corrected-Granmo TM as a multi-class gate over HeadOn/Linear/Circular/WallBounce/Accel, labelled by which expert's prediction was closest to the actual enemy position (an exact, supervised, per-shot label - no delayed credit). Offline it loses to the best of its OWN experts on essentially every fixture, and against DrussGT it cost real performance: baseline (path+relative) 7.56% real hit rate, damage 157 + power fix 7.47%, damage 239 + power fix + TM gun 5.59%, damage 133 The gun was selected on 806 ticks and fired 24 real shots at 4.2%. So the tree ships with EnableTmSelector = false: code and wiring kept intact for re-enabling, but it is not in the active rack. Worth recording from the clause dump: the gate DOES latch onto meaningful structure. On energy-threshold-turner, HeadOn's clauses key on the energy bits (the rule's own driving variable) while Circular keys on distance/velocity. So the TM is learning something real and interpretable - it simply cannot beat 'always pick the best expert'. Root cause (INFERRED): the closest-expert label is noisy because several experts are near-tied, and under the path metric the winner varies by power bin while the gate sees one shared per-tick input, so a one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit. (Zero-padding the 2-frame window was tried first and saturated every clause at 256-755 included literals; alternating the two real frames fixed that.) Also factors the corrected feedback into an exported tmLearnDir and exports the encoding/TM primitives; the Tsetlin tests still reproduce the documented mean=13.8 included literals, so the refactor is behaviour-preserving. Verified: 33/33 guard checks, tsetlin tests green, metric checks green, new power-selection guard green (13/14 selections change; relative bar still picks bin 1 and not bin 3 for a [30,25,12,5]% profile), 12/12 offline==online acceptance under the shipped default. |
||
|
|
dea4dcb574 |
feat(gun_harness): scale-aware selector thresholds; default = path + relative
The selection thresholds were calibrated for a rate scale that does not exist.
MEASURED on an exact offline replay of a fogged live WorldState vs DrussGT
(1397 selection ticks), the 0.10 absolute floor fires on 53.0% of point-metric
ticks and forces HeadOn, which has a REAL hit rate of 2.0-4.4% - worst or
near-worst of 13 guns. HeadOn's selection share: 69.1% (abs+point) -> 43.5%
(rel+point). My earlier claim that the floor fires ALWAYS is REFUTED - it is
53%, because bestRate is a max over gun x bin and a >=50-sample bin
occasionally clears 10%. The mechanism is confirmed; the literal statement was
not.
Scale-aware mode (GUN_SELECTOR_MODE, absolute|relative, default relative):
RelTieMargin = 0.20 dimensionless FRACTION of bestRate, replacing the
fixed 2pp band so the band scales with the metric
FloorPeakFrac = 0.25 the floor fires iff bestRate < 0.25 * peakRateRef,
SelectorWindow = 256 where peakRateRef is the field-best rate over the last
256 selection ticks - keeping the original 'don't trust
a collapsed field' purpose but only when the field is
bad RELATIVE TO ITS OWN RECENT BEST, and counting only
guns with >= MinObsBeforeCompete samples so cold-start
100% spikes cannot pin HeadOn
also pools the rate over power bins instead of taking the max over bins, so
one lucky bin no longer wins
absolute mode is preserved byte-for-byte for rollback.
A/B vs DrussGT, real server hit rate, 3 runs x 10 rounds per config, one frozen
binary:
absolute+point 3.66 / 2.45 / 5.01 pooled 3.76%
absolute+path 7.55 / 8.21 / 6.83 pooled 7.57%
relative+point 7.66 / 6.18 / 5.79 pooled 6.59%
relative+path 7.15 / 7.55 / 6.90 pooled 7.21%
absolute+point is SEPARATED from all three (p < 0.0001); the other three
OVERLAP each other (p = 0.18-0.64). So the METRIC is the dominant lever and
under path the two threshold models are statistically tied.
DEFAULT SET: metric = path, thresholds = relative. absolute+path was nominally
0.35pp higher but indistinguishable (p = 0.64); relative is the principled
scale-aware fix, is the only model that works under BOTH metrics, and prevents
the point-metric catastrophe if anyone switches back. Shipping absolute would
ship the accidental side-effect this work exists to remove.
STILL NOT SOLVED: the selector remains only a moderate ranker.
Spearman(virtual rank, real rank) is 0.52 for the winning config, 0.36 pooled
for path and 0.04 for point - and it is INCONSISTENT across run sets. The
metric switch won by de-selecting HeadOn, not by ranking guns better. That is
the next problem.
TASK B, report only: do NOT drive selection from raw real hit rates yet.
Only the selected gun fires, so unselected guns get near-zero real shots
(GuessFactor 20, Linear 24 vs HeadOn 733); noise is fatal (n=470 at p=10% gives
+/-2.8pp, most guns n<200 gives +/-5pp+ across a 3-15% spread); and real rate is
conditional on when the gun was selected. A blended signal with forced
exploration and shrinkage is defensible in principle but needs thousands of
shots per gun across many battles. Real rate is best used OFFLINE as the
evaluation metric - which is exactly what this A/B did.
RELATED BUG FLAGGED, not fixed: MinHitRate = 0.40 in bestPower is on the same
wrong scale - no bin ever clears 40%, so once every bin has data, power
selection falls back to bin 0 (power 1.0) late in a round.
Verified: 33/33 guard checks, 11/11 metric checks, tsetlin green, 12/12
offline==online acceptance under the shipped default, run_range rc=0 over 20
fixtures. Adds analyze_selector.nim to measure floor/tie/bestRate/HeadOn-share
per config on any fixture.
|
||
|
|
3b5d70b7c3 |
feat(gun_harness): runtime metric switch + A/B proving the point metric mis-selects
Adds GUN_VBULLET_METRIC (point|path, default point = unchanged behaviour) so the virtual-bullet hit model can be selected at runtime with no rebuild. Both the live tracker and the offline replay read the same value, so the 12/12 offline==online acceptance holds under EITHER setting (verified for both). A/B AGAINST THE LIVE BOSS, real server-side hit rate as ground truth, 5 battles x 12 rounds per metric on one frozen binary: point 4660 shots / 219 hits = 4.70% (per-run 3.16-5.53) path 4834 shots / 359 hits = 7.43% (per-run 6.55-8.24) The distributions DO NOT OVERLAP: path's worst run beats point's best run. +2.73pp, +58% relative, z = 5.56, p < 0.0001. Range distributions were identical (~460-478 px), so this is not a range confound. MECHANISM - and this is the important part. The gain is SELECTION, not better gun learning. Under the point model every gun's virtual rate is compressed into 0.6-4.4%, so HeadOn sits inside the 2pp tie margin and takes 72.6% of selection ticks / 76.9% of shots - while HeadOn is 11th of 13 by REAL hit rate (2.3%). The path model widens the band to 4.7-13.7% and ranks HeadOn 10th, so its shot share falls to 35.9% and Pattern/Accel/WallBounce get picked instead. Counterfactual: applying the point model's per-gun real rates to the path model's shot mix yields 7.65%, i.e. essentially the whole observed gain. So the selector, not the guns, is where the win lives. PER-GUN REAL HIT RATE vs DrussGT (path mix, the answer to 'which guns are worth keeping'): WallBounce 10.8, Pattern 10.5, Accel 10.0, Displace 9.3, Circular 9.2, AvgLead 8.5, KNN 5.7, StopShot 5.2, GuessFactor 3.7, Tsetlin 2.9. Per-gun N is small (hundreds of shots) so single-gun ordering is indicative, not definitive. TWO CAVEATS, recorded because they undercut a naive reading: 1. One adversary. DrussGT is a wave surfer and HeadOn is genuinely bad against surfers, so part of this may be matchup-specific. 2. The path model is NOT a better general ranker. Spearman(virtual rank, real rank) is 0.52 under point vs -0.04 under path. It wins by accidentally fixing HeadOn's mis-rank, not by ranking guns better. A more durable fix is to address the selection logic directly - which is the next job. Also adds a focused guard test (test_vbullet_metric) covering parsing/default, a receding-target point-miss/path-hit, a perpendicular-target path-miss, and replay determinism. Verified: 33 guard checks, 12/12 acceptance under both metrics, tsetlin tests green, range 34.3% (point, unchanged) / 50.8% (path). |
||
|
|
d5061ee215 |
test(range): restore the 12/12 offline==online proof; measure TM clause readability
Task 1 - the acceptance proof was unrunnable because RecordWorldState was a
compile-time const set to false. It is now a RUNTIME switch
(let RecordWorldState* = existsEnv("TR_RECORD_WORLDSTATE")), default OFF, so
ordinary runs write no fixture, and acceptance_offline_vs_online.nim enables
it for the battle it spawns and clears it afterwards. Restored and run twice:
12/12 deterministic guns match exactly (128-tick and 546-tick battles), with
Tsetlin reported separately as stochastic. Both nimble build variants clean.
Task 2 - does a compact encoding turn the TM's 99.35% into a READABLE rule?
Measured across window sizes (fixed seed, no tuning):
frames TEST acc eff.lits/clause firing clauses counterfactual low/high/mean
10 99.35% 152.8 37 100/24/62.4%
3 95.94% 54.9 35 96/20/58.6%
2 99.48% 39.6 38 95/25/60.7%
1 98.30% 19.2 45 100/24/62.3%
So 2 frames is strictly better than 10 on BOTH axes: +0.13 accuracy for 4x
smaller clauses. The 3-frame dip is non-monotonic and left unexplained rather
than smoothed over.
A readable rule WAS partially recovered. Five clauses carry the exact Gray
form !g10 ^ !g9 ^ !g8; g10 is inert in this data, so the effective rule is the
2-literal proposition !g9 ^ !g8, i.e. energy < 25.6. That is a genuine
threshold in readable propositional form - but at 25.6, NOT the labelled 30,
because 256 is a power-of-two Gray boundary expressible in two literals while
300 needs a longer conjunction. The TM found the nearest SIMPLE threshold.
The honest caveat: that threshold is not the ensemble's decision mechanism.
The counterfactual follow rate (high 24%, mean 62.3%) is statistically
identical at 1, 2 and 10 frames, so compactness did not make the model read
energy - its vote is carried by co-occurring bearing/velocity/heading/wall
literals. Also identified: clauses containing all 11 Gray energy bits are
satisfied at exactly one raw value (50, the dataset floor), so they are
'energy has hit the floor' detectors, not thresholds.
Methodological fix worth keeping: the earlier single-frame counterfactual
wrote energy into all 10 frame slots including the zeroed ones, reviving dead
clauses and producing a spurious 2% high-follow rate. setEnergyFrames now
rewrites only the exposed frames; the corrected figure is 24%.
|
||
|
|
89370008da |
fix(tsetlin): make the TM actually learn - saturation 714 -> 13.8 literals/clause
The gun has never contributed anything: Tsetlin.vHits was byte-for-byte equal to Linear.vHits in every measured round of every run, because its learned correction was always exactly 0. Six diagnosed defects fixed, plus one that was required to make the first one work: 1. Type I now conditions on the clause output. It previously rewarded included true literals unconditionally, omitting Granmo's (c=0, lk=1) -> toward Exclude counter-force, so true literals ratcheted toward Include forever. This was the root cause of the saturation. 2. Type II was unreachable dead code: its guard required cOut==1 AND lits[lit]==0 AND st>0 (included), but cOut==1 guarantees every included literal is 1. Its direction was wrong too - it should increment EXCLUDED false literals when the clause fires. 3. Resource allocation restored: Granmo's (T - clip(v,-T,T))/(2T) target replaces |error|/(2*RESID_MAX); TM_T was only an output normaliser. 4. Label baseline fixed - the factor-2 shrink. predX = linearX + cx, so the label was delta - cx while the learner's output IS cx, giving error = delta - 2cx and a fixed point of cx = delta/2: HALF the needed correction even with perfect feedback. TmTrace now stores linearX/linearY and training uses delta. 5. Hits no longer zero their label (a hit means |miss| < 18px, not 0). 6. The enemy-energy feature was duplicated - tmEncodeFrame passed state.selfEnergy with a stale comment claiming enemyEnergy was absent, while WorldState.enemyEnergy exists. Enemy-energy rules were literally unrepresentable. 7. REQUIRED EXTRA: tmEvalClause now implements Granmo Eq. 6 - an all-Exclude clause outputs 1 during learning and 0 during classification. Without it, fix #1 deadlocks every clause at empty. MEASURED EFFECT (energy-threshold-turner fixture, seed 1): mean included literals per active clause 714.0 -> 13.8 active clauses 100/100 -> 53/100 nonzero corrections 8/764 -> 708/764 Tsetlin virtual hits (Linear = 27/400) 27/400 -> 69/400 Divergence achieved: offline on 7/8 fixtures, and in a live gauntlet (RandomMover: Tsetlin 199/1200 vs Linear 288/1200, vDropped=vStarved=0). Tsetlin now LEARNS but is not yet competitive with Linear - the regression head is untuned, flagged as follow-up rather than claimed as a win. Also ignores compiled test harnesses that have no file extension, which the existing '**/tests/test_*' rule misses. |
||
|
|
974528d5cf |
feat(gun_harness): offline gun range, proven equivalent to live play
Gun evaluation previously required a full end-to-end battle (Java server + battle runner + websocket IPC to 2 bot processes, 50 rounds, ~3.4 min) and yielded only ~300-900 REAL shots across 13 guns -- far too few to rank guns, which is why tuning needed many repetitions. VirtualTracker is already a pure function of (WorldState stream, gun list); the only reason it needed Java was where WorldState came from. So the range replays a seq[WorldState] through the SAME tracker: offline and online scores are the same metric by construction, not an approximation. ACCEPTANCE TEST (the point of the whole thing): record one live round, replay it offline, compare per-gun virtual hit rates. 12/12 deterministic guns match EXACTLY, reproduced twice. Tsetlin is compared separately because tmLearnOne calls rand(). Getting to 12/12 exposed two real ordering quirks in the live loop: run() calls go() before the aim/fire block, so tickBullets resolves against the NEXT tick's scan while the prediction used the previous one; and if the target dies during that go() the final tick's spawn+resolution is skipped entirely. The recorder emits an end marker for the second case. The 5th (selected-gun) predict call was verified to be a no-op. Measured cost: 8 fixtures (1770 ticks, ~92k virtual bullets, 13 guns) replay in 2.9 s, ~32k virtual bullets/s -- roughly 70x faster and 100x more samples than a live gauntlet. Also adds a per-tick WorldState recorder behind const RecordWorldState (default off, mirrors the ShotLog idiom) which records the state the bot ACTUALLY builds, staleness included, rather than true positions -- recording the latter would hand the guns perfect information and produce flattering scores. 9 new guard checks (33 total, all passing), including fixture round-trip, replay determinism, stationary->HeadOn 100%, constant-velocity->Linear>HeadOn, and the energy-threshold turner crossing at t=41. |
||
|
|
3c90a5941d |
feat(selector): range-aware firing gate fitted to 2611 measured shots
Measured, not assumed. With the gate temporarily opened to 20 deg, every real shot was logged (tick, angle error at fire time, distance, power, hit) across 3 gauntlets: 2611 shots, 57.3% aggregate. Findings: - The geometric cone atan(BotRadius/d) is directionally confirmed but a WEAK lever: even at 0.0-0.1 deg error the hit rate at 400-600px is only ~53-57%, because PREDICTION error dominates alignment error. - Real effect of tightening the gate: 57.9% -> 68.0% aggregate hit rate (fixed 0.1 deg), not the 76.9% previously reported -- that was a high-variance draw (per-rep 62.8/66.4/77.2%). - The shipped range-aware gate (SafetyFactor 0.6) does NOT beat the fixed 2.0 deg gate on hit rate (55.8% vs 57.9%, ~1.5 sigma, inside noise). It fires 22-28% more shots and therefore lands more total hits (~509 vs ~434 per rep). No per-adversary score delta exceeded the 300-point run-to-run noise band, so no config is demonstrably better on score. Shipped anyway because it is strictly more expressive (a fixed threshold is the special case), tunable from one const, and physically motivated, but the honest verdict is recorded in-code: the gate is not the bottleneck. AimThresholdDeg is removed; shouldFire now takes distPx. Degenerate or NaN distance falls back to the ceiling rather than dividing by zero. Also adds a per-shot logger to ModularBot behind 'const ShotLog' so the measurement above is reproducible, and 10 new guard checks (24 total, all passing) covering monotonicity, clamping, formula, perfect alignment, gross misalignment and degenerate distance. Cross-checked against the server source: the gun fires BEFORE the turn is applied, so the logged angle error is the true departure error, and fireAssist auto-aim is off (unset by the Nim API and forced false by setAdjustRadarForGunTurn). |
||
|
|
e53690036b |
fix(guns): speed-sensitive caches, dead stop-shot branch, exact TM trace pairing
Four guns cached a whole prediction per tick while predict() is called once per power bin, so every bin after the first (and the real fired shot, which shares lastState) reused the power-1.0 lead. Fixed by caching only the speed-INDEPENDENT derived state and recomputing the lead per requested speed: - stop_shot: also fixes prevSpeed being written before it was read, which made abs(speed) < abs(prev) permanently false and the entire stop-prediction branch unreachable (it was just Linear). - displacement: the cache key included bulletSpeed, so the guard missed on all four bins and the 15-tick window advanced ~4x/tick, making the inferred velocity ~4x too small. - averaged_lead: tick cache removed outright. pattern_matcher: split into speed-independent match+path and per-call lead. FeedbackEvent gains fireTick/powerBin (additive; only virtual_bullets constructs one) so guns can pair feedback to the exact shot instead of guessing by coordinates. tsetlin uses it: traces are now keyed exactly by (fireTick, powerBin) with a 1024-slot ring, and the 10-frame window shifts at most once per tick (it was shifting ~4-5x/tick, so isWarmedUp tripped after ~2 ticks). KNOWN INCOMPLETE: tsetlin still does not diverge from Linear in battle. The two named bugs are fixed (a 600-tick sim shows trainedShots=2141, traceMisses=0, and a fixed-input probe converges to a 9.6px correction), but the TM's clause feedback itself is broken: ~131 of 1740 literals end up included per clause, so its conjunction never fires. Sweeping TM_S, TM_N_CLAUSES and a two-branch Type-I update did not change the correction from 0. Needs a real TM fix or removal, not another bug fix. First-ever guard tests for the gun selector: common_libs/tests/ test_gun_harness.nim (14 checks, headless, no Java). There were none before, which is how six broken guns survived a full analysis cycle. Against the previous HEAD, 5 of these checks FAIL - that is the regression guard. |