8dd9b3b3b5f81f5bbcd06ca9f1cfb118dbbc57dd
7 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f58d65d2e8 |
TMComposites gate: per-gun confidence faithful for 3 guns; no pair composes
Adds a per-sample intrinsic-confidence field (GunPrediction.confidence, threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF, KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that reproduce the paper's Figure 2 per gun and its Eq-8 composite. Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun): - FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak). - GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001). - No pair of guns specialises complementarily: the same gun dominates both high-confidence slices in every pair. - Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern 20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses. Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of competence is real but ~2pp short. Offline veto: design is dead. See docs/tmcomposites_gate.md. |
||
|
|
4657fe715e |
wave pairing: 36-58% of GF/DecayGF/KNN learning samples were MISLABELLED
The audit inferred (from code) that GF/DecayGF/KNN pop the OLDEST wave on resolution, while under bmPath bullets leave the arena in NON-FIFO order - so an outcome could be attached to the wrong wave. It also noted that `starved=0` does NOT rule this out. Both halves are now MEASURED. MISPAIRING RATE (10 DrussGT fixtures, real VirtualTracker, 344k resolutions/gun): gun bmPath mispair label err bmPoint mispair label err GuessFactor 36.48% 19.39% 18.24% 7.62% DecayGF 36.85% 19.52% 20.57% 8.64% KNN 57.91% 27.63% 29.75% 11.58% (starved = 0 everywhere, exactly as the audit predicted) So ~1 in 5 GF/DecayGF learning samples and ~1 in 4 KNN samples carried a WRONG guess-factor bin. This is a material corruption of the learning signal. FIX: the same fireTick-keyed ring scheme `tsetlin.nim`/`tm_selector.nim` already use - `slot = (fireTick*4 + bin) mod 1024` (period 256 ticks, longer than the ~91-tick max flight), looked up by exact key. Public interfaces unchanged; added `waveResolved`/`waveMispaired` integrity counters. AFTER: mispaired = 0 and starved = 0, both metrics, all three guns. EFFECT ON HIT RATE: SMALL AND NOT SIGNIFICANT. bmPath 4000 samples/gun: GuessFactor 23.20% -> 23.02% (-0.18pp, per-run sign-flip p=0.750) DecayGF 23.80% -> 24.25% (+0.45pp, p=0.625) KNN 18.27% -> 18.80% (+0.53pp, p=0.547) bmPoint: +0.05 / +0.33 / -0.15pp, p = 1.00 / 0.50 / 0.50. Per-run ranges overlap almost completely. A bullet-level z-test is anti-conservative (bullets within a fixture share a trajectory) and its KNN p=1.9e-16 cannot be trusted given ~10 effective independent runs. PLAIN READING: this is a CORRECTNESS fix, not a measurable hit-rate win. It removes a 36-58% mislabelling of the learning signal; the point estimates move by at most ~0.5pp, within run-to-run noise. Stated plainly rather than oversold. A REGRESSION IT CAUGHT IN ITSELF (and this explains the SIGSEGV another job saw and correctly attributed to a concurrent knn_gun.nim rewrite): the first implementation put an inline `array[1024, KNNWave]` (~100KB) inside each gun, which overflowed the default 8MB stack and made `test_power_selection` SIGSEGV. Causation was proven by stashing only the three gun files (test passed), then fixed by making the rings heap-backed `seq`. Verified: `test_power_selection` 3 PASS on the default stack, and zero inline `array[1024]` remain. Guards: test_wave_pairing 17 (new, pure), test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28. ModularBot compiles. Adds audit_wave_pairing.nim and compare_pairing.nim. |
||
|
|
e2ca2fc7d8 |
fix(guns): recover the DrussGT regression with a radial-fraction range blend
The previous fix (learn the residual against a constant-velocity base) was structurally right but cost us on real wave-surfing movement: GF 108 -> 55, KNN 101 -> 74 on the classic DrussGT captures. Root cause: the linear base is a poor model for a surfer, so the residual histogram is noisier than the old total-lead histogram. FIX: blend the RANGE between a radial-only forecast and the geometric one by radialFrac (the fraction of recent per-tick motion that is radial), keeping the constant-velocity bearing. dist = radialDist + rf*(linearDist - radialDist). New VelocityTracker in common_libs/guns/lead_forecast.nim; the window default is 32 and results were identical at 16 and 40, so it is not tightly tuned. Nine candidate bases were measured and rejected WITH NUMBERS rather than by argument, which is why I trust the winner: velocity scaling 0.8 recovers DrussGT but destroys wall-bounce 241 -> 20 radial-only range excellent DrussGT, wall-bounce 241 -> 140 short-window averaged vel worse than both bases outright hard reversal/speed gates help DrussGT, lose nothing, but weaker than blend radial-fraction blend best on BOTH <- shipped Result (hits per 2000; classic-5 = classic DrussGT captures, tr-5 = the new closed-loop TR captures, synth-10 = the rest): base classic-5 GF/DGF tr-5 GF/DGF synth-10 GF/DGF current(prefix) 108 / 108 41 / 39 1302 / 1302 linear(postfix) 55 / 76 9 / 4 2702 / 2692 BLEND 171 / 100 86 / 87 2717 / 2703 Strictly better than both on classic-5 GF and on every synthetic bucket. The one figure below the old base is classic-5 DecayGF (108 -> 100, -8/2000, within noise) and that is stated plainly rather than hidden. TASK B - enemy energy in learners. KNN gains an 8th feature, enemyEnergy/100, on a FIXED [0,1] scale (not min-max) because threshold behaviour keys off absolute energy. Honest result: it is NEUTRAL on the target fixture (77 vs 77) and roughly neutral in aggregate. The base change, not the feature, moved that fixture. Tsetlin already encoded enemyEnergy and now scores 88/400 on energy-threshold-turner against Linear's 43/400 - a 2x margin, which is the 'can a TM learn a high-level pattern' question answered in gun form. TASK C - is the virtual-bullet metric itself faithful? Quantified: scoring the bullet's PATH against BotRadius instead of the single point at aim distance raises every gun by +31% (GF) to +86% (HeadOn), so the current model is PESSIMISTIC, and it RE-RANKS materially: Linear 9th -> 6th, AvgLead 7th -> 3rd, GuessFactor 4th -> 9th, DecayGF 6th -> 12th. The 12/12 offline==online acceptance still holds under the path model (verified with a temporary env hook driving both sides), so no red flag. VERDICT: do NOT switch. The point model is the standard virtual-bullet PREDICTION-ACCURACY fitness - the bullet must arrive at the predicted point at the right time - while the path model measures hypothetical hit chance against a target that never dodges, and in open-loop fixtures it over-credits directional guns (HeadOn 35% on DrussGT, 100% on constant-velocity) for exactly that reason. The models differ materially but the current one is not shown to be unfaithful FOR ITS PURPOSE. Because the metric drives gun SELECTION, this is now being A/B'd against real hit rate versus the live DrussGT boss, which is the only ground truth we have. Verified: 20 fixtures 35636/104000 (34.3%); 33 guard checks; 12/12 acceptance; tsetlin tests green; live gauntlet 5/5. |
||
|
|
7f706e5b14 |
fix(guns): GF family aimed at the wrong RADIUS, not the wrong angle
The entire GuessFactor family scored 0% on clean circular and wall-bounce trajectories. Two hypotheses were on the table and BOTH were wrong: - MEA range too narrow / edge clamping: REFUTED. Measured 0 clamped shots out of 837/849/957, required offsets peak at ~33 deg against MEA 28.1-46.7 deg, and the 8 in arcsin(8/bulletSpeed) is correct (it is the max robot SPEED, not the hit radius). Changing it to BotRadius=18 would have coarsened resolution for nothing. - Peak selection: REFUTED. A sweep of every constant GF value showed the ORACLE-BEST constant offset on the original gun was only 6% circular, 4% wall-bounce, 7.5% random-walk. No peak choice could have done better. The learning path was fine too: ~850-960 observations per fixture, 0 starved waves, well-populated histograms. REAL CAUSE: the GF family aimed at the FIRE-TIME distance. The virtual-bullet metric resolves a bullet at the AIM-POINT distance and scores that single point against the enemy's position on that tick, so with any radial target motion the bullet stops at the wrong radius and misses even with a perfect angle. Angle-only prediction is structurally unscoreable under this metric. FIX: give the GF family a self-consistent constant-velocity forecast as its base reference (new common_libs/guns/lead_forecast.nim, which iterates the flight time to the same fixed point circular.nim uses), so the histogram learns the RESIDUAL against that forecast and the aim point lands at the right radius. Applied to guess_factor, decay_gf and knn_gun. Same defect fixed in Linear: it did a one-shot dist/bulletSpeed extrapolation and never iterated its flight time. The oracle sweep proves the structural fix, independently of tuning: the best achievable constant GF moved 6% -> 20% (circular), 4% -> 57% (wall-bounce), 7.5% -> 49% (random-walk). MEASURED, all 15 fixtures: total 39.0% -> 44.4% (30399 -> 34654 hits). circular GF 6 -> 23, DecayGF 6 -> 21 wall-bounce GF 0 -> 60.2, DecayGF 0 -> 60.2 constant-vel GF 26 -> 100, DecayGF 26 -> 100, KNN 26 -> 100, Linear 87 -> 100 random-walk GF 0 -> 53, DecayGF 0 -> 52, Linear 24 -> 53 StraightLine GF 8 -> 77, DecayGF 8 -> 77 Non-regression: 33 guard checks pass, the range's 12/12 offline==online acceptance still PASSES, tsetlin tests green, live gauntlet 5/5. HONEST TRADE-OFF, recorded rather than hidden: on the 5 real DrussGT wave-surfing captures the GF family REGRESSES - GuessFactor 108 -> 55, DecayGF 108 -> 76, KNN 101 -> 74 hits per 2000. The linear base is a poor model for a surfer, so the residual histogram is noisier than the old total-lead histogram. Linear itself improved there (95 -> 105). The synthetic range and the live gauntlet both improved, and the structural bug is provably fixed, so this was judged worth the cost - but recovering the DrussGT regression is the next job, not something to wave away. |
||
|
|
0cc682152d |
fix(guns): per-bin wave queues unbreak GF/DecayGF/KNN learning; fix vbullet drops
Wave queues (guess_factor, decay_gf, knn_gun): predict() stored ONE wave per tick while onResult() popped one per resolved bullet (~4/tick), so the queue drained to empty within a few dozen ticks, ~3 of every 4 resolutions returned without learning, and the survivor paired with a same-tick wave (bearingDelta ~= 0) pinning the histogram at centre. PROOF: GF.vHits == HeadOn.vHits and DecayGF.vHits == HeadOn.vHits byte-for-byte in every one of 50 rounds — the guns had degenerated to HeadOn. Now each gun keeps a per-bin FIFO with an O(1) head cursor. At most one push per (tick, bin) so the fire site's 5th predict() call is a no-op, and onResult pops the oldest wave of its OWN bin via e.bulletPower. Aiming math untouched (it was already correct: 0 deg = East, CCW+). maxBullets 2048 -> 8192: the rack spawns 52 bullets/tick so the ring wrapped every ~39 ticks while a long power-3 shot needs ~90, silently discarding unresolved bullets and biasing every measured hit rate by range. Added a droppedBullets counter so a future overflow is measurable, and wavePushes/ waveStarved counters on the three guns. After the fix: vDropped = 0 and vStarved = 0 across all 48 recorded rounds. fitnessFor is now exported, deterministic (enemies iterated in ascending id order) and shared by the selector and the stats dump, replacing a hand-rolled merge in ModularBot that never advanced its window head. Round lines gain additive keys: vDropped, vStarved. |
||
|
|
2cc2a3bd87 |
fix(ModularBot): ram loop prevention, dead-target guards, cleaner logging
- 30-tick cooldown after ghost-stuck/timeout ram exit prevents re-entry loop - enemy_tracker.update() skips dead bots to prevent same-tick scan resurrection - TFIL graphics cleared when ramming is active movement - [config] logs: white base with green-highlighted changes only - [ram:enter] logs trigger reason and key values on false→true transition - [death] and [target-invalid] logs retained for diagnostics |
||
|
|
c94ba1f2fe |
feat(ModularBot): decay-GF gun (recency-weighted), 13 guns total
- Decay-GF: exponential decay on GF histogram (0.998/tick, ~350-tick half-life) - Adapts faster to mid-battle strategy changes than standard GF - Battle-tested vs WaveSurfer, Crazy, RandomMover |