Files
SirRoboGarage/docs/tmcomposites_gate.md
T
SirStone f58d65d2e8 TMComposites gate: per-gun confidence faithful for 3 guns; no pair composes
Adds a per-sample intrinsic-confidence field (GunPrediction.confidence,
threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF,
KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that
reproduce the paper's Figure 2 per gun and its Eq-8 composite.

Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun):
- FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak).
- GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001).
- No pair of guns specialises complementarily: the same gun dominates both
  high-confidence slices in every pair.
- Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern
  20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses.
  Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of
  competence is real but ~2pp short. Offline veto: design is dead.

See docs/tmcomposites_gate.md.
2026-09-25 22:04:39 +02:00

18 KiB
Raw Blame History

TMComposites gate: is each gun's intrinsic confidence faithful, and are the guns complementary specialists?

Scope. Test the mechanism of TMComposites: Plug-and-Play Collaboration Between Specialized Tsetlin Machines (Granmo, arXiv:2309.04801v2, §3) on our gun rack: (A) is each gun's per-sample confidence faithful to its own accuracy? (B) are the guns complementary specialists? (C) does an Eq-8 alpha-normalised confidence-weighted composite beat the best single gun offline, and is any gain attributable to competence (shuffle control)? This is the one untested idea that could plausibly beat onlyPattern, whose selector was measured negative value (docs/selector_negative_value.md, docs/gun_rack_analysis.md, commit e0666a5) because it decides with a rolling hit-rate instead of a per-sample signal.

Evidence tags. [MEASURED] = read from the committed result common_libs/tests/fixtures/tmcomposites_gate.json, reproduced by the named command, or a source fact. [INFERRED] = reasoning from those measurements.


0. Direct answers

  1. Which guns are faithfully confident? [MEASURED] Pattern (weak), DecayGF (strong), KNN (moderate) are faithful: rank samples by the gun's own confidence and accuracy rises. GuessFactor is ANTI-faithful — it is more accurate when it is less confident. Tsetlin's class-sum max is useless (Spearman 0.001, p=0.68). The eight deterministic geometric guns (HeadOn, Linear, Circular, WallBounce, Accel, StopShot, Displace, AvgLead) expose no per-sample confidence at all.

  2. Do any pairs specialise complementarily? [MEASURED] No. In every one of the 10 pairs, on the slice where A's normalised confidence beats B's, A is not the more accurate gun — or the same gun dominates both slices (e.g. GuessFactor is more accurate than DecayGF and than KNN even on their own high-confidence slices). The paper's premise — one member's weakness is another's strength, decided by confidence — does not hold on our rack. Our guns are different estimators of the same target, not different specialists.

  3. Does the composite beat the best single gun? [MEASURED] No, and the design is dead offline. The Eq-8 composite scores 18.41% vs Pattern's 20.45% (McNemar p=3.1e-126); the faithful-only variant (Pattern + DecayGF + KNN) scores 18.68%, still losing to Pattern by 1.77pp (p=1.25e-133). Both lose on all three held-out battles. The confidence-shuffle control passes in the weak sense — the composite beats its own shuffle (18.41 vs 17.61, p=4.0e-14) — so the weighting carries a real but tiny competence signal; it is simply nowhere near enough to beat Pattern. Per the docs/offline_harness_trust.md rule (offline is veto-only), this negative kills the design; do not take a composite to a live A/B.


1. What was measured [MEASURED]

The offline gun range (common_libs/gun_harness/offline_range.nim + virtual_bullets.nim), driven by a new recorder common_libs/tests/measure_tmcomposites.nim. For every resolved virtual bullet it records (gun, tick, powerBin, confidence, hit, aim-bearing-relative-to-LOS, range). Ground truth is the shipped bmPath virtual-bullet metric (18 px hit radius — BotRadius), reproduced exactly per bullet, not a rolling rate.

fixtures ticks note
train (alpha calibration only) drussgt_vs_crazy, drussgt_vs_spinbot, drussgt_vs_ramfire, drussgt_vs_corners 22 242 classic-Robocode DrussGT
test (all reported numbers) tr_drussgt_vs_modularbot, tr_drussgt_vs_spinbot, tr_drussgt_vs_corners 33 425 Tank-Royale bridge captures

Each fixture is replayed with a fresh rack (as run_range does); the test battles are held out by battle, never by tick. The dump is 3 763 298 rows; per-gun test n ≈ 133 000 (33 425 ticks × 4 power bins).

Per-gun confidence signal defined in source [MEASURED] (the new GunPrediction.confidence field, common_libs/gun_harness/gun_interface.nim):

gun intrinsic per-sample signal source
GuessFactor peak GF bin weight max_i bins[i] (Eq 4 analogue) guess_factor.nim
DecayGF peak decayed GF bin weight decay_gf.nim
KNN peak Gaussian density over GF candidates (bestScore) knn_gun.nim
Pattern match quality 1/(1+bestMatchCost) pattern_matcher.nim
Tsetlin magnitude of the clamped clause-sum vote hypot(vx,vy) tsetlin.nim
TMHorizon side-class margin ` votes1-votes0
8 geometric guns none — confidence 0.0 head_on/linear/circular/wall_bounce/accel_predictor/stop_shot/displacement/averaged_lead

WARNING — this is exactly the mechanism class the task forbids: all five are per-sample statistics of the gun's own internal state at prediction time. None is a rolling accuracy or hit-rate average.


2. The mechanism (paper §3, Eqs 4/6/7/8)

  • Member t outputs class sums c^i_{t,d} (Eq 6); confidence is c_max (Eq 4).
  • Eq 7 normalises by alpha_t = max_{d,i}(c) - min_{d,i}(c).
  • Eq 8: y_d = argmax_i sum_t (1/alpha_t) c^i_{t,d}.

Adapted to an angle output: the class is a 0.5° aim bin relative to the fire-time line of sight ([-45°, +45°], 180 bins); each confident gun casts its alpha-normalised confidence into the bin of its own aim; the argmax bin wins, and the aim distance is the confidence-weighted mean distance of that bin's voters. alpha_t is calibrated on the train fixtures and frozen for test, so the test composite never sees test data when choosing its weights (Eq 7's X = train). Two composites are scored:

  • Composite = all five confidence guns {Tsetlin, GuessFactor, Pattern, DecayGF, KNN}.
  • CompositeF = only the guns §3 finds faithful {Pattern, DecayGF, KNN}.

3. Task A — confidence faithfulness [MEASURED]

Test samples ranked by each gun's own confidence. rho = Spearman(confidence, hit); bottom/top = accuracy in the lower/upper confidence half; the two-prop p is for top vs bottom.

gun n nonzero base rho rho p bottom% top% verdict
Accel 133 101 0 21.02 — — 14.47 27.57 NO SIGNAL
Pattern 133 100 132 860 20.45 +0.064 4.1e-121 18.61 22.28 FAITHFUL (weak)
WallBounce 133 103 0 20.15 — — 14.84 25.46 NO SIGNAL
AvgLead 133 100 0 20.08 — — 14.43 25.73 NO SIGNAL
StopShot 133 078 0 18.21 — — 12.82 23.61 NO SIGNAL
Tsetlin 132 964 105 353 18.18 +0.001 0.68 17.94 18.42 USELESS
Linear 133 104 0 17.31 — — 13.04 21.57 NO SIGNAL
Circular 133 107 0 17.27 — — 12.68 21.86 NO SIGNAL
GuessFactor 133 108 133 108 15.89 −0.067 5.1e-132 18.14 13.65 ANTI-FAITHFUL
DecayGF 133 104 133 104 14.83 +0.133 <1e-300 10.79 18.86 FAITHFUL (strong)
Displace 133 105 0 13.95 — — 11.42 16.47 NO SIGNAL
KNN 133 121 132 685 12.32 +0.090 2.1e-237 9.65 14.99 FAITHFUL
HeadOn 133 162 0 6.18 — — 7.09 5.27 NO SIGNAL

Accuracy-vs-confidence curves (deciles of confidence, low→high), the paper's Figure 2 reproduced per gun:

Pattern        14% 16% 19% 23% 22% 22% 21% 23% 22% 23%   rises, then flat
DecayGF        10%  9% 11% 15%  9% 11% 16% 21% 23% 23%   rises (noisy)
KNN            10%  9%  9% 10% 10% 12% 13% 15% 16% 19%   monotone rise
GuessFactor    16% 21% 21% 15% 18% 14% 18% 14% 13% 10%   FALLS
Tsetlin        13% 25% 12% 21% 19% 20% 16% 18% 20% 19%   flat/noise

Reading. DecayGF and KNN are the textbook faithful shapes (accuracy climbs with confidence). Pattern is faithful but weakly — its one prior is strong (18.61% → 22.28%) and then flattens, i.e. the match cost discriminates "no/poor match" from "match" but not much between matches. GuessFactor is the surprise and the most important negative: its histogram peak is anti-faithful, consistent with the earlier finding that the GF code path is degenerate on these tr-bridge captures. Tsetlin's c_max (clause-sum magnitude) carries no information about whether the shot hits — [INFERRED] because the TM's regression correction is trained on a residual and never learns the surfer (tsetlin.nim header: every variant sat at chance vs a shuffled-control).


4. Task B — complementary specialists? [MEASURED]

For each pair we split test samples by which gun has the higher alpha-normalised confidence and measure both guns on each slice. A|A = A's accuracy on the slice A wins, B|A = B's accuracy on that same slice; p is a paired McNemar.

A B A_wins A|A B|A p_A B_wins A|B B|B p_B complementary
DecayGF GuessFactor 43 493 19.1 20.0 4e-14 89 608 12.8 13.9 8e-46 No (GF dominates both)
DecayGF KNN 436 30.0 30.7 0.25 132 656 14.8 12.3 4e-99 No (DecayGF wins 0.3% only)
DecayGF Pattern 26 788 11.0 14.8 1e-51 106 279 15.8 21.9 0 No (Pattern dominates)
DecayGF Tsetlin 121 163 14.7 18.4 3e-211 11 795 16.2 15.4 0.045 No
GuessFactor KNN 44 521 12.3 8.5 2e-84 88 570 17.7 14.3 9e-115 No (GF dominates both)
GuessFactor Pattern 40 605 10.6 14.3 7e-69 92 459 18.2 23.2 2e-207 No
GuessFactor Tsetlin 117 799 15.5 17.9 6e-92 15 156 19.2 20.4 0.0012 No
KNN Pattern 38 485 9.7 15.8 6e-155 94 344 13.4 22.4 0 No
KNN Tsetlin 131 820 12.2 18.1 0 804 21.6 30.0 6e-06 No (KNN wins 0.6% only)
Pattern Tsetlin 121 591 21.1 18.8 2e-58 11 221 13.6 11.6 2e-07 No (Pattern dominates)

No pair is complementary. The pattern in nearly every row is that the more accurate gun is the same one on both slices — the confidence orderings are not aligned across guns, so "A is confident here" does not mean "A is the expert here". The two pairs with an unequal split (DecayGF vs KNN, KNN vs Tsetlin) are ones where the "loser" wins only 0.3–0.6% of samples — not a usable slice. A pair with no complementary slices cannot form a useful composite, and none does. [INFERRED] the mechanism: all five guns estimate the same quantity (the enemy's intercept) from the same recorded trajectory; differing booleanisations create different noise, not different competence regions, so their confidence rankings carry no cross-gun information.


5. Task C — composite vs best single, with the shuffle control [MEASURED]

Held-out test battles, paired per sample. member_oracle = accuracy if a perfect per-sample selector could pick any confidence member (the ceiling for any member-picking composite; a free-aim oracle could only be higher).

arm accuracy n
Accel (best single overall — no confidence) 21.02% 133 101
Pattern (best single with confidence / incumbent) 20.45% 133 100
WallBounce 20.15% 133 103
AvgLead 20.08% 133 100
CompositeF (faithful only: Pattern+DecayGF+KNN) 18.68% 133 126
Composite (all 5) 18.41% 133 094
Tsetlin 18.18% 132 964
CompositeFShuf (control) 18.06% 133 119
CompositeShuf (control) 17.61% 133 103
GuessFactor 15.89% 133 108
DecayGF 14.83% 133 104
KNN 12.32% 133 121
member_oracle (perfect member picker) 41.72% 133 094

Paired comparisons (McNemar on discordant pairs):

comparison hits opponents p
Composite vs Pattern 24 488 27 215 3.1e-126 (composite loses)
CompositeF vs Pattern 24 854 27 215 1.25e-133 (composite loses)
Composite vs its shuffle 24 496 23 443 4.0e-14 (composite wins)
CompositeF vs its shuffle 24 864 24 040 4.0e-15 (composite wins)
CompositeF vs Composite — — (faithful-only is marginally better, +0.27pp)

Per held-out battle (composite vs Pattern): tr_drussgt_vs_modularbot 12.2% vs 13.0%; tr_drussgt_vs_spinbot 29.0% vs 33.9%; tr_drussgt_vs_corners 22.3% vs 22.4%. The composite loses all three.

Verdict: the design is dead offline. The alpha-normalised confidence-weighted composite does not beat the best single gun; it loses to Pattern by ~1.8–2.0pp with p≈1e-130, and to the best single overall (Accel) by more. The shuffle control passes in the weak sense: the composite is genuinely (p≈1e-14) better than a version with the same weighting distribution but shuffled competences, so the confidence signal is not pure noise — but the effect is ~0.6–0.8pp versus a ~2pp deficit, i.e. real but far too small. The member_oracle of 41.72% shows enormous headroom exists — it is not reachable from these confidence signals.


6. Why it fails [INFERRED]

  1. Faithfulness is not competence. A gun can be perfectly confidence-ordered and still be worse than another gun everywhere. DecayGF is more faithful than Pattern (rho 0.133 vs 0.064) yet 5.6pp less accurate, so its (correct) ordering contributes weak votes against a stronger member.
  2. The members are not specialists. §4 shows the confidence orderings do not identify competence regions across guns; every gun attacks the whole input space. The paper's win comes from members that are good on disjoint subsets.
  3. The signal is diluted by anti-faithful members. GuessFactor is anti-faithful and Tsetlin is useless; CompositeF (faithful only) is better than Composite (18.68 vs 18.41) and more faithful (rho 0.114 vs 0.009), which confirms the poison — but removing it still leaves the composite below Pattern.
  4. Alpha normalisation is unstable for online-learning guns. Eq 7 sets alpha_t to the train range; GuessFactor's histogram accumulates unboundedly (alpha≈1.5e4 here), so its normalised confidence is tiny on test and it almost never casts a decisive vote. [INFERRED] this mutes the very member whose confidence was measured anti-faithful.

7. Limits and caveats

  • Open-loop corpus. [MEASURED]/[INFERRED] The fixtures are recorded trajectories; the enemy never reacts to the composite's (or any) bullets, and the tr-bridge battles were recorded while Pattern's selector was shooting. A gun that behaves like Pattern is therefore favoured in framing. This cannot rescue the composite: it loses to Pattern and to Accel/AvgLead/WallBounce, so the deficit is not a framing artefact. Per docs/offline_harness_trust.md, this is exactly the closed-loop class of question where the offline harness is veto-only — a negative kills the design, a positive would have proved nothing.
  • Coverage gap (documented LIMIT). [MEASURED] The offline range builds guns 0..13; TMPATTERN (14) and TMHORIZON (15) are not in it (docs/offline_harness_trust.md §2.8). TMHorizon's sideConf is instrumented in source but not measured here. Since Tsetlin's c_max is already useless and no pair composes, the omission does not change the verdict.
  • One composite architecture. [INFERRED] The vote uses each gun's scalar confidence at the bin of its own aim. The paper's members also expose a full class distribution (GF/KNN do). Feeding those full distributions could refine the composite, but §4 shows the core premise — cross-gun competence specialisation — is absent, so a refined vote has no signal to exploit.

8. Recommendation

Do not build or A/B the composite. The offline gate is a veto and it vetoes.

Two by-products worth keeping:

  • GuessFactor / anti-faithful warning. GuessFactor's own histogram peak is anti-faithful on this corpus; do not use it as a competence signal (e.g. for a confidence-gated firing decision or a power policy). Pattern, DecayGF and KNN confidences are faithful and could gate a shot ("do not fire when not confident") — a different, narrower mechanism than the composite, and still subject to the live veto.
  • The selector conclusion is reinforced. The rack's problem is not the decision statistic the selector uses (rolling rate vs per-sample confidence): even a per-sample intrinsic confidence, applied cross-gun, cannot beat Pattern. The rack does not contain complementary specialists.

9. Reproduction [MEASURED]

# 1. record per-sample confidence + hit (≈ 6.5 min; fresh rack per fixture)
nim c -d:release --nimcache:/tmp/nc_j104 --path:common_libs \
    -o:/tmp/measure_tmc common_libs/tests/measure_tmcomposites.nim
/tmp/measure_tmc --out /tmp/tmc_full.jsonl \
  --train tools/fixtures/drussgt_vs_crazy.jsonl tools/fixtures/drussgt_vs_spinbot.jsonl \
          tools/fixtures/drussgt_vs_ramfire.jsonl tools/fixtures/drussgt_vs_corners.jsonl \
  --test  tools/fixtures/tr_drussgt_vs_modularbot.jsonl tools/fixtures/tr_drussgt_vs_spinbot.jsonl \
          tools/fixtures/tr_drussgt_vs_corners.jsonl

# 2. analyse (≈ 40 s); committed result is common_libs/tests/fixtures/tmcomposites_gate.json
python3 common_libs/tests/analyze_tmcomposites.py \
  --input /tmp/tmc_full.jsonl --json common_libs/tests/fixtures/tmcomposites_gate.json

The recorder threads a new per-sample confidence field through GunPrediction/FeedbackEvent/VirtualBullet (gun_interface.nim, virtual_bullets.nim) and populates it in guess_factor.nim, decay_gf.nim, knn_gun.nim, pattern_matcher.nim, tsetlin.nim, tm_horizon.nim; the field defaults to 0.0, so every existing caller and test is unchanged. Verified test_gun_harness, test_vbullet_metric, test_wave_pairing, test_tm_horizon, test_pattern_radial_offset, test_range_rack_parity, test_selector_tiebreak all pass.