Adds a per-sample intrinsic-confidence field (GunPrediction.confidence, threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF, KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that reproduce the paper's Figure 2 per gun and its Eq-8 composite. Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun): - FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak). - GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001). - No pair of guns specialises complementarily: the same gun dominates both high-confidence slices in every pair. - Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern 20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses. Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of competence is real but ~2pp short. Offline veto: design is dead. See docs/tmcomposites_gate.md.
18 KiB
TMComposites gate: is each gun's intrinsic confidence faithful, and are the guns complementary specialists?
Scope. Test the mechanism of TMComposites: Plug-and-Play Collaboration
Between Specialized Tsetlin Machines (Granmo, arXiv:2309.04801v2, §3) on our
gun rack: (A) is each gun's per-sample confidence faithful to its own accuracy?
(B) are the guns complementary specialists? (C) does an Eq-8 alpha-normalised
confidence-weighted composite beat the best single gun offline, and is any gain
attributable to competence (shuffle control)? This is the one untested idea that
could plausibly beat onlyPattern, whose selector was measured negative value
(docs/selector_negative_value.md, docs/gun_rack_analysis.md, commit e0666a5)
because it decides with a rolling hit-rate instead of a per-sample signal.
Evidence tags. [MEASURED] = read from the committed result
common_libs/tests/fixtures/tmcomposites_gate.json, reproduced by the named
command, or a source fact. [INFERRED] = reasoning from those measurements.
0. Direct answers
-
Which guns are faithfully confident?
[MEASURED]Pattern (weak), DecayGF (strong), KNN (moderate) are faithful: rank samples by the gun's own confidence and accuracy rises. GuessFactor is ANTI-faithful — it is more accurate when it is less confident. Tsetlin's class-sum max is useless (Spearman 0.001, p=0.68). The eight deterministic geometric guns (HeadOn, Linear, Circular, WallBounce, Accel, StopShot, Displace, AvgLead) expose no per-sample confidence at all. -
Do any pairs specialise complementarily?
[MEASURED]No. In every one of the 10 pairs, on the slice where A's normalised confidence beats B's, A is not the more accurate gun — or the same gun dominates both slices (e.g. GuessFactor is more accurate than DecayGF and than KNN even on their own high-confidence slices). The paper's premise — one member's weakness is another's strength, decided by confidence — does not hold on our rack. Our guns are different estimators of the same target, not different specialists. -
Does the composite beat the best single gun?
[MEASURED]No, and the design is dead offline. The Eq-8 composite scores 18.41% vs Pattern's 20.45% (McNemar p=3.1e-126); the faithful-only variant (Pattern + DecayGF + KNN) scores 18.68%, still losing to Pattern by 1.77pp (p=1.25e-133). Both lose on all three held-out battles. The confidence-shuffle control passes in the weak sense — the composite beats its own shuffle (18.41 vs 17.61, p=4.0e-14) — so the weighting carries a real but tiny competence signal; it is simply nowhere near enough to beat Pattern. Per thedocs/offline_harness_trust.mdrule (offline is veto-only), this negative kills the design; do not take a composite to a live A/B.
1. What was measured [MEASURED]
The offline gun range (common_libs/gun_harness/offline_range.nim +
virtual_bullets.nim), driven by a new recorder
common_libs/tests/measure_tmcomposites.nim. For every resolved virtual bullet
it records (gun, tick, powerBin, confidence, hit, aim-bearing-relative-to-LOS, range). Ground truth is the shipped bmPath virtual-bullet metric (18 px hit
radius — BotRadius), reproduced exactly per bullet, not a rolling rate.
| fixtures | ticks | note | |
|---|---|---|---|
| train (alpha calibration only) | drussgt_vs_crazy, drussgt_vs_spinbot, drussgt_vs_ramfire, drussgt_vs_corners |
22 242 | classic-Robocode DrussGT |
| test (all reported numbers) | tr_drussgt_vs_modularbot, tr_drussgt_vs_spinbot, tr_drussgt_vs_corners |
33 425 | Tank-Royale bridge captures |
Each fixture is replayed with a fresh rack (as run_range does); the test
battles are held out by battle, never by tick. The dump is 3 763 298 rows;
per-gun test n ≈ 133 000 (33 425 ticks × 4 power bins).
Per-gun confidence signal defined in source [MEASURED] (the new
GunPrediction.confidence field, common_libs/gun_harness/gun_interface.nim):
| gun | intrinsic per-sample signal | source |
|---|---|---|
| GuessFactor | peak GF bin weight max_i bins[i] (Eq 4 analogue) |
guess_factor.nim |
| DecayGF | peak decayed GF bin weight | decay_gf.nim |
| KNN | peak Gaussian density over GF candidates (bestScore) |
knn_gun.nim |
| Pattern | match quality 1/(1+bestMatchCost) |
pattern_matcher.nim |
| Tsetlin | magnitude of the clamped clause-sum vote hypot(vx,vy) |
tsetlin.nim |
| TMHorizon | side-class margin ` | votes1-votes0 |
| 8 geometric guns | none — confidence 0.0 | head_on/linear/circular/wall_bounce/accel_predictor/stop_shot/displacement/averaged_lead |
WARNING — this is exactly the mechanism class the task forbids: all five are per-sample statistics of the gun's own internal state at prediction time. None is a rolling accuracy or hit-rate average.
2. The mechanism (paper §3, Eqs 4/6/7/8)
- Member
toutputs class sumsc^i_{t,d}(Eq 6); confidence isc_max(Eq 4). - Eq 7 normalises by
alpha_t = max_{d,i}(c) - min_{d,i}(c). - Eq 8:
y_d = argmax_i sum_t (1/alpha_t) c^i_{t,d}.
Adapted to an angle output: the class is a 0.5° aim bin relative to the
fire-time line of sight ([-45°, +45°], 180 bins); each confident gun casts its
alpha-normalised confidence into the bin of its own aim; the argmax bin wins, and
the aim distance is the confidence-weighted mean distance of that bin's voters.
alpha_t is calibrated on the train fixtures and frozen for test, so the test
composite never sees test data when choosing its weights (Eq 7's X = train).
Two composites are scored:
- Composite = all five confidence guns {Tsetlin, GuessFactor, Pattern, DecayGF, KNN}.
- CompositeF = only the guns §3 finds faithful {Pattern, DecayGF, KNN}.
3. Task A — confidence faithfulness [MEASURED]
Test samples ranked by each gun's own confidence. rho = Spearman(confidence,
hit); bottom/top = accuracy in the lower/upper confidence half; the two-prop
p is for top vs bottom.
| gun | n | nonzero | base | rho | rho p | bottom% | top% | verdict |
|---|---|---|---|---|---|---|---|---|
| Accel | 133 101 | 0 | 21.02 | — | — | 14.47 | 27.57 | NO SIGNAL |
| Pattern | 133 100 | 132 860 | 20.45 | +0.064 | 4.1e-121 | 18.61 | 22.28 | FAITHFUL (weak) |
| WallBounce | 133 103 | 0 | 20.15 | — | — | 14.84 | 25.46 | NO SIGNAL |
| AvgLead | 133 100 | 0 | 20.08 | — | — | 14.43 | 25.73 | NO SIGNAL |
| StopShot | 133 078 | 0 | 18.21 | — | — | 12.82 | 23.61 | NO SIGNAL |
| Tsetlin | 132 964 | 105 353 | 18.18 | +0.001 | 0.68 | 17.94 | 18.42 | USELESS |
| Linear | 133 104 | 0 | 17.31 | — | — | 13.04 | 21.57 | NO SIGNAL |
| Circular | 133 107 | 0 | 17.27 | — | — | 12.68 | 21.86 | NO SIGNAL |
| GuessFactor | 133 108 | 133 108 | 15.89 | −0.067 | 5.1e-132 | 18.14 | 13.65 | ANTI-FAITHFUL |
| DecayGF | 133 104 | 133 104 | 14.83 | +0.133 | <1e-300 | 10.79 | 18.86 | FAITHFUL (strong) |
| Displace | 133 105 | 0 | 13.95 | — | — | 11.42 | 16.47 | NO SIGNAL |
| KNN | 133 121 | 132 685 | 12.32 | +0.090 | 2.1e-237 | 9.65 | 14.99 | FAITHFUL |
| HeadOn | 133 162 | 0 | 6.18 | — | — | 7.09 | 5.27 | NO SIGNAL |
Accuracy-vs-confidence curves (deciles of confidence, low→high), the paper's Figure 2 reproduced per gun:
Pattern 14% 16% 19% 23% 22% 22% 21% 23% 22% 23% rises, then flat
DecayGF 10% 9% 11% 15% 9% 11% 16% 21% 23% 23% rises (noisy)
KNN 10% 9% 9% 10% 10% 12% 13% 15% 16% 19% monotone rise
GuessFactor 16% 21% 21% 15% 18% 14% 18% 14% 13% 10% FALLS
Tsetlin 13% 25% 12% 21% 19% 20% 16% 18% 20% 19% flat/noise
Reading. DecayGF and KNN are the textbook faithful shapes (accuracy climbs
with confidence). Pattern is faithful but weakly — its one prior is strong
(18.61% → 22.28%) and then flattens, i.e. the match cost discriminates
"no/poor match" from "match" but not much between matches. GuessFactor is the
surprise and the most important negative: its histogram peak is anti-faithful,
consistent with the earlier finding that the GF code path is degenerate on these
tr-bridge captures. Tsetlin's c_max (clause-sum magnitude) carries no
information about whether the shot hits — [INFERRED] because the TM's
regression correction is trained on a residual and never learns the surfer
(tsetlin.nim header: every variant sat at chance vs a shuffled-control).
4. Task B — complementary specialists? [MEASURED]
For each pair we split test samples by which gun has the higher alpha-normalised
confidence and measure both guns on each slice. A|A = A's accuracy on the
slice A wins, B|A = B's accuracy on that same slice; p is a paired McNemar.
| A | B | A_wins | A|A | B|A | p_A | B_wins | A|B | B|B | p_B | complementary |
|---|---|---|---|---|---|---|---|---|---|---|
| DecayGF | GuessFactor | 43 493 | 19.1 | 20.0 | 4e-14 | 89 608 | 12.8 | 13.9 | 8e-46 | No (GF dominates both) |
| DecayGF | KNN | 436 | 30.0 | 30.7 | 0.25 | 132 656 | 14.8 | 12.3 | 4e-99 | No (DecayGF wins 0.3% only) |
| DecayGF | Pattern | 26 788 | 11.0 | 14.8 | 1e-51 | 106 279 | 15.8 | 21.9 | 0 | No (Pattern dominates) |
| DecayGF | Tsetlin | 121 163 | 14.7 | 18.4 | 3e-211 | 11 795 | 16.2 | 15.4 | 0.045 | No |
| GuessFactor | KNN | 44 521 | 12.3 | 8.5 | 2e-84 | 88 570 | 17.7 | 14.3 | 9e-115 | No (GF dominates both) |
| GuessFactor | Pattern | 40 605 | 10.6 | 14.3 | 7e-69 | 92 459 | 18.2 | 23.2 | 2e-207 | No |
| GuessFactor | Tsetlin | 117 799 | 15.5 | 17.9 | 6e-92 | 15 156 | 19.2 | 20.4 | 0.0012 | No |
| KNN | Pattern | 38 485 | 9.7 | 15.8 | 6e-155 | 94 344 | 13.4 | 22.4 | 0 | No |
| KNN | Tsetlin | 131 820 | 12.2 | 18.1 | 0 | 804 | 21.6 | 30.0 | 6e-06 | No (KNN wins 0.6% only) |
| Pattern | Tsetlin | 121 591 | 21.1 | 18.8 | 2e-58 | 11 221 | 13.6 | 11.6 | 2e-07 | No (Pattern dominates) |
No pair is complementary. The pattern in nearly every row is that the more
accurate gun is the same one on both slices — the confidence orderings are not
aligned across guns, so "A is confident here" does not mean "A is the expert
here". The two pairs with an unequal split (DecayGF vs KNN, KNN vs Tsetlin) are
ones where the "loser" wins only 0.3–0.6% of samples — not a usable slice. A
pair with no complementary slices cannot form a useful composite, and none does.
[INFERRED] the mechanism: all five guns estimate the same quantity (the
enemy's intercept) from the same recorded trajectory; differing booleanisations
create different noise, not different competence regions, so their confidence
rankings carry no cross-gun information.
5. Task C — composite vs best single, with the shuffle control [MEASURED]
Held-out test battles, paired per sample. member_oracle = accuracy if a
perfect per-sample selector could pick any confidence member (the ceiling for
any member-picking composite; a free-aim oracle could only be higher).
| arm | accuracy | n |
|---|---|---|
| Accel (best single overall — no confidence) | 21.02% | 133 101 |
| Pattern (best single with confidence / incumbent) | 20.45% | 133 100 |
| WallBounce | 20.15% | 133 103 |
| AvgLead | 20.08% | 133 100 |
| CompositeF (faithful only: Pattern+DecayGF+KNN) | 18.68% | 133 126 |
| Composite (all 5) | 18.41% | 133 094 |
| Tsetlin | 18.18% | 132 964 |
| CompositeFShuf (control) | 18.06% | 133 119 |
| CompositeShuf (control) | 17.61% | 133 103 |
| GuessFactor | 15.89% | 133 108 |
| DecayGF | 14.83% | 133 104 |
| KNN | 12.32% | 133 121 |
| member_oracle (perfect member picker) | 41.72% | 133 094 |
Paired comparisons (McNemar on discordant pairs):
| comparison | hits | opponents | p |
|---|---|---|---|
| Composite vs Pattern | 24 488 | 27 215 | 3.1e-126 (composite loses) |
| CompositeF vs Pattern | 24 854 | 27 215 | 1.25e-133 (composite loses) |
| Composite vs its shuffle | 24 496 | 23 443 | 4.0e-14 (composite wins) |
| CompositeF vs its shuffle | 24 864 | 24 040 | 4.0e-15 (composite wins) |
| CompositeF vs Composite | — | — | (faithful-only is marginally better, +0.27pp) |
Per held-out battle (composite vs Pattern): tr_drussgt_vs_modularbot 12.2% vs
13.0%; tr_drussgt_vs_spinbot 29.0% vs 33.9%; tr_drussgt_vs_corners 22.3% vs
22.4%. The composite loses all three.
Verdict: the design is dead offline. The alpha-normalised confidence-weighted
composite does not beat the best single gun; it loses to Pattern by ~1.8–2.0pp
with p≈1e-130, and to the best single overall (Accel) by more. The shuffle control
passes in the weak sense: the composite is genuinely (p≈1e-14) better than a
version with the same weighting distribution but shuffled competences, so the
confidence signal is not pure noise — but the effect is ~0.6–0.8pp versus a ~2pp
deficit, i.e. real but far too small. The member_oracle of 41.72% shows
enormous headroom exists — it is not reachable from these confidence signals.
6. Why it fails [INFERRED]
- Faithfulness is not competence. A gun can be perfectly confidence-ordered and still be worse than another gun everywhere. DecayGF is more faithful than Pattern (rho 0.133 vs 0.064) yet 5.6pp less accurate, so its (correct) ordering contributes weak votes against a stronger member.
- The members are not specialists. §4 shows the confidence orderings do not identify competence regions across guns; every gun attacks the whole input space. The paper's win comes from members that are good on disjoint subsets.
- The signal is diluted by anti-faithful members. GuessFactor is anti-faithful and Tsetlin is useless; CompositeF (faithful only) is better than Composite (18.68 vs 18.41) and more faithful (rho 0.114 vs 0.009), which confirms the poison — but removing it still leaves the composite below Pattern.
- Alpha normalisation is unstable for online-learning guns. Eq 7 sets
alpha_tto the train range; GuessFactor's histogram accumulates unboundedly (alpha≈1.5e4 here), so its normalised confidence is tiny on test and it almost never casts a decisive vote.[INFERRED]this mutes the very member whose confidence was measured anti-faithful.
7. Limits and caveats
- Open-loop corpus.
[MEASURED]/[INFERRED]The fixtures are recorded trajectories; the enemy never reacts to the composite's (or any) bullets, and thetr-bridgebattles were recorded while Pattern's selector was shooting. A gun that behaves like Pattern is therefore favoured in framing. This cannot rescue the composite: it loses to Pattern and to Accel/AvgLead/WallBounce, so the deficit is not a framing artefact. Perdocs/offline_harness_trust.md, this is exactly the closed-loop class of question where the offline harness is veto-only — a negative kills the design, a positive would have proved nothing. - Coverage gap (documented LIMIT).
[MEASURED]The offline range builds guns 0..13; TMPATTERN (14) and TMHORIZON (15) are not in it (docs/offline_harness_trust.md§2.8). TMHorizon'ssideConfis instrumented in source but not measured here. Since Tsetlin's c_max is already useless and no pair composes, the omission does not change the verdict. - One composite architecture.
[INFERRED]The vote uses each gun's scalar confidence at the bin of its own aim. The paper's members also expose a full class distribution (GF/KNN do). Feeding those full distributions could refine the composite, but §4 shows the core premise — cross-gun competence specialisation — is absent, so a refined vote has no signal to exploit.
8. Recommendation
Do not build or A/B the composite. The offline gate is a veto and it vetoes.
Two by-products worth keeping:
- GuessFactor / anti-faithful warning. GuessFactor's own histogram peak is anti-faithful on this corpus; do not use it as a competence signal (e.g. for a confidence-gated firing decision or a power policy). Pattern, DecayGF and KNN confidences are faithful and could gate a shot ("do not fire when not confident") — a different, narrower mechanism than the composite, and still subject to the live veto.
- The selector conclusion is reinforced. The rack's problem is not the decision statistic the selector uses (rolling rate vs per-sample confidence): even a per-sample intrinsic confidence, applied cross-gun, cannot beat Pattern. The rack does not contain complementary specialists.
9. Reproduction [MEASURED]
# 1. record per-sample confidence + hit (≈ 6.5 min; fresh rack per fixture)
nim c -d:release --nimcache:/tmp/nc_j104 --path:common_libs \
-o:/tmp/measure_tmc common_libs/tests/measure_tmcomposites.nim
/tmp/measure_tmc --out /tmp/tmc_full.jsonl \
--train tools/fixtures/drussgt_vs_crazy.jsonl tools/fixtures/drussgt_vs_spinbot.jsonl \
tools/fixtures/drussgt_vs_ramfire.jsonl tools/fixtures/drussgt_vs_corners.jsonl \
--test tools/fixtures/tr_drussgt_vs_modularbot.jsonl tools/fixtures/tr_drussgt_vs_spinbot.jsonl \
tools/fixtures/tr_drussgt_vs_corners.jsonl
# 2. analyse (≈ 40 s); committed result is common_libs/tests/fixtures/tmcomposites_gate.json
python3 common_libs/tests/analyze_tmcomposites.py \
--input /tmp/tmc_full.jsonl --json common_libs/tests/fixtures/tmcomposites_gate.json
The recorder threads a new per-sample confidence field through
GunPrediction/FeedbackEvent/VirtualBullet (gun_interface.nim,
virtual_bullets.nim) and populates it in guess_factor.nim, decay_gf.nim,
knn_gun.nim, pattern_matcher.nim, tsetlin.nim, tm_horizon.nim; the field
defaults to 0.0, so every existing caller and test is unchanged. Verified
test_gun_harness, test_vbullet_metric, test_wave_pairing,
test_tm_horizon, test_pattern_radial_offset, test_range_rack_parity,
test_selector_tiebreak all pass.