Files
SirRoboGarage/docs/tmcomposites_gate.md
T
SirStone f58d65d2e8 TMComposites gate: per-gun confidence faithful for 3 guns; no pair composes
Adds a per-sample intrinsic-confidence field (GunPrediction.confidence,
threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF,
KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that
reproduce the paper's Figure 2 per gun and its Eq-8 composite.

Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun):
- FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak).
- GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001).
- No pair of guns specialises complementarily: the same gun dominates both
  high-confidence slices in every pair.
- Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern
  20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses.
  Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of
  competence is real but ~2pp short. Offline veto: design is dead.

See docs/tmcomposites_gate.md.
2026-09-25 22:04:39 +02:00

319 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TMComposites gate: is each gun's intrinsic confidence faithful, and are the guns complementary specialists?
**Scope.** Test the mechanism of *TMComposites: Plug-and-Play Collaboration
Between Specialized Tsetlin Machines* (Granmo, arXiv:2309.04801v2, §3) on our
gun rack: (A) is each gun's per-sample confidence faithful to its own accuracy?
(B) are the guns complementary specialists? (C) does an Eq-8 alpha-normalised
confidence-weighted composite beat the best single gun offline, and is any gain
attributable to competence (shuffle control)? This is the one untested idea that
could plausibly beat `onlyPattern`, whose selector was measured **negative value**
(`docs/selector_negative_value.md`, `docs/gun_rack_analysis.md`, commit `e0666a5`)
because it decides with a rolling hit-rate instead of a per-sample signal.
**Evidence tags.** `[MEASURED]` = read from the committed result
`common_libs/tests/fixtures/tmcomposites_gate.json`, reproduced by the named
command, or a source fact. `[INFERRED]` = reasoning from those measurements.
---
## 0. Direct answers
1. **Which guns are faithfully confident?** `[MEASURED]`
**Pattern (weak), DecayGF (strong), KNN (moderate)** are faithful: rank samples
by the gun's own confidence and accuracy rises. **GuessFactor is
ANTI-faithful** — it is more accurate when it is *less* confident. **Tsetlin's
class-sum max is useless** (Spearman 0.001, p=0.68). The eight deterministic
geometric guns (HeadOn, Linear, Circular, WallBounce, Accel, StopShot,
Displace, AvgLead) expose **no per-sample confidence at all**.
2. **Do any pairs specialise complementarily?** `[MEASURED]` **No.** In every one
of the 10 pairs, on the slice where A's normalised confidence beats B's, A is
*not* the more accurate gun — or the same gun dominates both slices (e.g.
GuessFactor is more accurate than DecayGF and than KNN even on their own
high-confidence slices). The paper's premise — one member's weakness is
another's strength, decided by confidence — does not hold on our rack. Our
guns are different *estimators of the same target*, not different *specialists*.
3. **Does the composite beat the best single gun?** `[MEASURED]` **No, and the
design is dead offline.** The Eq-8 composite scores **18.41%** vs Pattern's
**20.45%** (McNemar p=3.1e-126); the faithful-only variant (Pattern + DecayGF +
KNN) scores **18.68%**, still losing to Pattern by 1.77pp (p=1.25e-133). Both
lose on all three held-out battles. The **confidence-shuffle control passes in
the weak sense** — the composite beats its own shuffle (18.41 vs 17.61,
p=4.0e-14) — so the weighting carries a *real but tiny* competence signal; it
is simply nowhere near enough to beat Pattern. Per the `docs/offline_harness_trust.md`
rule (offline is veto-only), this negative **kills the design**; do not take a
composite to a live A/B.
---
## 1. What was measured `[MEASURED]`
The offline gun range (`common_libs/gun_harness/offline_range.nim` +
`virtual_bullets.nim`), driven by a new recorder
`common_libs/tests/measure_tmcomposites.nim`. For every resolved virtual bullet
it records `(gun, tick, powerBin, confidence, hit, aim-bearing-relative-to-LOS,
range)`. Ground truth is the shipped `bmPath` virtual-bullet metric (18 px hit
radius — `BotRadius`), reproduced exactly per bullet, not a rolling rate.
| | fixtures | ticks | note |
|---|---|---:|---|
| **train** (alpha calibration only) | `drussgt_vs_crazy`, `drussgt_vs_spinbot`, `drussgt_vs_ramfire`, `drussgt_vs_corners` | 22 242 | classic-Robocode DrussGT |
| **test** (all reported numbers) | `tr_drussgt_vs_modularbot`, `tr_drussgt_vs_spinbot`, `tr_drussgt_vs_corners` | 33 425 | Tank-Royale bridge captures |
Each fixture is replayed with a **fresh rack** (as `run_range` does); the test
battles are **held out by battle**, never by tick. The dump is 3 763 298 rows;
per-gun test `n ≈ 133 000` (33 425 ticks × 4 power bins).
**Per-gun confidence signal defined in source** `[MEASURED]` (the new
`GunPrediction.confidence` field, `common_libs/gun_harness/gun_interface.nim`):
| gun | intrinsic per-sample signal | source |
|---|---|---|
| GuessFactor | peak GF bin weight `max_i bins[i]` (Eq 4 analogue) | `guess_factor.nim` |
| DecayGF | peak decayed GF bin weight | `decay_gf.nim` |
| KNN | peak Gaussian density over GF candidates (`bestScore`) | `knn_gun.nim` |
| Pattern | match quality `1/(1+bestMatchCost)` | `pattern_matcher.nim` |
| Tsetlin | magnitude of the clamped clause-sum vote `hypot(vx,vy)` | `tsetlin.nim` |
| TMHorizon | side-class margin `|votes1-votes0|/(2*half)` | `tm_horizon.nim` (instrumented; see §7) |
| 8 geometric guns | none — confidence 0.0 | `head_on/linear/circular/wall_bounce/accel_predictor/stop_shot/displacement/averaged_lead` |
**WARNING — this is exactly the mechanism class the task forbids:** all five are
per-sample statistics of the gun's own internal state at prediction time. None is
a rolling accuracy or hit-rate average.
---
## 2. The mechanism (paper §3, Eqs 4/6/7/8)
- Member `t` outputs class sums `c^i_{t,d}` (Eq 6); confidence is `c_max` (Eq 4).
- Eq 7 normalises by `alpha_t = max_{d,i}(c) - min_{d,i}(c)`.
- Eq 8: `y_d = argmax_i sum_t (1/alpha_t) c^i_{t,d}`.
Adapted to an **angle output**: the class is a 0.5° aim bin relative to the
fire-time line of sight (`[-45°, +45°]`, 180 bins); each confident gun casts its
alpha-normalised confidence into the bin of its own aim; the argmax bin wins, and
the aim distance is the confidence-weighted mean distance of that bin's voters.
`alpha_t` is calibrated on the **train** fixtures and frozen for test, so the test
composite never sees test data when choosing its weights (Eq 7's `X` = train).
Two composites are scored:
- **Composite** = all five confidence guns {Tsetlin, GuessFactor, Pattern, DecayGF, KNN}.
- **CompositeF** = only the guns §3 finds faithful {Pattern, DecayGF, KNN}.
---
## 3. Task A — confidence faithfulness `[MEASURED]`
Test samples ranked by each gun's own confidence. `rho` = Spearman(confidence,
hit); `bottom/top` = accuracy in the lower/upper confidence half; the two-prop
p is for top vs bottom.
| gun | n | nonzero | base | rho | rho p | bottom% | top% | verdict |
|---|---:|---:|---:|---:|---:|---:|---:|---|
| Accel | 133 101 | 0 | 21.02 | — | — | 14.47 | 27.57 | NO SIGNAL |
| **Pattern** | 133 100 | 132 860 | 20.45 | **+0.064** | 4.1e-121 | 18.61 | **22.28** | **FAITHFUL (weak)** |
| WallBounce | 133 103 | 0 | 20.15 | — | — | 14.84 | 25.46 | NO SIGNAL |
| AvgLead | 133 100 | 0 | 20.08 | — | — | 14.43 | 25.73 | NO SIGNAL |
| StopShot | 133 078 | 0 | 18.21 | — | — | 12.82 | 23.61 | NO SIGNAL |
| Tsetlin | 132 964 | 105 353 | 18.18 | +0.001 | 0.68 | 17.94 | 18.42 | **USELESS** |
| Linear | 133 104 | 0 | 17.31 | — | — | 13.04 | 21.57 | NO SIGNAL |
| Circular | 133 107 | 0 | 17.27 | — | — | 12.68 | 21.86 | NO SIGNAL |
| **GuessFactor** | 133 108 | 133 108 | 15.89 | **−0.067** | 5.1e-132 | 18.14 | **13.65** | **ANTI-FAITHFUL** |
| **DecayGF** | 133 104 | 133 104 | 14.83 | **+0.133** | <1e-300 | 10.79 | **18.86** | **FAITHFUL (strong)** |
| Displace | 133 105 | 0 | 13.95 | — | — | 11.42 | 16.47 | NO SIGNAL |
| **KNN** | 133 121 | 132 685 | 12.32 | **+0.090** | 2.1e-237 | 9.65 | **14.99** | **FAITHFUL** |
| HeadOn | 133 162 | 0 | 6.18 | — | — | 7.09 | 5.27 | NO SIGNAL |
Accuracy-vs-confidence curves (deciles of confidence, low→high), the paper's
Figure 2 reproduced per gun:
```
Pattern 14% 16% 19% 23% 22% 22% 21% 23% 22% 23% rises, then flat
DecayGF 10% 9% 11% 15% 9% 11% 16% 21% 23% 23% rises (noisy)
KNN 10% 9% 9% 10% 10% 12% 13% 15% 16% 19% monotone rise
GuessFactor 16% 21% 21% 15% 18% 14% 18% 14% 13% 10% FALLS
Tsetlin 13% 25% 12% 21% 19% 20% 16% 18% 20% 19% flat/noise
```
**Reading.** DecayGF and KNN are the textbook faithful shapes (accuracy climbs
with confidence). Pattern is faithful but weakly — its one prior is strong
(`18.61% → 22.28%`) and then flattens, i.e. the match cost discriminates
"no/poor match" from "match" but not much *between* matches. GuessFactor is the
surprise and the most important negative: its histogram peak is **anti-faithful**,
consistent with the earlier finding that the GF code path is degenerate on these
`tr-bridge` captures. Tsetlin's `c_max` (clause-sum magnitude) carries **no**
information about whether the shot hits — `[INFERRED]` because the TM's
regression correction is trained on a residual and never learns the surfer
(`tsetlin.nim` header: every variant sat at chance vs a shuffled-control).
---
## 4. Task B — complementary specialists? `[MEASURED]`
For each pair we split test samples by which gun has the higher **alpha-normalised**
confidence and measure **both** guns on each slice. `A|A` = A's accuracy on the
slice A wins, `B|A` = B's accuracy on that same slice; `p` is a paired McNemar.
| A | B | A_wins | A\|A | B\|A | p_A | B_wins | A\|B | B\|B | p_B | complementary |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---|
| DecayGF | GuessFactor | 43 493 | 19.1 | **20.0** | 4e-14 | 89 608 | 12.8 | **13.9** | 8e-46 | **No** (GF dominates both) |
| DecayGF | KNN | 436 | 30.0 | **30.7** | 0.25 | 132 656 | **14.8** | 12.3 | 4e-99 | **No** (DecayGF wins 0.3% only) |
| DecayGF | Pattern | 26 788 | 11.0 | **14.8** | 1e-51 | 106 279 | 15.8 | **21.9** | 0 | **No** (Pattern dominates) |
| DecayGF | Tsetlin | 121 163 | 14.7 | **18.4** | 3e-211 | 11 795 | **16.2** | 15.4 | 0.045 | **No** |
| GuessFactor | KNN | 44 521 | **12.3** | 8.5 | 2e-84 | 88 570 | **17.7** | 14.3 | 9e-115 | **No** (GF dominates both) |
| GuessFactor | Pattern | 40 605 | 10.6 | **14.3** | 7e-69 | 92 459 | 18.2 | **23.2** | 2e-207 | **No** |
| GuessFactor | Tsetlin | 117 799 | 15.5 | **17.9** | 6e-92 | 15 156 | 19.2 | **20.4** | 0.0012 | **No** |
| KNN | Pattern | 38 485 | 9.7 | **15.8** | 6e-155 | 94 344 | 13.4 | **22.4** | 0 | **No** |
| KNN | Tsetlin | 131 820 | 12.2 | **18.1** | 0 | 804 | 21.6 | **30.0** | 6e-06 | **No** (KNN wins 0.6% only) |
| Pattern | Tsetlin | 121 591 | **21.1** | 18.8 | 2e-58 | 11 221 | **13.6** | 11.6 | 2e-07 | **No** (Pattern dominates) |
**No pair is complementary.** The pattern in nearly every row is that the more
accurate gun is the same one on *both* slices — the confidence orderings are not
aligned across guns, so "A is confident here" does not mean "A is the expert
here". The two pairs with an unequal split (DecayGF vs KNN, KNN vs Tsetlin) are
ones where the "loser" wins *only* 0.3–0.6% of samples — not a usable slice. A
pair with no complementary slices cannot form a useful composite, and none does.
`[INFERRED]` the mechanism: all five guns estimate the *same* quantity (the
enemy's intercept) from the *same* recorded trajectory; differing booleanisations
create different *noise*, not different *competence regions*, so their confidence
rankings carry no cross-gun information.
---
## 5. Task C — composite vs best single, with the shuffle control `[MEASURED]`
Held-out test battles, paired per sample. `member_oracle` = accuracy if a
perfect per-sample selector could pick any confidence member (the ceiling for
any member-picking composite; a free-aim oracle could only be higher).
| arm | accuracy | n |
|---|---:|---:|
| **Accel** (best single overall — no confidence) | **21.02%** | 133 101 |
| **Pattern** (best single *with* confidence / incumbent) | **20.45%** | 133 100 |
| WallBounce | 20.15% | 133 103 |
| AvgLead | 20.08% | 133 100 |
| **CompositeF** (faithful only: Pattern+DecayGF+KNN) | **18.68%** | 133 126 |
| **Composite** (all 5) | **18.41%** | 133 094 |
| Tsetlin | 18.18% | 132 964 |
| **CompositeFShuf** (control) | **18.06%** | 133 119 |
| **CompositeShuf** (control) | **17.61%** | 133 103 |
| GuessFactor | 15.89% | 133 108 |
| DecayGF | 14.83% | 133 104 |
| KNN | 12.32% | 133 121 |
| **member_oracle** (perfect member picker) | **41.72%** | 133 094 |
Paired comparisons (McNemar on discordant pairs):
| comparison | hits | opponents | p |
|---|---:|---:|---:|
| Composite vs **Pattern** | 24 488 | 27 215 | **3.1e-126** (composite loses) |
| CompositeF vs **Pattern** | 24 854 | 27 215 | **1.25e-133** (composite loses) |
| Composite vs its **shuffle** | 24 496 | 23 443 | 4.0e-14 (composite wins) |
| CompositeF vs its **shuffle** | 24 864 | 24 040 | 4.0e-15 (composite wins) |
| CompositeF vs Composite | — | — | (faithful-only is marginally better, +0.27pp) |
Per held-out battle (composite vs Pattern): `tr_drussgt_vs_modularbot` 12.2% vs
13.0%; `tr_drussgt_vs_spinbot` 29.0% vs 33.9%; `tr_drussgt_vs_corners` 22.3% vs
22.4%. The composite loses **all three**.
**Verdict: the design is dead offline.** The alpha-normalised confidence-weighted
composite does **not** beat the best single gun; it loses to Pattern by ~1.8–2.0pp
with p≈1e-130, and to the best single overall (Accel) by more. The shuffle control
**passes in the weak sense**: the composite is genuinely (p≈1e-14) better than a
version with the same weighting distribution but shuffled competences, so the
confidence signal is not pure noise — but the effect is ~0.6–0.8pp versus a ~2pp
deficit, i.e. real but far too small. The `member_oracle` of **41.72%** shows
enormous headroom exists — it is not reachable from these confidence signals.
---
## 6. Why it fails `[INFERRED]`
1. **Faithfulness is not competence.** A gun can be perfectly confidence-ordered
and still be worse than another gun everywhere. DecayGF is *more* faithful than
Pattern (rho 0.133 vs 0.064) yet 5.6pp less accurate, so its (correct) ordering
contributes weak votes against a stronger member.
2. **The members are not specialists.** §4 shows the confidence orderings do not
identify competence regions across guns; every gun attacks the whole input
space. The paper's win comes from members that are *good on disjoint subsets*.
3. **The signal is diluted by anti-faithful members.** GuessFactor is anti-faithful
and Tsetlin is useless; CompositeF (faithful only) is better than Composite
(18.68 vs 18.41) and more faithful (rho 0.114 vs 0.009), which confirms the
poison — but removing it still leaves the composite below Pattern.
4. **Alpha normalisation is unstable for online-learning guns.** Eq 7 sets
`alpha_t` to the train range; GuessFactor's histogram accumulates unboundedly
(alpha≈1.5e4 here), so its normalised confidence is tiny on test and it almost
never casts a decisive vote. `[INFERRED]` this mutes the very member whose
confidence was measured anti-faithful.
---
## 7. Limits and caveats
- **Open-loop corpus.** `[MEASURED]`/`[INFERRED]` The fixtures are recorded
trajectories; the enemy never reacts to the composite's (or any) bullets, and
the `tr-bridge` battles were recorded while **Pattern's selector** was shooting.
A gun that behaves like Pattern is therefore favoured in *framing*. This cannot
rescue the composite: it loses to Pattern **and** to Accel/AvgLead/WallBounce, so
the deficit is not a framing artefact. Per `docs/offline_harness_trust.md`, this
is exactly the closed-loop class of question where the offline harness is
**veto-only** — a negative kills the design, a positive would have proved nothing.
- **Coverage gap (documented LIMIT).** `[MEASURED]` The offline range builds guns
0..13; TMPATTERN (14) and TMHORIZON (15) are not in it
(`docs/offline_harness_trust.md` §2.8). TMHorizon's `sideConf` is instrumented in
source but not measured here. Since Tsetlin's c_max is already useless and no pair
composes, the omission does not change the verdict.
- **One composite architecture.** `[INFERRED]` The vote uses each gun's scalar
confidence at the bin of its own aim. The paper's members also expose a full
class distribution (GF/KNN do). Feeding those full distributions could refine the
composite, but §4 shows the core premise — cross-gun competence specialisation —
is absent, so a refined vote has no signal to exploit.
---
## 8. Recommendation
**Do not build or A/B the composite.** The offline gate is a veto and it vetoes.
Two by-products worth keeping:
- **GuessFactor / anti-faithful warning.** GuessFactor's own histogram peak is
anti-faithful on this corpus; do **not** use it as a competence signal (e.g. for
a confidence-gated firing decision or a power policy). Pattern, DecayGF and KNN
confidences are faithful and *could* gate a shot ("do not fire when not
confident") — a different, narrower mechanism than the composite, and still
subject to the live veto.
- **The selector conclusion is reinforced.** The rack's problem is not the
*decision statistic* the selector uses (rolling rate vs per-sample confidence):
even a per-sample intrinsic confidence, applied cross-gun, cannot beat Pattern.
The rack does not contain complementary specialists.
---
## 9. Reproduction `[MEASURED]`
```bash
# 1. record per-sample confidence + hit (≈ 6.5 min; fresh rack per fixture)
nim c -d:release --nimcache:/tmp/nc_j104 --path:common_libs \
-o:/tmp/measure_tmc common_libs/tests/measure_tmcomposites.nim
/tmp/measure_tmc --out /tmp/tmc_full.jsonl \
--train tools/fixtures/drussgt_vs_crazy.jsonl tools/fixtures/drussgt_vs_spinbot.jsonl \
tools/fixtures/drussgt_vs_ramfire.jsonl tools/fixtures/drussgt_vs_corners.jsonl \
--test tools/fixtures/tr_drussgt_vs_modularbot.jsonl tools/fixtures/tr_drussgt_vs_spinbot.jsonl \
tools/fixtures/tr_drussgt_vs_corners.jsonl
# 2. analyse (≈ 40 s); committed result is common_libs/tests/fixtures/tmcomposites_gate.json
python3 common_libs/tests/analyze_tmcomposites.py \
--input /tmp/tmc_full.jsonl --json common_libs/tests/fixtures/tmcomposites_gate.json
```
The recorder threads a new per-sample `confidence` field through
`GunPrediction`/`FeedbackEvent`/`VirtualBullet` (`gun_interface.nim`,
`virtual_bullets.nim`) and populates it in `guess_factor.nim`, `decay_gf.nim`,
`knn_gun.nim`, `pattern_matcher.nim`, `tsetlin.nim`, `tm_horizon.nim`; the field
defaults to 0.0, so every existing caller and test is unchanged. Verified
`test_gun_harness`, `test_vbullet_metric`, `test_wave_pairing`,
`test_tm_horizon`, `test_pattern_radial_offset`, `test_range_rack_parity`,
`test_selector_tiebreak` all pass.