f58d65d2e8
Adds a per-sample intrinsic-confidence field (GunPrediction.confidence, threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF, KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that reproduce the paper's Figure 2 per gun and its Eq-8 composite. Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun): - FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak). - GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001). - No pair of guns specialises complementarily: the same gun dominates both high-confidence slices in every pair. - Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern 20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses. Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of competence is real but ~2pp short. Offline veto: design is dead. See docs/tmcomposites_gate.md.
319 lines
18 KiB
Markdown
319 lines
18 KiB
Markdown
# TMComposites gate: is each gun's intrinsic confidence faithful, and are the guns complementary specialists?
|
||
|
||
**Scope.** Test the mechanism of *TMComposites: Plug-and-Play Collaboration
|
||
Between Specialized Tsetlin Machines* (Granmo, arXiv:2309.04801v2, §3) on our
|
||
gun rack: (A) is each gun's per-sample confidence faithful to its own accuracy?
|
||
(B) are the guns complementary specialists? (C) does an Eq-8 alpha-normalised
|
||
confidence-weighted composite beat the best single gun offline, and is any gain
|
||
attributable to competence (shuffle control)? This is the one untested idea that
|
||
could plausibly beat `onlyPattern`, whose selector was measured **negative value**
|
||
(`docs/selector_negative_value.md`, `docs/gun_rack_analysis.md`, commit `e0666a5`)
|
||
because it decides with a rolling hit-rate instead of a per-sample signal.
|
||
|
||
**Evidence tags.** `[MEASURED]` = read from the committed result
|
||
`common_libs/tests/fixtures/tmcomposites_gate.json`, reproduced by the named
|
||
command, or a source fact. `[INFERRED]` = reasoning from those measurements.
|
||
|
||
---
|
||
|
||
## 0. Direct answers
|
||
|
||
1. **Which guns are faithfully confident?** `[MEASURED]`
|
||
**Pattern (weak), DecayGF (strong), KNN (moderate)** are faithful: rank samples
|
||
by the gun's own confidence and accuracy rises. **GuessFactor is
|
||
ANTI-faithful** — it is more accurate when it is *less* confident. **Tsetlin's
|
||
class-sum max is useless** (Spearman 0.001, p=0.68). The eight deterministic
|
||
geometric guns (HeadOn, Linear, Circular, WallBounce, Accel, StopShot,
|
||
Displace, AvgLead) expose **no per-sample confidence at all**.
|
||
|
||
2. **Do any pairs specialise complementarily?** `[MEASURED]` **No.** In every one
|
||
of the 10 pairs, on the slice where A's normalised confidence beats B's, A is
|
||
*not* the more accurate gun — or the same gun dominates both slices (e.g.
|
||
GuessFactor is more accurate than DecayGF and than KNN even on their own
|
||
high-confidence slices). The paper's premise — one member's weakness is
|
||
another's strength, decided by confidence — does not hold on our rack. Our
|
||
guns are different *estimators of the same target*, not different *specialists*.
|
||
|
||
3. **Does the composite beat the best single gun?** `[MEASURED]` **No, and the
|
||
design is dead offline.** The Eq-8 composite scores **18.41%** vs Pattern's
|
||
**20.45%** (McNemar p=3.1e-126); the faithful-only variant (Pattern + DecayGF +
|
||
KNN) scores **18.68%**, still losing to Pattern by 1.77pp (p=1.25e-133). Both
|
||
lose on all three held-out battles. The **confidence-shuffle control passes in
|
||
the weak sense** — the composite beats its own shuffle (18.41 vs 17.61,
|
||
p=4.0e-14) — so the weighting carries a *real but tiny* competence signal; it
|
||
is simply nowhere near enough to beat Pattern. Per the `docs/offline_harness_trust.md`
|
||
rule (offline is veto-only), this negative **kills the design**; do not take a
|
||
composite to a live A/B.
|
||
|
||
---
|
||
|
||
## 1. What was measured `[MEASURED]`
|
||
|
||
The offline gun range (`common_libs/gun_harness/offline_range.nim` +
|
||
`virtual_bullets.nim`), driven by a new recorder
|
||
`common_libs/tests/measure_tmcomposites.nim`. For every resolved virtual bullet
|
||
it records `(gun, tick, powerBin, confidence, hit, aim-bearing-relative-to-LOS,
|
||
range)`. Ground truth is the shipped `bmPath` virtual-bullet metric (18 px hit
|
||
radius — `BotRadius`), reproduced exactly per bullet, not a rolling rate.
|
||
|
||
| | fixtures | ticks | note |
|
||
|---|---|---:|---|
|
||
| **train** (alpha calibration only) | `drussgt_vs_crazy`, `drussgt_vs_spinbot`, `drussgt_vs_ramfire`, `drussgt_vs_corners` | 22 242 | classic-Robocode DrussGT |
|
||
| **test** (all reported numbers) | `tr_drussgt_vs_modularbot`, `tr_drussgt_vs_spinbot`, `tr_drussgt_vs_corners` | 33 425 | Tank-Royale bridge captures |
|
||
|
||
Each fixture is replayed with a **fresh rack** (as `run_range` does); the test
|
||
battles are **held out by battle**, never by tick. The dump is 3 763 298 rows;
|
||
per-gun test `n ≈ 133 000` (33 425 ticks × 4 power bins).
|
||
|
||
**Per-gun confidence signal defined in source** `[MEASURED]` (the new
|
||
`GunPrediction.confidence` field, `common_libs/gun_harness/gun_interface.nim`):
|
||
|
||
| gun | intrinsic per-sample signal | source |
|
||
|---|---|---|
|
||
| GuessFactor | peak GF bin weight `max_i bins[i]` (Eq 4 analogue) | `guess_factor.nim` |
|
||
| DecayGF | peak decayed GF bin weight | `decay_gf.nim` |
|
||
| KNN | peak Gaussian density over GF candidates (`bestScore`) | `knn_gun.nim` |
|
||
| Pattern | match quality `1/(1+bestMatchCost)` | `pattern_matcher.nim` |
|
||
| Tsetlin | magnitude of the clamped clause-sum vote `hypot(vx,vy)` | `tsetlin.nim` |
|
||
| TMHorizon | side-class margin `|votes1-votes0|/(2*half)` | `tm_horizon.nim` (instrumented; see §7) |
|
||
| 8 geometric guns | none — confidence 0.0 | `head_on/linear/circular/wall_bounce/accel_predictor/stop_shot/displacement/averaged_lead` |
|
||
|
||
**WARNING — this is exactly the mechanism class the task forbids:** all five are
|
||
per-sample statistics of the gun's own internal state at prediction time. None is
|
||
a rolling accuracy or hit-rate average.
|
||
|
||
---
|
||
|
||
## 2. The mechanism (paper §3, Eqs 4/6/7/8)
|
||
|
||
- Member `t` outputs class sums `c^i_{t,d}` (Eq 6); confidence is `c_max` (Eq 4).
|
||
- Eq 7 normalises by `alpha_t = max_{d,i}(c) - min_{d,i}(c)`.
|
||
- Eq 8: `y_d = argmax_i sum_t (1/alpha_t) c^i_{t,d}`.
|
||
|
||
Adapted to an **angle output**: the class is a 0.5° aim bin relative to the
|
||
fire-time line of sight (`[-45°, +45°]`, 180 bins); each confident gun casts its
|
||
alpha-normalised confidence into the bin of its own aim; the argmax bin wins, and
|
||
the aim distance is the confidence-weighted mean distance of that bin's voters.
|
||
`alpha_t` is calibrated on the **train** fixtures and frozen for test, so the test
|
||
composite never sees test data when choosing its weights (Eq 7's `X` = train).
|
||
Two composites are scored:
|
||
|
||
- **Composite** = all five confidence guns {Tsetlin, GuessFactor, Pattern, DecayGF, KNN}.
|
||
- **CompositeF** = only the guns §3 finds faithful {Pattern, DecayGF, KNN}.
|
||
|
||
---
|
||
|
||
## 3. Task A — confidence faithfulness `[MEASURED]`
|
||
|
||
Test samples ranked by each gun's own confidence. `rho` = Spearman(confidence,
|
||
hit); `bottom/top` = accuracy in the lower/upper confidence half; the two-prop
|
||
p is for top vs bottom.
|
||
|
||
| gun | n | nonzero | base | rho | rho p | bottom% | top% | verdict |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|---|
|
||
| Accel | 133 101 | 0 | 21.02 | — | — | 14.47 | 27.57 | NO SIGNAL |
|
||
| **Pattern** | 133 100 | 132 860 | 20.45 | **+0.064** | 4.1e-121 | 18.61 | **22.28** | **FAITHFUL (weak)** |
|
||
| WallBounce | 133 103 | 0 | 20.15 | — | — | 14.84 | 25.46 | NO SIGNAL |
|
||
| AvgLead | 133 100 | 0 | 20.08 | — | — | 14.43 | 25.73 | NO SIGNAL |
|
||
| StopShot | 133 078 | 0 | 18.21 | — | — | 12.82 | 23.61 | NO SIGNAL |
|
||
| Tsetlin | 132 964 | 105 353 | 18.18 | +0.001 | 0.68 | 17.94 | 18.42 | **USELESS** |
|
||
| Linear | 133 104 | 0 | 17.31 | — | — | 13.04 | 21.57 | NO SIGNAL |
|
||
| Circular | 133 107 | 0 | 17.27 | — | — | 12.68 | 21.86 | NO SIGNAL |
|
||
| **GuessFactor** | 133 108 | 133 108 | 15.89 | **−0.067** | 5.1e-132 | 18.14 | **13.65** | **ANTI-FAITHFUL** |
|
||
| **DecayGF** | 133 104 | 133 104 | 14.83 | **+0.133** | <1e-300 | 10.79 | **18.86** | **FAITHFUL (strong)** |
|
||
| Displace | 133 105 | 0 | 13.95 | — | — | 11.42 | 16.47 | NO SIGNAL |
|
||
| **KNN** | 133 121 | 132 685 | 12.32 | **+0.090** | 2.1e-237 | 9.65 | **14.99** | **FAITHFUL** |
|
||
| HeadOn | 133 162 | 0 | 6.18 | — | — | 7.09 | 5.27 | NO SIGNAL |
|
||
|
||
Accuracy-vs-confidence curves (deciles of confidence, low→high), the paper's
|
||
Figure 2 reproduced per gun:
|
||
|
||
```
|
||
Pattern 14% 16% 19% 23% 22% 22% 21% 23% 22% 23% rises, then flat
|
||
DecayGF 10% 9% 11% 15% 9% 11% 16% 21% 23% 23% rises (noisy)
|
||
KNN 10% 9% 9% 10% 10% 12% 13% 15% 16% 19% monotone rise
|
||
GuessFactor 16% 21% 21% 15% 18% 14% 18% 14% 13% 10% FALLS
|
||
Tsetlin 13% 25% 12% 21% 19% 20% 16% 18% 20% 19% flat/noise
|
||
```
|
||
|
||
**Reading.** DecayGF and KNN are the textbook faithful shapes (accuracy climbs
|
||
with confidence). Pattern is faithful but weakly — its one prior is strong
|
||
(`18.61% → 22.28%`) and then flattens, i.e. the match cost discriminates
|
||
"no/poor match" from "match" but not much *between* matches. GuessFactor is the
|
||
surprise and the most important negative: its histogram peak is **anti-faithful**,
|
||
consistent with the earlier finding that the GF code path is degenerate on these
|
||
`tr-bridge` captures. Tsetlin's `c_max` (clause-sum magnitude) carries **no**
|
||
information about whether the shot hits — `[INFERRED]` because the TM's
|
||
regression correction is trained on a residual and never learns the surfer
|
||
(`tsetlin.nim` header: every variant sat at chance vs a shuffled-control).
|
||
|
||
---
|
||
|
||
## 4. Task B — complementary specialists? `[MEASURED]`
|
||
|
||
For each pair we split test samples by which gun has the higher **alpha-normalised**
|
||
confidence and measure **both** guns on each slice. `A|A` = A's accuracy on the
|
||
slice A wins, `B|A` = B's accuracy on that same slice; `p` is a paired McNemar.
|
||
|
||
| A | B | A_wins | A\|A | B\|A | p_A | B_wins | A\|B | B\|B | p_B | complementary |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---|
|
||
| DecayGF | GuessFactor | 43 493 | 19.1 | **20.0** | 4e-14 | 89 608 | 12.8 | **13.9** | 8e-46 | **No** (GF dominates both) |
|
||
| DecayGF | KNN | 436 | 30.0 | **30.7** | 0.25 | 132 656 | **14.8** | 12.3 | 4e-99 | **No** (DecayGF wins 0.3% only) |
|
||
| DecayGF | Pattern | 26 788 | 11.0 | **14.8** | 1e-51 | 106 279 | 15.8 | **21.9** | 0 | **No** (Pattern dominates) |
|
||
| DecayGF | Tsetlin | 121 163 | 14.7 | **18.4** | 3e-211 | 11 795 | **16.2** | 15.4 | 0.045 | **No** |
|
||
| GuessFactor | KNN | 44 521 | **12.3** | 8.5 | 2e-84 | 88 570 | **17.7** | 14.3 | 9e-115 | **No** (GF dominates both) |
|
||
| GuessFactor | Pattern | 40 605 | 10.6 | **14.3** | 7e-69 | 92 459 | 18.2 | **23.2** | 2e-207 | **No** |
|
||
| GuessFactor | Tsetlin | 117 799 | 15.5 | **17.9** | 6e-92 | 15 156 | 19.2 | **20.4** | 0.0012 | **No** |
|
||
| KNN | Pattern | 38 485 | 9.7 | **15.8** | 6e-155 | 94 344 | 13.4 | **22.4** | 0 | **No** |
|
||
| KNN | Tsetlin | 131 820 | 12.2 | **18.1** | 0 | 804 | 21.6 | **30.0** | 6e-06 | **No** (KNN wins 0.6% only) |
|
||
| Pattern | Tsetlin | 121 591 | **21.1** | 18.8 | 2e-58 | 11 221 | **13.6** | 11.6 | 2e-07 | **No** (Pattern dominates) |
|
||
|
||
**No pair is complementary.** The pattern in nearly every row is that the more
|
||
accurate gun is the same one on *both* slices — the confidence orderings are not
|
||
aligned across guns, so "A is confident here" does not mean "A is the expert
|
||
here". The two pairs with an unequal split (DecayGF vs KNN, KNN vs Tsetlin) are
|
||
ones where the "loser" wins *only* 0.3–0.6% of samples — not a usable slice. A
|
||
pair with no complementary slices cannot form a useful composite, and none does.
|
||
`[INFERRED]` the mechanism: all five guns estimate the *same* quantity (the
|
||
enemy's intercept) from the *same* recorded trajectory; differing booleanisations
|
||
create different *noise*, not different *competence regions*, so their confidence
|
||
rankings carry no cross-gun information.
|
||
|
||
---
|
||
|
||
## 5. Task C — composite vs best single, with the shuffle control `[MEASURED]`
|
||
|
||
Held-out test battles, paired per sample. `member_oracle` = accuracy if a
|
||
perfect per-sample selector could pick any confidence member (the ceiling for
|
||
any member-picking composite; a free-aim oracle could only be higher).
|
||
|
||
| arm | accuracy | n |
|
||
|---|---:|---:|
|
||
| **Accel** (best single overall — no confidence) | **21.02%** | 133 101 |
|
||
| **Pattern** (best single *with* confidence / incumbent) | **20.45%** | 133 100 |
|
||
| WallBounce | 20.15% | 133 103 |
|
||
| AvgLead | 20.08% | 133 100 |
|
||
| **CompositeF** (faithful only: Pattern+DecayGF+KNN) | **18.68%** | 133 126 |
|
||
| **Composite** (all 5) | **18.41%** | 133 094 |
|
||
| Tsetlin | 18.18% | 132 964 |
|
||
| **CompositeFShuf** (control) | **18.06%** | 133 119 |
|
||
| **CompositeShuf** (control) | **17.61%** | 133 103 |
|
||
| GuessFactor | 15.89% | 133 108 |
|
||
| DecayGF | 14.83% | 133 104 |
|
||
| KNN | 12.32% | 133 121 |
|
||
| **member_oracle** (perfect member picker) | **41.72%** | 133 094 |
|
||
|
||
Paired comparisons (McNemar on discordant pairs):
|
||
|
||
| comparison | hits | opponents | p |
|
||
|---|---:|---:|---:|
|
||
| Composite vs **Pattern** | 24 488 | 27 215 | **3.1e-126** (composite loses) |
|
||
| CompositeF vs **Pattern** | 24 854 | 27 215 | **1.25e-133** (composite loses) |
|
||
| Composite vs its **shuffle** | 24 496 | 23 443 | 4.0e-14 (composite wins) |
|
||
| CompositeF vs its **shuffle** | 24 864 | 24 040 | 4.0e-15 (composite wins) |
|
||
| CompositeF vs Composite | — | — | (faithful-only is marginally better, +0.27pp) |
|
||
|
||
Per held-out battle (composite vs Pattern): `tr_drussgt_vs_modularbot` 12.2% vs
|
||
13.0%; `tr_drussgt_vs_spinbot` 29.0% vs 33.9%; `tr_drussgt_vs_corners` 22.3% vs
|
||
22.4%. The composite loses **all three**.
|
||
|
||
**Verdict: the design is dead offline.** The alpha-normalised confidence-weighted
|
||
composite does **not** beat the best single gun; it loses to Pattern by ~1.8–2.0pp
|
||
with p≈1e-130, and to the best single overall (Accel) by more. The shuffle control
|
||
**passes in the weak sense**: the composite is genuinely (p≈1e-14) better than a
|
||
version with the same weighting distribution but shuffled competences, so the
|
||
confidence signal is not pure noise — but the effect is ~0.6–0.8pp versus a ~2pp
|
||
deficit, i.e. real but far too small. The `member_oracle` of **41.72%** shows
|
||
enormous headroom exists — it is not reachable from these confidence signals.
|
||
|
||
---
|
||
|
||
## 6. Why it fails `[INFERRED]`
|
||
|
||
1. **Faithfulness is not competence.** A gun can be perfectly confidence-ordered
|
||
and still be worse than another gun everywhere. DecayGF is *more* faithful than
|
||
Pattern (rho 0.133 vs 0.064) yet 5.6pp less accurate, so its (correct) ordering
|
||
contributes weak votes against a stronger member.
|
||
2. **The members are not specialists.** §4 shows the confidence orderings do not
|
||
identify competence regions across guns; every gun attacks the whole input
|
||
space. The paper's win comes from members that are *good on disjoint subsets*.
|
||
3. **The signal is diluted by anti-faithful members.** GuessFactor is anti-faithful
|
||
and Tsetlin is useless; CompositeF (faithful only) is better than Composite
|
||
(18.68 vs 18.41) and more faithful (rho 0.114 vs 0.009), which confirms the
|
||
poison — but removing it still leaves the composite below Pattern.
|
||
4. **Alpha normalisation is unstable for online-learning guns.** Eq 7 sets
|
||
`alpha_t` to the train range; GuessFactor's histogram accumulates unboundedly
|
||
(alpha≈1.5e4 here), so its normalised confidence is tiny on test and it almost
|
||
never casts a decisive vote. `[INFERRED]` this mutes the very member whose
|
||
confidence was measured anti-faithful.
|
||
|
||
---
|
||
|
||
## 7. Limits and caveats
|
||
|
||
- **Open-loop corpus.** `[MEASURED]`/`[INFERRED]` The fixtures are recorded
|
||
trajectories; the enemy never reacts to the composite's (or any) bullets, and
|
||
the `tr-bridge` battles were recorded while **Pattern's selector** was shooting.
|
||
A gun that behaves like Pattern is therefore favoured in *framing*. This cannot
|
||
rescue the composite: it loses to Pattern **and** to Accel/AvgLead/WallBounce, so
|
||
the deficit is not a framing artefact. Per `docs/offline_harness_trust.md`, this
|
||
is exactly the closed-loop class of question where the offline harness is
|
||
**veto-only** — a negative kills the design, a positive would have proved nothing.
|
||
- **Coverage gap (documented LIMIT).** `[MEASURED]` The offline range builds guns
|
||
0..13; TMPATTERN (14) and TMHORIZON (15) are not in it
|
||
(`docs/offline_harness_trust.md` §2.8). TMHorizon's `sideConf` is instrumented in
|
||
source but not measured here. Since Tsetlin's c_max is already useless and no pair
|
||
composes, the omission does not change the verdict.
|
||
- **One composite architecture.** `[INFERRED]` The vote uses each gun's scalar
|
||
confidence at the bin of its own aim. The paper's members also expose a full
|
||
class distribution (GF/KNN do). Feeding those full distributions could refine the
|
||
composite, but §4 shows the core premise — cross-gun competence specialisation —
|
||
is absent, so a refined vote has no signal to exploit.
|
||
|
||
---
|
||
|
||
## 8. Recommendation
|
||
|
||
**Do not build or A/B the composite.** The offline gate is a veto and it vetoes.
|
||
|
||
Two by-products worth keeping:
|
||
|
||
- **GuessFactor / anti-faithful warning.** GuessFactor's own histogram peak is
|
||
anti-faithful on this corpus; do **not** use it as a competence signal (e.g. for
|
||
a confidence-gated firing decision or a power policy). Pattern, DecayGF and KNN
|
||
confidences are faithful and *could* gate a shot ("do not fire when not
|
||
confident") — a different, narrower mechanism than the composite, and still
|
||
subject to the live veto.
|
||
- **The selector conclusion is reinforced.** The rack's problem is not the
|
||
*decision statistic* the selector uses (rolling rate vs per-sample confidence):
|
||
even a per-sample intrinsic confidence, applied cross-gun, cannot beat Pattern.
|
||
The rack does not contain complementary specialists.
|
||
|
||
---
|
||
|
||
## 9. Reproduction `[MEASURED]`
|
||
|
||
```bash
|
||
# 1. record per-sample confidence + hit (≈ 6.5 min; fresh rack per fixture)
|
||
nim c -d:release --nimcache:/tmp/nc_j104 --path:common_libs \
|
||
-o:/tmp/measure_tmc common_libs/tests/measure_tmcomposites.nim
|
||
/tmp/measure_tmc --out /tmp/tmc_full.jsonl \
|
||
--train tools/fixtures/drussgt_vs_crazy.jsonl tools/fixtures/drussgt_vs_spinbot.jsonl \
|
||
tools/fixtures/drussgt_vs_ramfire.jsonl tools/fixtures/drussgt_vs_corners.jsonl \
|
||
--test tools/fixtures/tr_drussgt_vs_modularbot.jsonl tools/fixtures/tr_drussgt_vs_spinbot.jsonl \
|
||
tools/fixtures/tr_drussgt_vs_corners.jsonl
|
||
|
||
# 2. analyse (≈ 40 s); committed result is common_libs/tests/fixtures/tmcomposites_gate.json
|
||
python3 common_libs/tests/analyze_tmcomposites.py \
|
||
--input /tmp/tmc_full.jsonl --json common_libs/tests/fixtures/tmcomposites_gate.json
|
||
```
|
||
|
||
The recorder threads a new per-sample `confidence` field through
|
||
`GunPrediction`/`FeedbackEvent`/`VirtualBullet` (`gun_interface.nim`,
|
||
`virtual_bullets.nim`) and populates it in `guess_factor.nim`, `decay_gf.nim`,
|
||
`knn_gun.nim`, `pattern_matcher.nim`, `tsetlin.nim`, `tm_horizon.nim`; the field
|
||
defaults to 0.0, so every existing caller and test is unchanged. Verified
|
||
`test_gun_harness`, `test_vbullet_metric`, `test_wave_pairing`,
|
||
`test_tm_horizon`, `test_pattern_radial_offset`, `test_range_rack_parity`,
|
||
`test_selector_tiebreak` all pass.
|