TMComposites gate: per-gun confidence faithful for 3 guns; no pair composes

Adds a per-sample intrinsic-confidence field (GunPrediction.confidence,
threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF,
KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that
reproduce the paper's Figure 2 per gun and its Eq-8 composite.

Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun):
- FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak).
- GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001).
- No pair of guns specialises complementarily: the same gun dominates both
  high-confidence slices in every pair.
- Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern
  20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses.
  Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of
  competence is real but ~2pp short. Offline veto: design is dead.

See docs/tmcomposites_gate.md.
This commit is contained in:
2026-09-25 22:02:28 +02:00
parent d0750ab020
commit f58d65d2e8
12 changed files with 2666 additions and 11 deletions
+8 -1
View File
@@ -413,7 +413,14 @@ proc predict*(g: var TsetlinGun, state: WorldState, bulletSpeed: float): GunPred
alive: true,
)
GunPrediction(x: predX, y: predY)
GunPrediction(
x: predX,
y: predY,
# TMComposites Eq 4: the two output teams' clamped clause sums are (vx, vy);
# their magnitude is how hard the machine is voting to move the correction.
# Warm-up fallback (window not full) leaves the default 0.0 = no vote.
confidence: hypot(vx, vy),
)
proc onResult*(g: var TsetlinGun, e: FeedbackEvent) =
inc g.shotCount