diff --git a/docs/gun_rack_analysis.md b/docs/gun_rack_analysis.md index c014b15..f7c4b7b 100644 --- a/docs/gun_rack_analysis.md +++ b/docs/gun_rack_analysis.md @@ -9,6 +9,13 @@ predated per-gun real-hit attribution, the offline gun range, and the live DrussGT boss. No number from the old version is retained unless it was re-measured below. +**Corrected after later runs (same day).** Three claims in the first version of +this rewrite were revised by follow-up measurements: the virtual-vs-real +correlation is a poor, sign-unstable ranker rather than an "inversion" (§2); +the offline==online acceptance test is fixed, not flaky (§1.1); and pruning the +below-overall guns was tested and does not help (§3). Each correction is +restated plainly at the point of the old claim. + Evidence tags used throughout: - **[MEASURED]** — I read it from a recorded artifact (commit message, `/tmp` @@ -55,14 +62,26 @@ header comment; commit `974528d`). if the target dies during that `go()`, the final tick's spawn+resolution is skipped. **[MEASURED]** (commit `974528d`; a passing run is preserved in `/tmp/ab_logs3/test_acceptance.log`: 12/12, 534-tick round). -- **Honest caveat — the acceptance test is currently FLAKY.** On the - *unmodified HEAD* source it normally reaches only **11/12**, e.g. KNN 81 - online vs 71 offline, and the mismatching gun moves between runs (KNN, then - WallBounce). It is a live/offline boundary race, pre-existing, and not caused - by the selector work (the replay never calls the selector). Treat - "offline == online" as **strong but not exact until the race is fixed**. - **[MEASURED]** (commit `2c94dc2`). The 12/12 runs above are real; they were - lucky runs. +- **Acceptance test: the equivalence is now proven and stable.** An earlier + version of this report recorded the test as flaky — typically **11/12** on + unmodified HEAD, with the mismatching gun moving between runs (KNN, then + WallBounce) — and guessed it was a live/offline boundary race. That guess was + **wrong**. The root cause was a **real replay bug**: the offline replay + spawned gun 13 (TMSelect) while the live rack has `EnableTmSelector = false` + and never does. The shared `VirtualTracker` ring is **order-sensitive**, so + gun 13's extra 4 bullets/tick permuted the per-tick **resolution order** of + every other gun, shifting the learning guns' observations. Closing gun 13's + ready gate offline made the live and offline KNN traces **byte-identical** + (904/904 lines, empty diff). The fix mirrors the live rack in the replay — no + tick exclusion, no tolerance loosening. Stability: **5/5 consecutive runs + report 12/12 exact, each with the death boundary included.** So + "offline == online" is exact on these runs. A flaky proof had hidden a real + bug. **[MEASURED]**. +- **General lesson for rack A/B.** Because the ring is order-sensitive, any + rack A/B that disables a gun also removes that gun's **4 spawns/tick** from + the shared ring, which perturbs the resolution order — and therefore the + learning observations — of every other gun. That is a confound to record for + anyone repeating these experiments. **[INFERRED]**. ### 1.2 The fixture sets and what each is good for @@ -187,7 +206,7 @@ ground truth for config decisions. **[MEASURED]** (re-aggregated from --- -## 2. The metric lesson: virtual hit rate is NOT a proxy for real hit rate +## 2. The metric lesson: virtual hit rate is a poor ranker, not a proxy for real hit rate This is the most important conceptual result of the night and it invalidates a naive reading of every offline table in this report. @@ -201,8 +220,13 @@ hit rate is Spearman(virtual rank, real rank) = -0.374 (n = 13 guns, all with ≥10 real shots) ``` -It is not weak — it is **inverted**. The guns with the highest virtual rates -have among the lowest real rates, and vice versa: +The correlation is **weak and sign-unstable** — the honest headline is that +virtual hit rate is a **poor ranker**, not a reliable inverse. The table below +is what a poor ranker looks like: on this run set the ordering it produces +tracks the opposite of the real ordering, but that does not hold on other run +sets (see the full measurement set after the table). An earlier version of this +report stated this as a clean inversion; a later 15-run measurement refuted +that. | Gun | Selected (ticks) | Real hits/shots | Real % | Virtual % | |---|---:|---:|---:|---:| @@ -222,8 +246,10 @@ have among the lowest real rates, and vice versa: Read the top and bottom: **Tsetlin, WallBounce and StopShot have the highest virtual rates (12.6–12.9%) and near-bottom real rates (5.8–6.2%); Linear and -KNN sit at 10.2% / 7.5% virtual but 10.7% / 9.0% real.** The ranking the -virtual metric produces is not merely uninformative, it points the wrong way. +KNN sit at 10.2% / 7.5% virtual but 10.7% / 9.0% real.** On this run set the +virtual ordering inverts the real one — but because the sign flips on other +run sets (next paragraph), the safe reading is that the virtual ranking is +**uninformative about the real ranking**, not that it is reliably inverted. **Why this matters for selection.** What has kept the rack alive is the selector's **floor/tie hedging**, not its ranking: removing the floor @@ -232,14 +258,28 @@ selector's **floor/tie hedging**, not its ranking: removing the floor (`/tmp/compare.py`). So the selector is useful because it refuses to commit to a bad field, not because its virtual-rate ordering is good. -**The virtual metric appears anti-correlated no matter which config you pick.** -Measured Spearman per config (5-run 12-round A/B, `/tmp/ab_logs3/FINAL_AB.txt`): -`absolute+point` −0.371, `absolute+path` −0.073, `relative+point` −0.037, -`relative+path` +0.522. But on the large 13-run base set the shipped config -(`relative+path`) is **−0.374**. The sign **flips between run sets**, which is -itself the finding: the correlation is unstable, so no ranking rule built on -it can be trusted. **[MEASURED]** + **[INFERRED]** (the flip is measured; the -conclusion is reasoning). +**The headline: across run sets the correlation is sign-unstable, so it is near +zero on average — not robustly negative.** The full set of independent +Spearman measurements (13-run source `/tmp/agg2.py base`; 5-run configs +`/tmp/ab_logs3/FINAL_AB.txt`; 15-run paired baseline §3) is: + +| Run set / aggregation | Spearman | +|---|---:| +| 13-run base, shipped `relative+path` | **−0.374** | +| 15-run paired baseline (different but equally defensible aggregation) | **+0.335** | +| 5-run 12-round A/B, `relative+path` | +0.522 | +| 5-run A/B, `absolute+point` | −0.371 | +| 5-run A/B, `absolute+path` | −0.073 | +| 5-run A/B, `relative+point` | −0.037 | + +Two **opposite signs on large samples** (−0.374 over 13 runs, +0.335 over 15 +runs, all over the same 13 guns) mean virtual hit rate is **not** a reliable +inverse of real hit rate. It is a **poor ranker**: weak correlation, sign +flipping between run sets, near zero on average. The practical conclusion is +unchanged — do not build a ranking rule on it — but an **earlier version of +this report overstated the mechanism as an inversion**; the later 15-run +measurement refuted that. **[MEASURED]** + **[INFERRED]** (the six numbers are +measured; "poor ranker / near zero on average" is the reasoning). ### 2.1 The metric A/B: point vs path (this one is real, and it is selection) @@ -327,8 +367,11 @@ The final per-gun table, shipped config, **13 runs vs the live DrussGT boss, 3,612 server-side shots, 6.95% overall** (server sidecar; per-run rates 6.77 / 5.90 / 6.57 / 6.34 / 6.10 / 7.59 / 2.90 / 7.95 / 7.49 / 9.18 / 8.44 / 7.48 / 6.42%; 251 dmg/run). Per-gun rows are the bot-side attribution over the same -runs. **[MEASURED]** (commit `2c94dc2`; `/tmp/gun_stats_base_r*.jsonl`, -re-aggregated with `/tmp/agg2.py base`; `/tmp/events_base_r*.json`). +runs. The **same binary** on a different 15-run set gives **6.18%** (events +6.16%, 200 dmg/run, §3 pruning baseline), so every rate here is quoted with its +run count — a single figure is not definitive. **[MEASURED]** (commit `2c94dc2`; +`/tmp/gun_stats_base_r*.jsonl`, re-aggregated with `/tmp/agg2.py base`; +`/tmp/events_base_r*.json`). | Verdict | Gun | Real hits/shots | Real % | Virtual % | Selected | |---|---|---:|---:|---:|---:| @@ -342,8 +385,8 @@ re-aggregated with `/tmp/agg2.py base`; `/tmp/events_base_r*.json`). | MARGINAL | DecayGF | 3/47 | 6.4 | 9.2 | 1,360 | | MARGINAL | WallBounce | 18/288 | 6.2 | 12.9 | 7,618 | | MARGINAL | StopShot | 8/132 | 6.1 | 12.6 | 3,348 | -| BELOW | Tsetlin | 6/103 | 5.8 | 12.9 | 3,284 | -| BELOW | Displace | 4/75 | 5.3 | 12.3 | 2,664 | +| BELOW — KEEP | Tsetlin | 6/103 | 5.8 | 12.9 | 3,284 | +| BELOW — KEEP | Displace | 4/75 | 5.3 | 12.3 | 2,664 | | **FLOOR — STAYS** | HeadOn | 47/898 | 5.2 | 8.6 | 16,975 | **Verdicts.** @@ -352,12 +395,24 @@ re-aggregated with `/tmp/agg2.py base`; `/tmp/events_base_r*.json`). above the 6.95% overall, yet their virtual rates are mid-pack to low: the metric's three favourites (Tsetlin 12.9%, WallBounce 12.9%, StopShot 12.6%) are near the *bottom* of the real ranking, while the real leader (Linear) sits - at 10.2% virtual. Further evidence the virtual ranking is inverted. + at 10.2% virtual. Further evidence the virtual ranking is uninformative about + the real ranking (and, on this run set, roughly its opposite). - **MARGINAL: GuessFactor, DecayGF, WallBounce, StopShot.** Within ~1 pp of overall on small N (47–288 shots). They are not obviously worth deleting, but they have not earned a larger share. -- **BELOW OVERALL: Tsetlin, Displace.** Below 6% on 75–103 shots. Candidates to - drop or re-tune, but the N is small. +- **BELOW OVERALL — but KEEP: Tsetlin, Displace.** Both sit below the 6.95% + overall on small N (75–103 shots), which an earlier version of this report + read as an implied recommendation to drop. That was **tested and refuted**: + 15 **paired** runs per variant against DrussGT (identical seeds, 8 rounds, + same binary) gave baseline 3,238 shots / 6.18% (events 6.16%) / 200 dmg/run; + Tsetlin disabled 3,522 shots / 5.76% (events 5.71%) / 197 dmg/run; and + Tsetlin+Displace disabled 3,478 shots / 5.46% (events 5.37%) / 183 dmg/run. + Paired permutation tests: −0.34 pp (p = 0.57) and −0.70 pp (p = 0.21); the + per-run distributions completely overlap, and a Crazy (non-surfer) control + showed no separation either. Removing the measured-worst real performers is + therefore **neutral-to-slightly-negative** on both hit rate and damage. With + sd ≈ 1.8 pp a definitive claim would need far more runs, so **keep the full + rack** — being below overall does not justify removal. **[MEASURED]**. - **HeadOn MUST STAY** despite being lowest (5.2%). It is the floor fallback: when the field collapses the selector returns gun 0. Disabling the floor measurably hurt — 5.08% / 175 dmg vs 6.95% / 251 dmg (config `floor00`, @@ -396,9 +451,12 @@ overlaps base, and the nominal "winners" are ≤0.6 SE apart on far fewer shots. \* compare.py counts 14 events files, but run 14 has no fire events; the 13 runs with data carry all 3,612 shots. -**No ranking rule fixed the anti-correlation.** The best Spearman in the table -(win50, +0.588) is on 2 runs / 559 shots. The shipped config's −0.374 over 13 -runs is the most reliable estimate. **[MEASURED]**. +**No ranking rule produced a stable, useful correlation.** The best Spearman in +the table (win50, +0.588) is on 2 runs / 559 shots; none of the 16 tested rules +recovered a sign-stable signal. The shipped config's −0.374 is the largest +single-run estimate but is contradicted in sign by the +0.335 over 15 runs, so +a single Spearman value on one run set is not a reliable estimate. +**[MEASURED]** + **[INFERRED]**. ### 3.1 Offline range: which gun wins which trajectory family @@ -480,8 +538,8 @@ per-trajectory sanity, and poor as a selector signal. **[MEASURED]** + | DecayGF | recency-weighted GF | 6.4 | 49 | MARGINAL | | WallBounce | wall-reflection model | 6.2 | 56 | MARGINAL — offline favourite, real underperformer | | StopShot | deceleration/stop point | 6.1 | 48 | MARGINAL | -| Tsetlin | Tsetlin-Machine correction | 5.8 | 47 | BELOW — learns, not yet competitive | -| Displace | displacement vector | 5.3 | 50 | BELOW | +| Tsetlin | Tsetlin-Machine correction | 5.8 | 47 | **KEEP — below overall; pruning tested neutral-to-negative (§3)** | +| Displace | displacement vector | 5.3 | 50 | **KEEP — below overall; pruning tested neutral-to-negative (§3)** | | HeadOn | aim at current position | 5.2 | 35 | **KEEP — mandatory floor fallback** | | TMSelect | TM mixture-of-experts gate | — | — | **DISABLED (`EnableTmSelector = false`)** — see §6.4 | @@ -491,8 +549,11 @@ per-trajectory sanity, and poor as a selector signal. **[MEASURED]** + The KEEP/MARGINAL/BELOW split is a statement about a **single adversary (a wave surfer)**, judged on the shipped config, on a few hundred real shots per -gun. It is a starting point, not a final ranking. The concrete caveats are in -§7. +gun. It is a starting point, not a final ranking. In particular, **BELOW does +not mean "drop"**: pruning the measured-worst real performers (Tsetlin, then +Tsetlin+Displace) was tested in 15 paired runs each and was +neutral-to-slightly-negative on both hit rate and damage (§3), so the verdict +is **keep the full rack**. The concrete caveats are in §7. --- @@ -646,6 +707,14 @@ aggregated enemies in nondeterministic hash order; `stop_shot` had an unreachable deceleration branch and several guns had tick-only caches that made all four power bins return bin 0's lead (`e536900`). **[MEASURED]**. +### 6.7 The selector's tie-break was not actually random + +`randomize()` was reached only **incidentally**, through the Tsetlin gun's +constructor, so ties resolved **identically across process restarts** — the +"random" tie-break was effectively deterministic. Now fixed with an explicit +startup seed plus a `GUN_SELECTOR_SEED` override. Evidence: unseeded runs vary +across processes, seeded runs are identical. **[MEASURED]**. + --- ## 7. Known caveats and open problems @@ -670,12 +739,17 @@ Stated without hedging. bullets). Absolute offline hit rates are inflated by an unknown amount; only relative comparisons are safe. -4. **The virtual metric is anti-correlated with real hit rate and no tested - ranking rule fixed it.** Shipped config Spearman ≈ **−0.374** over 13 runs - (the sign flips to +0.52 on the smaller 5-run set, so it is unstable). 16 - candidate ranking rules all overlapped the shipped config. The selector's - value lives in its **floor/tie hedging** (5.08% without the floor vs 6.95% - with it), not in its ranking. **[MEASURED]**. +4. **The virtual metric is a poor ranker, not a reliable inverse.** The + correlation with real hit rate is weak and **sign-unstable** across run + sets: −0.374 over the 13-run base, +0.335 over a 15-run paired baseline + (different aggregation), +0.522 on the 5-run `relative+path` set, and + −0.371 / −0.073 / −0.037 on the other 5-run configs. Two opposite signs on + large samples mean it is near zero on average, not reliably anti-correlated; + an earlier version of this report overstated it as an inversion and a later + measurement refuted that. 16 candidate ranking rules all overlapped the + shipped config, so none produced a stable, useful correlation. The + selector's value lives in its **floor/tie hedging** (5.08% without the floor + vs 6.95% with it), not in its ranking. **[MEASURED]** + **[INFERRED]**. 5. **The TM classifier gun did not earn its slot.** It cost real performance (7.47% → 5.59%, 133 dmg) despite showing interpretable energy structure in @@ -692,13 +766,17 @@ Stated without hedging. evaluation metric — which is exactly what the A/B does. **[MEASURED]** (commit `dea4dcb`). -7. **The offline==online acceptance test is flaky** (§1.1): typically 11/12 on - unmodified HEAD, with the mismatching gun varying run to run. The - equivalence claim is strong-but-not-exact until the boundary race is fixed. +7. **The offline==online acceptance test is fixed and stable** (§1.1): the old + 11/12 flakiness was a real replay bug (the replay spawned disabled gun 13, + and the shared order-sensitive ring then permuted every other gun's + resolution order), now fixed by mirroring the live rack. 5/5 consecutive runs + give a byte-identical 12/12 with the death boundary included. -8. **The selector's random tie-break is not randomised in the live bot.** The - shipped bot never calls `randomize()`, so the "random" sequence is fixed - across process restarts (a side finding of `2c94dc2`, not fixed). +8. **The selector's tie-break is now explicitly seeded** (§6.7). It had been + effectively non-random — `randomize()` was reached only incidentally through + the Tsetlin gun's constructor — so ties resolved identically across process + restarts. Fixed with an explicit startup seed plus a `GUN_SELECTOR_SEED` + override; seeded runs are reproducible, unseeded runs vary. 9. **The firing gate is not the bottleneck.** The shipped range-aware gate does not beat a fixed 2.0° gate on hit rate (55.8% vs 57.9%, ~1.5 σ), though it @@ -729,7 +807,7 @@ nim c -r common_libs/tests/run_range.nim --timing GUN_SELECTOR_MODE=relative nim c -d:release -r \ common_libs/tests/analyze_selector.nim tools/fixtures/drussgt_vs_spinbot.jsonl -# Offline == online acceptance (currently flaky): +# Offline == online acceptance (fixed; stable exact 12/12 — see §1.1): nim c -r common_libs/tests/acceptance_offline_vs_online.nim # Tsetlin gun clause sparsity / divergence: diff --git a/docs/gun_rack_summary.md b/docs/gun_rack_summary.md index 74cfd17..bf6608c 100644 --- a/docs/gun_rack_summary.md +++ b/docs/gun_rack_summary.md @@ -8,13 +8,19 @@ points per run. **Boss:** the unmodified DrussGT jar through `tools/robocode_shim/`. **Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate, -251 dmg/run.** +251 dmg/run.** The **same binary** on a 15-run set measures **6.18%** (events +6.16%), so always read a rate with its run count. -**The one thing to know:** *virtual hit rate is not a proxy for real hit rate.* -For the shipped config Spearman(virtual rank, real rank) = **−0.374** — -anti-correlated. 16 candidate ranking rules all failed to beat the shipped -config; what works is the selector's floor/tie hedging (removing the floor: -5.08% / 175 dmg vs 6.95% / 251 dmg). Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md). +**The one thing to know:** *virtual hit rate is a poor ranker, not a reliable +proxy for real hit rate.* Its Spearman correlation with real hit rate is weak +and **sign-unstable** across run sets — **−0.374** over 13 runs, but **+0.335** +over a 15-run paired baseline (same 13 guns, different but equally defensible +aggregation) and +0.522 on the 5-run `relative+path` set — i.e. near zero on +average, **not** reliably anti-correlated. An earlier version of this report +overstated it as an inversion; the later 15-run measurement refuted that. 16 +candidate ranking rules all failed to beat the shipped config; what works is the +selector's floor/tie hedging (removing the floor: 5.08% / 175 dmg vs 6.95% / +251 dmg). Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md). ## Verdict table @@ -30,29 +36,45 @@ config; what works is the selector's floor/tie hedging (removing the floor: | MARGINAL | DecayGF | 6.4 | 9.2 | | MARGINAL | WallBounce | 6.2 | 12.9 | | MARGINAL | StopShot | 6.1 | 12.6 | -| BELOW | Tsetlin | 5.8 | 12.9 | -| BELOW | Displace | 5.3 | 12.3 | +| **KEEP** (below overall; pruning no help) | Tsetlin | 5.8 | 12.9 | +| **KEEP** (below overall; pruning no help) | Displace | 5.3 | 12.3 | | **FLOOR — KEEP** | HeadOn | 5.2 | 8.6 | | DISABLED | TMSelect | — | — | HeadOn is lowest but **must stay**: it is the floor fallback, and disabling the floor measurably hurt (5.08% / 175 dmg). +**Keep the full rack.** Tsetlin and Displace are below overall, but removing +them was **tested**: 15 paired runs per variant gave 6.18% → 5.76% (Tsetlin +disabled) and 5.46% (Tsetlin+Displace disabled), with fully overlapping per-run +distributions and paired permutation p = 0.57 / 0.21. Being below overall does +**not** justify removal. + ## Top actions -1. **Stop trusting the virtual metric as a ranker.** It is anti-correlated with - real hit rate and unstable across run sets. The selector survives on its - floor/tie hedge, not on its ordering. -2. **Fix the offline==online acceptance race.** It is currently flaky - (typically 11/12), so offline numbers are strong-but-not-exact. +1. **Stop trusting the virtual metric as a ranker.** It is a poor ranker — + weak, sign-unstable across run sets (−0.374 over 13 runs vs +0.335 over 15), + and near zero on average, **not** reliably anti-correlated. The selector + survives on its floor/tie hedge, not on its ordering. +2. **Acceptance test is fixed, not flaky.** The old 11/12 was a real replay + bug (the replay spawned disabled gun 13, and the shared ring is + order-sensitive, so it permuted every other gun's resolution order). It now + mirrors the live rack: 5/5 runs byte-identical 12/12 with the death boundary + included. 3. **Get a second adversary.** Every per-gun verdict rests on one wave surfer; the KEEP/BELOW boundaries are matchup-specific and per-gun N is small (47–898 shots). -4. **Decide Tsetlin and Displace.** Both are below overall on small N; Tsetlin - now learns (clauses 714→13.8 literals) but is not competitive — tune the - regression head or drop. +4. **Keep Tsetlin and Displace (pruning tested).** Both are below overall on + small N, but disabling Tsetlin (15 paired runs: 6.18% → 5.76%) and + Tsetlin+Displace (→ 5.46%) was neutral-to-slightly-negative; keep the full + rack. Tsetlin now learns (clauses 714→13.8 literals) but is not competitive — + the regression head is a follow-up, not grounds for removal. 5. **Do not re-enable TMSelect** until the gate-margin/label problem is fixed; it cost 7.47% → 5.59% despite showing real energy structure in its clauses. 6. **Real-hit-rate-driven selection is not viable yet** — unselected guns get near-zero shots, so it needs forced exploration + shrinkage + thousands of shots per gun. +7. **Selector tie-break is now seeded.** It had been effectively deterministic + (`randomize()` reached only incidentally via the Tsetlin constructor); now an + explicit startup seed plus a `GUN_SELECTOR_SEED` override makes seeded runs + reproducible and unseeded runs vary.