docs: correct the overstated 'inverted metric' claim; record tonight's fixes
Three corrections, all prompted by later measurements: 1. The virtual-vs-real rank correlation is NOT robustly negative. Six independent Spearman measurements now exist (-0.374, +0.335, +0.522, -0.371, -0.073, -0.037) and the sign flips on large samples, so it is near zero on average. The honest headline is that virtual hit rate is a POOR RANKER, not an inverted one. The report said 'not weak - it is inverted' in six places; it now says so in none. The practical conclusion (do not trust it for ranking) is unchanged; the mechanism claimed was wrong. 2. The offline==online acceptance is FIXED, not flaky. Root cause was that the replay spawned gun 13 (TMSelect) while the live rack has it disabled, and the shared VirtualTracker ring is ORDER-SENSITIVE, so gun 13's extra 4 bullets/tick permuted the per-tick resolution order for every other gun and shifted the learning guns' observations. After closing gun 13's ready gate offline the live and offline KNN traces are byte-identical (904/904 lines, empty diff). 5/5 consecutive runs now report 12/12 exact with the death boundary included. Recorded with the lesson: a flaky proof was hiding a real bug. Also records the general A/B confound - disabling a gun removes its 4 spawns/tick from the shared ring, perturbing resolution order for the rest. 3. Pruning was tested and does NOT help, so the verdict for Tsetlin and Displace changes from an implied drop to BELOW OVERALL - KEEP. 15 paired runs: baseline 6.18%, Tsetlin-off 5.76%, Tsetlin+Displace-off 5.46%; paired permutation p=0.57 and p=0.21; distributions completely overlap; a non-surfer control showed no separation. Being below average does not justify removal. Also records the tie-break randomness fix, and quotes run counts with every rate (6.95% over 13 runs vs 6.18% over 15 runs, same binary) rather than presenting a single figure as definitive.
This commit is contained in:
+126
-48
@@ -9,6 +9,13 @@ predated per-gun real-hit attribution, the offline gun range, and the live
|
||||
DrussGT boss. No number from the old version is retained unless it was
|
||||
re-measured below.
|
||||
|
||||
**Corrected after later runs (same day).** Three claims in the first version of
|
||||
this rewrite were revised by follow-up measurements: the virtual-vs-real
|
||||
correlation is a poor, sign-unstable ranker rather than an "inversion" (§2);
|
||||
the offline==online acceptance test is fixed, not flaky (§1.1); and pruning the
|
||||
below-overall guns was tested and does not help (§3). Each correction is
|
||||
restated plainly at the point of the old claim.
|
||||
|
||||
Evidence tags used throughout:
|
||||
|
||||
- **[MEASURED]** — I read it from a recorded artifact (commit message, `/tmp`
|
||||
@@ -55,14 +62,26 @@ header comment; commit `974528d`).
|
||||
if the target dies during that `go()`, the final tick's spawn+resolution is
|
||||
skipped. **[MEASURED]** (commit `974528d`; a passing run is preserved in
|
||||
`/tmp/ab_logs3/test_acceptance.log`: 12/12, 534-tick round).
|
||||
- **Honest caveat — the acceptance test is currently FLAKY.** On the
|
||||
*unmodified HEAD* source it normally reaches only **11/12**, e.g. KNN 81
|
||||
online vs 71 offline, and the mismatching gun moves between runs (KNN, then
|
||||
WallBounce). It is a live/offline boundary race, pre-existing, and not caused
|
||||
by the selector work (the replay never calls the selector). Treat
|
||||
"offline == online" as **strong but not exact until the race is fixed**.
|
||||
**[MEASURED]** (commit `2c94dc2`). The 12/12 runs above are real; they were
|
||||
lucky runs.
|
||||
- **Acceptance test: the equivalence is now proven and stable.** An earlier
|
||||
version of this report recorded the test as flaky — typically **11/12** on
|
||||
unmodified HEAD, with the mismatching gun moving between runs (KNN, then
|
||||
WallBounce) — and guessed it was a live/offline boundary race. That guess was
|
||||
**wrong**. The root cause was a **real replay bug**: the offline replay
|
||||
spawned gun 13 (TMSelect) while the live rack has `EnableTmSelector = false`
|
||||
and never does. The shared `VirtualTracker` ring is **order-sensitive**, so
|
||||
gun 13's extra 4 bullets/tick permuted the per-tick **resolution order** of
|
||||
every other gun, shifting the learning guns' observations. Closing gun 13's
|
||||
ready gate offline made the live and offline KNN traces **byte-identical**
|
||||
(904/904 lines, empty diff). The fix mirrors the live rack in the replay — no
|
||||
tick exclusion, no tolerance loosening. Stability: **5/5 consecutive runs
|
||||
report 12/12 exact, each with the death boundary included.** So
|
||||
"offline == online" is exact on these runs. A flaky proof had hidden a real
|
||||
bug. **[MEASURED]**.
|
||||
- **General lesson for rack A/B.** Because the ring is order-sensitive, any
|
||||
rack A/B that disables a gun also removes that gun's **4 spawns/tick** from
|
||||
the shared ring, which perturbs the resolution order — and therefore the
|
||||
learning observations — of every other gun. That is a confound to record for
|
||||
anyone repeating these experiments. **[INFERRED]**.
|
||||
|
||||
### 1.2 The fixture sets and what each is good for
|
||||
|
||||
@@ -187,7 +206,7 @@ ground truth for config decisions. **[MEASURED]** (re-aggregated from
|
||||
|
||||
---
|
||||
|
||||
## 2. The metric lesson: virtual hit rate is NOT a proxy for real hit rate
|
||||
## 2. The metric lesson: virtual hit rate is a poor ranker, not a proxy for real hit rate
|
||||
|
||||
This is the most important conceptual result of the night and it invalidates a
|
||||
naive reading of every offline table in this report.
|
||||
@@ -201,8 +220,13 @@ hit rate is
|
||||
Spearman(virtual rank, real rank) = -0.374 (n = 13 guns, all with ≥10 real shots)
|
||||
```
|
||||
|
||||
It is not weak — it is **inverted**. The guns with the highest virtual rates
|
||||
have among the lowest real rates, and vice versa:
|
||||
The correlation is **weak and sign-unstable** — the honest headline is that
|
||||
virtual hit rate is a **poor ranker**, not a reliable inverse. The table below
|
||||
is what a poor ranker looks like: on this run set the ordering it produces
|
||||
tracks the opposite of the real ordering, but that does not hold on other run
|
||||
sets (see the full measurement set after the table). An earlier version of this
|
||||
report stated this as a clean inversion; a later 15-run measurement refuted
|
||||
that.
|
||||
|
||||
| Gun | Selected (ticks) | Real hits/shots | Real % | Virtual % |
|
||||
|---|---:|---:|---:|---:|
|
||||
@@ -222,8 +246,10 @@ have among the lowest real rates, and vice versa:
|
||||
|
||||
Read the top and bottom: **Tsetlin, WallBounce and StopShot have the highest
|
||||
virtual rates (12.6–12.9%) and near-bottom real rates (5.8–6.2%); Linear and
|
||||
KNN sit at 10.2% / 7.5% virtual but 10.7% / 9.0% real.** The ranking the
|
||||
virtual metric produces is not merely uninformative, it points the wrong way.
|
||||
KNN sit at 10.2% / 7.5% virtual but 10.7% / 9.0% real.** On this run set the
|
||||
virtual ordering inverts the real one — but because the sign flips on other
|
||||
run sets (next paragraph), the safe reading is that the virtual ranking is
|
||||
**uninformative about the real ranking**, not that it is reliably inverted.
|
||||
|
||||
**Why this matters for selection.** What has kept the rack alive is the
|
||||
selector's **floor/tie hedging**, not its ranking: removing the floor
|
||||
@@ -232,14 +258,28 @@ selector's **floor/tie hedging**, not its ranking: removing the floor
|
||||
(`/tmp/compare.py`). So the selector is useful because it refuses to commit to
|
||||
a bad field, not because its virtual-rate ordering is good.
|
||||
|
||||
**The virtual metric appears anti-correlated no matter which config you pick.**
|
||||
Measured Spearman per config (5-run 12-round A/B, `/tmp/ab_logs3/FINAL_AB.txt`):
|
||||
`absolute+point` −0.371, `absolute+path` −0.073, `relative+point` −0.037,
|
||||
`relative+path` +0.522. But on the large 13-run base set the shipped config
|
||||
(`relative+path`) is **−0.374**. The sign **flips between run sets**, which is
|
||||
itself the finding: the correlation is unstable, so no ranking rule built on
|
||||
it can be trusted. **[MEASURED]** + **[INFERRED]** (the flip is measured; the
|
||||
conclusion is reasoning).
|
||||
**The headline: across run sets the correlation is sign-unstable, so it is near
|
||||
zero on average — not robustly negative.** The full set of independent
|
||||
Spearman measurements (13-run source `/tmp/agg2.py base`; 5-run configs
|
||||
`/tmp/ab_logs3/FINAL_AB.txt`; 15-run paired baseline §3) is:
|
||||
|
||||
| Run set / aggregation | Spearman |
|
||||
|---|---:|
|
||||
| 13-run base, shipped `relative+path` | **−0.374** |
|
||||
| 15-run paired baseline (different but equally defensible aggregation) | **+0.335** |
|
||||
| 5-run 12-round A/B, `relative+path` | +0.522 |
|
||||
| 5-run A/B, `absolute+point` | −0.371 |
|
||||
| 5-run A/B, `absolute+path` | −0.073 |
|
||||
| 5-run A/B, `relative+point` | −0.037 |
|
||||
|
||||
Two **opposite signs on large samples** (−0.374 over 13 runs, +0.335 over 15
|
||||
runs, all over the same 13 guns) mean virtual hit rate is **not** a reliable
|
||||
inverse of real hit rate. It is a **poor ranker**: weak correlation, sign
|
||||
flipping between run sets, near zero on average. The practical conclusion is
|
||||
unchanged — do not build a ranking rule on it — but an **earlier version of
|
||||
this report overstated the mechanism as an inversion**; the later 15-run
|
||||
measurement refuted that. **[MEASURED]** + **[INFERRED]** (the six numbers are
|
||||
measured; "poor ranker / near zero on average" is the reasoning).
|
||||
|
||||
### 2.1 The metric A/B: point vs path (this one is real, and it is selection)
|
||||
|
||||
@@ -327,8 +367,11 @@ The final per-gun table, shipped config, **13 runs vs the live DrussGT boss,
|
||||
3,612 server-side shots, 6.95% overall** (server sidecar; per-run rates 6.77 /
|
||||
5.90 / 6.57 / 6.34 / 6.10 / 7.59 / 2.90 / 7.95 / 7.49 / 9.18 / 8.44 / 7.48 /
|
||||
6.42%; 251 dmg/run). Per-gun rows are the bot-side attribution over the same
|
||||
runs. **[MEASURED]** (commit `2c94dc2`; `/tmp/gun_stats_base_r*.jsonl`,
|
||||
re-aggregated with `/tmp/agg2.py base`; `/tmp/events_base_r*.json`).
|
||||
runs. The **same binary** on a different 15-run set gives **6.18%** (events
|
||||
6.16%, 200 dmg/run, §3 pruning baseline), so every rate here is quoted with its
|
||||
run count — a single figure is not definitive. **[MEASURED]** (commit `2c94dc2`;
|
||||
`/tmp/gun_stats_base_r*.jsonl`, re-aggregated with `/tmp/agg2.py base`;
|
||||
`/tmp/events_base_r*.json`).
|
||||
|
||||
| Verdict | Gun | Real hits/shots | Real % | Virtual % | Selected |
|
||||
|---|---|---:|---:|---:|---:|
|
||||
@@ -342,8 +385,8 @@ re-aggregated with `/tmp/agg2.py base`; `/tmp/events_base_r*.json`).
|
||||
| MARGINAL | DecayGF | 3/47 | 6.4 | 9.2 | 1,360 |
|
||||
| MARGINAL | WallBounce | 18/288 | 6.2 | 12.9 | 7,618 |
|
||||
| MARGINAL | StopShot | 8/132 | 6.1 | 12.6 | 3,348 |
|
||||
| BELOW | Tsetlin | 6/103 | 5.8 | 12.9 | 3,284 |
|
||||
| BELOW | Displace | 4/75 | 5.3 | 12.3 | 2,664 |
|
||||
| BELOW — KEEP | Tsetlin | 6/103 | 5.8 | 12.9 | 3,284 |
|
||||
| BELOW — KEEP | Displace | 4/75 | 5.3 | 12.3 | 2,664 |
|
||||
| **FLOOR — STAYS** | HeadOn | 47/898 | 5.2 | 8.6 | 16,975 |
|
||||
|
||||
**Verdicts.**
|
||||
@@ -352,12 +395,24 @@ re-aggregated with `/tmp/agg2.py base`; `/tmp/events_base_r*.json`).
|
||||
above the 6.95% overall, yet their virtual rates are mid-pack to low: the
|
||||
metric's three favourites (Tsetlin 12.9%, WallBounce 12.9%, StopShot 12.6%)
|
||||
are near the *bottom* of the real ranking, while the real leader (Linear) sits
|
||||
at 10.2% virtual. Further evidence the virtual ranking is inverted.
|
||||
at 10.2% virtual. Further evidence the virtual ranking is uninformative about
|
||||
the real ranking (and, on this run set, roughly its opposite).
|
||||
- **MARGINAL: GuessFactor, DecayGF, WallBounce, StopShot.** Within ~1 pp of
|
||||
overall on small N (47–288 shots). They are not obviously worth deleting, but
|
||||
they have not earned a larger share.
|
||||
- **BELOW OVERALL: Tsetlin, Displace.** Below 6% on 75–103 shots. Candidates to
|
||||
drop or re-tune, but the N is small.
|
||||
- **BELOW OVERALL — but KEEP: Tsetlin, Displace.** Both sit below the 6.95%
|
||||
overall on small N (75–103 shots), which an earlier version of this report
|
||||
read as an implied recommendation to drop. That was **tested and refuted**:
|
||||
15 **paired** runs per variant against DrussGT (identical seeds, 8 rounds,
|
||||
same binary) gave baseline 3,238 shots / 6.18% (events 6.16%) / 200 dmg/run;
|
||||
Tsetlin disabled 3,522 shots / 5.76% (events 5.71%) / 197 dmg/run; and
|
||||
Tsetlin+Displace disabled 3,478 shots / 5.46% (events 5.37%) / 183 dmg/run.
|
||||
Paired permutation tests: −0.34 pp (p = 0.57) and −0.70 pp (p = 0.21); the
|
||||
per-run distributions completely overlap, and a Crazy (non-surfer) control
|
||||
showed no separation either. Removing the measured-worst real performers is
|
||||
therefore **neutral-to-slightly-negative** on both hit rate and damage. With
|
||||
sd ≈ 1.8 pp a definitive claim would need far more runs, so **keep the full
|
||||
rack** — being below overall does not justify removal. **[MEASURED]**.
|
||||
- **HeadOn MUST STAY** despite being lowest (5.2%). It is the floor fallback:
|
||||
when the field collapses the selector returns gun 0. Disabling the floor
|
||||
measurably hurt — 5.08% / 175 dmg vs 6.95% / 251 dmg (config `floor00`,
|
||||
@@ -396,9 +451,12 @@ overlaps base, and the nominal "winners" are ≤0.6 SE apart on far fewer shots.
|
||||
\* compare.py counts 14 events files, but run 14 has no fire events; the 13
|
||||
runs with data carry all 3,612 shots.
|
||||
|
||||
**No ranking rule fixed the anti-correlation.** The best Spearman in the table
|
||||
(win50, +0.588) is on 2 runs / 559 shots. The shipped config's −0.374 over 13
|
||||
runs is the most reliable estimate. **[MEASURED]**.
|
||||
**No ranking rule produced a stable, useful correlation.** The best Spearman in
|
||||
the table (win50, +0.588) is on 2 runs / 559 shots; none of the 16 tested rules
|
||||
recovered a sign-stable signal. The shipped config's −0.374 is the largest
|
||||
single-run estimate but is contradicted in sign by the +0.335 over 15 runs, so
|
||||
a single Spearman value on one run set is not a reliable estimate.
|
||||
**[MEASURED]** + **[INFERRED]**.
|
||||
|
||||
### 3.1 Offline range: which gun wins which trajectory family
|
||||
|
||||
@@ -480,8 +538,8 @@ per-trajectory sanity, and poor as a selector signal. **[MEASURED]** +
|
||||
| DecayGF | recency-weighted GF | 6.4 | 49 | MARGINAL |
|
||||
| WallBounce | wall-reflection model | 6.2 | 56 | MARGINAL — offline favourite, real underperformer |
|
||||
| StopShot | deceleration/stop point | 6.1 | 48 | MARGINAL |
|
||||
| Tsetlin | Tsetlin-Machine correction | 5.8 | 47 | BELOW — learns, not yet competitive |
|
||||
| Displace | displacement vector | 5.3 | 50 | BELOW |
|
||||
| Tsetlin | Tsetlin-Machine correction | 5.8 | 47 | **KEEP — below overall; pruning tested neutral-to-negative (§3)** |
|
||||
| Displace | displacement vector | 5.3 | 50 | **KEEP — below overall; pruning tested neutral-to-negative (§3)** |
|
||||
| HeadOn | aim at current position | 5.2 | 35 | **KEEP — mandatory floor fallback** |
|
||||
| TMSelect | TM mixture-of-experts gate | — | — | **DISABLED (`EnableTmSelector = false`)** — see §6.4 |
|
||||
|
||||
@@ -491,8 +549,11 @@ per-trajectory sanity, and poor as a selector signal. **[MEASURED]** +
|
||||
|
||||
The KEEP/MARGINAL/BELOW split is a statement about a **single adversary (a
|
||||
wave surfer)**, judged on the shipped config, on a few hundred real shots per
|
||||
gun. It is a starting point, not a final ranking. The concrete caveats are in
|
||||
§7.
|
||||
gun. It is a starting point, not a final ranking. In particular, **BELOW does
|
||||
not mean "drop"**: pruning the measured-worst real performers (Tsetlin, then
|
||||
Tsetlin+Displace) was tested in 15 paired runs each and was
|
||||
neutral-to-slightly-negative on both hit rate and damage (§3), so the verdict
|
||||
is **keep the full rack**. The concrete caveats are in §7.
|
||||
|
||||
---
|
||||
|
||||
@@ -646,6 +707,14 @@ aggregated enemies in nondeterministic hash order; `stop_shot` had an
|
||||
unreachable deceleration branch and several guns had tick-only caches that made
|
||||
all four power bins return bin 0's lead (`e536900`). **[MEASURED]**.
|
||||
|
||||
### 6.7 The selector's tie-break was not actually random
|
||||
|
||||
`randomize()` was reached only **incidentally**, through the Tsetlin gun's
|
||||
constructor, so ties resolved **identically across process restarts** — the
|
||||
"random" tie-break was effectively deterministic. Now fixed with an explicit
|
||||
startup seed plus a `GUN_SELECTOR_SEED` override. Evidence: unseeded runs vary
|
||||
across processes, seeded runs are identical. **[MEASURED]**.
|
||||
|
||||
---
|
||||
|
||||
## 7. Known caveats and open problems
|
||||
@@ -670,12 +739,17 @@ Stated without hedging.
|
||||
bullets). Absolute offline hit rates are inflated by an unknown amount; only
|
||||
relative comparisons are safe.
|
||||
|
||||
4. **The virtual metric is anti-correlated with real hit rate and no tested
|
||||
ranking rule fixed it.** Shipped config Spearman ≈ **−0.374** over 13 runs
|
||||
(the sign flips to +0.52 on the smaller 5-run set, so it is unstable). 16
|
||||
candidate ranking rules all overlapped the shipped config. The selector's
|
||||
value lives in its **floor/tie hedging** (5.08% without the floor vs 6.95%
|
||||
with it), not in its ranking. **[MEASURED]**.
|
||||
4. **The virtual metric is a poor ranker, not a reliable inverse.** The
|
||||
correlation with real hit rate is weak and **sign-unstable** across run
|
||||
sets: −0.374 over the 13-run base, +0.335 over a 15-run paired baseline
|
||||
(different aggregation), +0.522 on the 5-run `relative+path` set, and
|
||||
−0.371 / −0.073 / −0.037 on the other 5-run configs. Two opposite signs on
|
||||
large samples mean it is near zero on average, not reliably anti-correlated;
|
||||
an earlier version of this report overstated it as an inversion and a later
|
||||
measurement refuted that. 16 candidate ranking rules all overlapped the
|
||||
shipped config, so none produced a stable, useful correlation. The
|
||||
selector's value lives in its **floor/tie hedging** (5.08% without the floor
|
||||
vs 6.95% with it), not in its ranking. **[MEASURED]** + **[INFERRED]**.
|
||||
|
||||
5. **The TM classifier gun did not earn its slot.** It cost real performance
|
||||
(7.47% → 5.59%, 133 dmg) despite showing interpretable energy structure in
|
||||
@@ -692,13 +766,17 @@ Stated without hedging.
|
||||
evaluation metric — which is exactly what the A/B does. **[MEASURED]**
|
||||
(commit `dea4dcb`).
|
||||
|
||||
7. **The offline==online acceptance test is flaky** (§1.1): typically 11/12 on
|
||||
unmodified HEAD, with the mismatching gun varying run to run. The
|
||||
equivalence claim is strong-but-not-exact until the boundary race is fixed.
|
||||
7. **The offline==online acceptance test is fixed and stable** (§1.1): the old
|
||||
11/12 flakiness was a real replay bug (the replay spawned disabled gun 13,
|
||||
and the shared order-sensitive ring then permuted every other gun's
|
||||
resolution order), now fixed by mirroring the live rack. 5/5 consecutive runs
|
||||
give a byte-identical 12/12 with the death boundary included.
|
||||
|
||||
8. **The selector's random tie-break is not randomised in the live bot.** The
|
||||
shipped bot never calls `randomize()`, so the "random" sequence is fixed
|
||||
across process restarts (a side finding of `2c94dc2`, not fixed).
|
||||
8. **The selector's tie-break is now explicitly seeded** (§6.7). It had been
|
||||
effectively non-random — `randomize()` was reached only incidentally through
|
||||
the Tsetlin gun's constructor — so ties resolved identically across process
|
||||
restarts. Fixed with an explicit startup seed plus a `GUN_SELECTOR_SEED`
|
||||
override; seeded runs are reproducible, unseeded runs vary.
|
||||
|
||||
9. **The firing gate is not the bottleneck.** The shipped range-aware gate does
|
||||
not beat a fixed 2.0° gate on hit rate (55.8% vs 57.9%, ~1.5 σ), though it
|
||||
@@ -729,7 +807,7 @@ nim c -r common_libs/tests/run_range.nim --timing
|
||||
GUN_SELECTOR_MODE=relative nim c -d:release -r \
|
||||
common_libs/tests/analyze_selector.nim tools/fixtures/drussgt_vs_spinbot.jsonl
|
||||
|
||||
# Offline == online acceptance (currently flaky):
|
||||
# Offline == online acceptance (fixed; stable exact 12/12 — see §1.1):
|
||||
nim c -r common_libs/tests/acceptance_offline_vs_online.nim
|
||||
|
||||
# Tsetlin gun clause sparsity / divergence:
|
||||
|
||||
+38
-16
@@ -8,13 +8,19 @@ points per run.
|
||||
|
||||
**Boss:** the unmodified DrussGT jar through `tools/robocode_shim/`.
|
||||
**Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate,
|
||||
251 dmg/run.**
|
||||
251 dmg/run.** The **same binary** on a 15-run set measures **6.18%** (events
|
||||
6.16%), so always read a rate with its run count.
|
||||
|
||||
**The one thing to know:** *virtual hit rate is not a proxy for real hit rate.*
|
||||
For the shipped config Spearman(virtual rank, real rank) = **−0.374** —
|
||||
anti-correlated. 16 candidate ranking rules all failed to beat the shipped
|
||||
config; what works is the selector's floor/tie hedging (removing the floor:
|
||||
5.08% / 175 dmg vs 6.95% / 251 dmg). Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md).
|
||||
**The one thing to know:** *virtual hit rate is a poor ranker, not a reliable
|
||||
proxy for real hit rate.* Its Spearman correlation with real hit rate is weak
|
||||
and **sign-unstable** across run sets — **−0.374** over 13 runs, but **+0.335**
|
||||
over a 15-run paired baseline (same 13 guns, different but equally defensible
|
||||
aggregation) and +0.522 on the 5-run `relative+path` set — i.e. near zero on
|
||||
average, **not** reliably anti-correlated. An earlier version of this report
|
||||
overstated it as an inversion; the later 15-run measurement refuted that. 16
|
||||
candidate ranking rules all failed to beat the shipped config; what works is the
|
||||
selector's floor/tie hedging (removing the floor: 5.08% / 175 dmg vs 6.95% /
|
||||
251 dmg). Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md).
|
||||
|
||||
## Verdict table
|
||||
|
||||
@@ -30,29 +36,45 @@ config; what works is the selector's floor/tie hedging (removing the floor:
|
||||
| MARGINAL | DecayGF | 6.4 | 9.2 |
|
||||
| MARGINAL | WallBounce | 6.2 | 12.9 |
|
||||
| MARGINAL | StopShot | 6.1 | 12.6 |
|
||||
| BELOW | Tsetlin | 5.8 | 12.9 |
|
||||
| BELOW | Displace | 5.3 | 12.3 |
|
||||
| **KEEP** (below overall; pruning no help) | Tsetlin | 5.8 | 12.9 |
|
||||
| **KEEP** (below overall; pruning no help) | Displace | 5.3 | 12.3 |
|
||||
| **FLOOR — KEEP** | HeadOn | 5.2 | 8.6 |
|
||||
| DISABLED | TMSelect | — | — |
|
||||
|
||||
HeadOn is lowest but **must stay**: it is the floor fallback, and disabling the
|
||||
floor measurably hurt (5.08% / 175 dmg).
|
||||
|
||||
**Keep the full rack.** Tsetlin and Displace are below overall, but removing
|
||||
them was **tested**: 15 paired runs per variant gave 6.18% → 5.76% (Tsetlin
|
||||
disabled) and 5.46% (Tsetlin+Displace disabled), with fully overlapping per-run
|
||||
distributions and paired permutation p = 0.57 / 0.21. Being below overall does
|
||||
**not** justify removal.
|
||||
|
||||
## Top actions
|
||||
|
||||
1. **Stop trusting the virtual metric as a ranker.** It is anti-correlated with
|
||||
real hit rate and unstable across run sets. The selector survives on its
|
||||
floor/tie hedge, not on its ordering.
|
||||
2. **Fix the offline==online acceptance race.** It is currently flaky
|
||||
(typically 11/12), so offline numbers are strong-but-not-exact.
|
||||
1. **Stop trusting the virtual metric as a ranker.** It is a poor ranker —
|
||||
weak, sign-unstable across run sets (−0.374 over 13 runs vs +0.335 over 15),
|
||||
and near zero on average, **not** reliably anti-correlated. The selector
|
||||
survives on its floor/tie hedge, not on its ordering.
|
||||
2. **Acceptance test is fixed, not flaky.** The old 11/12 was a real replay
|
||||
bug (the replay spawned disabled gun 13, and the shared ring is
|
||||
order-sensitive, so it permuted every other gun's resolution order). It now
|
||||
mirrors the live rack: 5/5 runs byte-identical 12/12 with the death boundary
|
||||
included.
|
||||
3. **Get a second adversary.** Every per-gun verdict rests on one wave surfer;
|
||||
the KEEP/BELOW boundaries are matchup-specific and per-gun N is small
|
||||
(47–898 shots).
|
||||
4. **Decide Tsetlin and Displace.** Both are below overall on small N; Tsetlin
|
||||
now learns (clauses 714→13.8 literals) but is not competitive — tune the
|
||||
regression head or drop.
|
||||
4. **Keep Tsetlin and Displace (pruning tested).** Both are below overall on
|
||||
small N, but disabling Tsetlin (15 paired runs: 6.18% → 5.76%) and
|
||||
Tsetlin+Displace (→ 5.46%) was neutral-to-slightly-negative; keep the full
|
||||
rack. Tsetlin now learns (clauses 714→13.8 literals) but is not competitive —
|
||||
the regression head is a follow-up, not grounds for removal.
|
||||
5. **Do not re-enable TMSelect** until the gate-margin/label problem is fixed;
|
||||
it cost 7.47% → 5.59% despite showing real energy structure in its clauses.
|
||||
6. **Real-hit-rate-driven selection is not viable yet** — unselected guns get
|
||||
near-zero shots, so it needs forced exploration + shrinkage + thousands of
|
||||
shots per gun.
|
||||
7. **Selector tie-break is now seeded.** It had been effectively deterministic
|
||||
(`randomize()` reached only incidentally via the Tsetlin constructor); now an
|
||||
explicit startup seed plus a `GUN_SELECTOR_SEED` override makes seeded runs
|
||||
reproducible and unseeded runs vary.
|
||||
|
||||
Reference in New Issue
Block a user