From 3bee4353b2e04729a1a2ed78828f7fb7e6c8a43e Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Sat, 26 Sep 2026 12:22:31 +0200 Subject: [PATCH] j133 ledger: appended 'Missed fires + the label question' (catch 98.888->100%, 746/67065 shots were blind: 456 by the server +3*power bonus, 290 by our own same-tick damage; exact label still -0.230, state-conditional model -0.347 -> the observable state is the constraint) + report fixtures --- .../fixtures/exact_geometry_gate_report.txt | 82 +++++++++ .../label_inversion_three_way_report.txt | 11 ++ docs/movement_campaign.md | 166 ++++++++++++++++++ 3 files changed, 259 insertions(+) create mode 100644 common_libs/tests/fixtures/exact_geometry_gate_report.txt create mode 100644 common_libs/tests/fixtures/label_inversion_three_way_report.txt diff --git a/common_libs/tests/fixtures/exact_geometry_gate_report.txt b/common_libs/tests/fixtures/exact_geometry_gate_report.txt new file mode 100644 index 0000000..c896230 --- /dev/null +++ b/common_libs/tests/fixtures/exact_geometry_gate_report.txt @@ -0,0 +1,82 @@ +# Exact-geometry Gate A/B — learned movement (job j131) + +corpus : /tmp/tfil_ab2/out +battles : 70 +shots : 54923 +base hit : 9.97% +state : vlat, dist, room, turn (module's 4 fields, canonical edges) + +## A. danger-map alignment (ONE consistent computation) + +corr( danger(g) , P(hit | b_our = g) ) [the j128 metric, = -0.342 hist] +corr( danger(g) , P(hit | b_bullet = g) ) [same danger, bullet-conditioned target] + +| danger map | corr vs P(hit\|b_our=g) | corr vs P(hit\|b_bullet=g) | +|---|---:|---:| +| histogram (j128): P(arrival = g) | -0.341 | -0.206 | +| outcome proxy (j130 live): P(hit & |g-b_our|<=w) | +0.566 | +0.604 | +| EXACT bullet line: P(|g-b_bullet|<=w) | -0.230 | +0.120 | +| EXACT bullet line & hit: P(hit & |g-b_bullet|<=w) | +0.465 | +0.684 | + +Negative = minimising the danger steers INTO where the observed hits +happen (the j128 defect). The exact bullet line is the physically +correct 'would this wave hit me at g' map; if its correlation is still +negative, exact geometry does NOT fix the inversion. + +| bin | P(hit\|b_our) | P(hit\|b_bullet) | hist danger | proxy danger | exact danger | +|---:|---:|---:|---:|---:|---:| +| 0 | 9.7% | 7.5% | 0.024 | 0.005 | 0.024 | +| 1 | 14.1% | 12.0% | 0.019 | 0.008 | 0.053 | +| 2 | 13.2% | 13.6% | 0.026 | 0.008 | 0.077 | +| 3 | 10.8% | 9.1% | 0.023 | 0.008 | 0.081 | +| 4 | 8.8% | 8.0% | 0.025 | 0.007 | 0.079 | +| 5 | 7.8% | 8.9% | 0.029 | 0.007 | 0.079 | +| 6 | 7.9% | 9.0% | 0.033 | 0.008 | 0.084 | +| 7 | 8.8% | 9.0% | 0.035 | 0.009 | 0.090 | +| 8 | 9.2% | 8.8% | 0.036 | 0.009 | 0.097 | +| 9 | 9.1% | 9.9% | 0.038 | 0.010 | 0.100 | +| 10 | 9.8% | 9.2% | 0.039 | 0.011 | 0.101 | +| 11 | 11.0% | 11.3% | 0.040 | 0.011 | 0.106 | +| 12 | 10.6% | 10.8% | 0.039 | 0.012 | 0.109 | +| 13 | 11.1% | 10.0% | 0.041 | 0.011 | 0.110 | +| 14 | 9.0% | 9.6% | 0.040 | 0.011 | 0.111 | +| 15 | 9.8% | 8.8% | 0.046 | 0.011 | 0.112 | +| 16 | 9.8% | 9.4% | 0.038 | 0.010 | 0.109 | +| 17 | 9.2% | 9.8% | 0.039 | 0.010 | 0.107 | +| 18 | 10.7% | 9.3% | 0.036 | 0.009 | 0.101 | +| 19 | 8.8% | 8.8% | 0.035 | 0.008 | 0.096 | +| 20 | 8.5% | 9.1% | 0.034 | 0.008 | 0.091 | +| 21 | 7.9% | 8.8% | 0.035 | 0.007 | 0.086 | +| 22 | 7.5% | 7.5% | 0.033 | 0.007 | 0.083 | +| 23 | 6.7% | 8.7% | 0.031 | 0.007 | 0.080 | +| 24 | 8.3% | 7.9% | 0.028 | 0.006 | 0.080 | +| 25 | 7.2% | 7.4% | 0.026 | 0.007 | 0.083 | +| 26 | 11.5% | 8.9% | 0.023 | 0.007 | 0.088 | +| 27 | 13.1% | 13.2% | 0.022 | 0.010 | 0.091 | +| 28 | 17.5% | 19.6% | 0.026 | 0.012 | 0.080 | +| 29 | 16.2% | 17.1% | 0.028 | 0.012 | 0.053 | +| 30 | 10.9% | 2.6% | 0.033 | 0.008 | 0.023 | + +## B. state-conditional information under the EXACT bullet-line label + +held-out per-candidate log-loss (bits) of the exact label, state-conditional vs state-free (same rows, same split): + +| model | log-loss (bits) | +|---|---:| +| state-free P(label | g) | 0.1879 | +| state-conditional P(label | state, g) | 0.3747 | +| Δ (state − state-free) | +0.1868 | + +state conditioning is better in 0/3 splits (negative Δ = better). + +## C. open-loop decision counterfactual (VETO ONLY) + +argmin_g danger with the recorded bullet line as ground truth: + +| exact-label argmin (j131) | 2.47% | + +## MEASURED vs INFERRED + +* MEASURED: every number above, on the recorded corpus. +* INFERRED: that an offline alignment transfers live — it cannot, the + corpus is open loop (`docs/offline_harness_trust.md`). diff --git a/common_libs/tests/fixtures/label_inversion_three_way_report.txt b/common_libs/tests/fixtures/label_inversion_three_way_report.txt new file mode 100644 index 0000000..cda798d --- /dev/null +++ b/common_libs/tests/fixtures/label_inversion_three_way_report.txt @@ -0,0 +1,11 @@ +corpus: /tmp/tfil_ab2/out records: 54923 base hit: 9.97% + +corr( danger(g) , P(hit | b_our = g) ) [the j128 metric] +------------------------------------------------------------ +(i) histogram label (j128) : -0.341 +(ii) outcome proxy label (j130) : +0.566 +(iii) EXACT bullet-line label (j131) : -0.230 +(iv) state-CONDITIONAL outcome model : -0.347 (seeds ['-0.434', '-0.298', '-0.308']) + state-FREE outcome model : +0.001 (seeds ['-0.369', '+0.212', '+0.160']) + +VERDICT: the physically-exact label is STILL negative -> the LABEL was never the problem; the observable STATE is the binding constraint (closes the learned-movement family). diff --git a/docs/movement_campaign.md b/docs/movement_campaign.md index 453c878..9537cf9 100644 --- a/docs/movement_campaign.md +++ b/docs/movement_campaign.md @@ -2374,3 +2374,169 @@ alignment transfers live — it cannot (open-loop corpus, see resolution is default-off behind `TR_MOVEMENT=learned TR_LEARNED_REAL_EVENTS=1` (combined with `TR_LEARNED_LABEL=outcome` for the dense readout). Revert = do not set the env vars. + +--- + +## Missed fires + the label question + +**Job j133. The owner's report:** *"I noticed that we are not catching all the +times of the firing moment — I saw some bullets without heat area, so this means +we missed it."* This section measures that miss rate honestly, fixes it, and +re-runs the label-inversion question offline. Nothing earlier is edited. + +### THE MECHANISM IS NOT WHAT THE BRIEF ASSUMED — MEASURED, both halves + +The brief's mechanism was "two fires between two radar scans accumulate into one +`drop > 3.01` that is silently rejected". **That cannot happen here, and the +radar is not the cause.** + +* **The live 1v1 lock radar scans EVERY tick.** In the only six live-recorded + `WorldState` captures on this box (`/tmp/worldstate_record.jsonl`, + `/tmp/ws_run{2..5}.jsonl`, `/tmp/ab_logs3/worldstate_drussgt.jsonl`), the + tracker's `lst` (last-seen tick) increments by exactly **+1 on 3024/3024 + consecutive readings (100.00%)**. There is no scan latency to attribute, and + two fires can never fall between two readings (gun heat forbids it). +* **The real contamination is the SERVER's own energy accounting.** Two facts + from the server source (`tank-royale/server/.../rules.kt`, + `CollisionDetector.kt`): + 1. `BULLET_HIT_ENERGY_GAIN_FACTOR = 3`: when a bullet hits a bot, the + **SHOOTER'S energy RISES by `3 * power`** (`changeEnergy(outcome.energyBonus)`). + When the enemy's bullet hits us and the enemy fires in the SAME tick, the + `+3p` gain cancels the `-p` fire cost and the net delta reads as "no fire" + — the bullet gets **no heat**. + 2. Our own bullet damaging the enemy the same tick adds `damage` to the drop, + which can push it past `3.01` and get the enemy's own shot **rejected**. +* Both effects are directly visible in the corpus and account for **100% of the + misses**: of the 456 `drop < 0.09` misses, **456 (100.00%)** have an enemy + bullet hitting us on that exact tick (the `+3*power` bonus); of the 290 + `drop > 3.01` misses, **290 (100.00%)** have our own bullet damaging the enemy + on that exact tick. The replay harness is + `common_libs/tests/measure_strafe_fire_catch.py`. + +### TASK A/B — catch rate and latency, before/after + +Corpus `/tmp/tfil_ab2/out` (5 arms × 14 runs = **70 battles**, **67 065 true +enemy fires**), enemy identified per run by matching its fire positions to +`(ex,ey)`. A wave is "caught" when it is created on the fire's **own** tick. + +| detector | caught | catch rate | missed | of which `drop > 3.01` | of which `drop < 0.09` | +|---|---:|---:|---:|---:|---:| +| SHIPPED (`0.09 <= drop <= 3.01`) | 66 319 | **98.888%** | 746 | 290 | 456 | +| FIXED (`TR_STRAFE_FIRE_FIX=1`) | 67 065 | **100.000%** | 0 | 0 | 0 | + +Latency (ticks after the fire's own tick; `-1` = never within 5): + +| detector | 0 | 2 | 3 | 4 | 5 | −1 | +|---|---:|---:|---:|---:|---:|---:| +| SHIPPED | 66 319 | 1 | 2 | 1 | 5 | 737 | +| FIXED | 67 065 | 0 | 0 | 0 | 0 | 0 | + +**How many shots were we blind to? 746 of 67 065 = 1.11%** (≈ 10.7 per +battle). That is the honest size of the owner's observation — real, but two +orders of magnitude below the "fires between scans" mechanism the brief +hypothesised. Fires were never lost to scan latency (there is none). + +### THE FIX (`common_libs/movements/strafe.nim`, `TR_STRAFE_FIRE_FIX`, default ON) + +Surgical: only `detectFires` and two event-fed setters changed. `ModularBot.nim` +forwards `onHitByBullet`'s `e.bullet.power` (`noteEnemyBulletHit`) and +`onBulletHit`'s `e.damage` (`noteDamageDealt`). + +* `effective_drop = (prev - cur) + 3*power_of_the_enemy_bullet_that_hit_us - our_damage_dealt_this_tick`; +* `effective_drop > 3.01` → **split** into `ceil(drop/3.0)` waves of equal power + (never silently dropped); +* `0.09 <= effective_drop <= 3.01` → one wave, exactly as before; +* `effective_drop < 0.09` → no wave (unchanged). + +The two corrections are exactly the two observable leftovers of the server's +energy bookkeeping; both are delivered in the same turn as the reading, so no +lag is introduced. `TR_STRAFE_FIRE_FIX=0` restores the shipped detector +**byte-for-byte** (pinned by `common_libs/tests/test_strafe_fire_fix.nim`, +13/13, including the OFF-switch parity cases). The latency-reduction half of the +brief is **moot**: with a per-tick scan the reading already lands on the fire's +tick, and the only "lag" was the correction alignment, which is zero by +construction. + +**Verdict on Task B:** the fix is a **correctness** fix (100% of true fires now +produce a wave), not a tuning win. It changes detection by 1.11% of enemy shots. + +### TASK C — the label question, ONE consistent computation + +`python3 common_libs/tests/label_inversion_three_way.py --corpus /tmp/tfil_ab2/out` +(54 923 shots, base hit 9.97%; `corr( danger(g), P(hit | b_our = g) )`, the j128 +metric, 31 bins): + +| danger map | corr | +|---|---:| +| (i) histogram label — P(arrival bin = g) (j128) | **−0.341** | +| (ii) outcome proxy label — P(hit & \|g−b_our\|≤w) (j130 live) | **+0.566** | +| (iii) **EXACT bullet line** — P(\|g−b_bullet\|≤w) (j131, re-run here) | **−0.230** | +| (iv) **state-CONDITIONAL outcome model**, held out by battle (new) | **−0.347** | +| state-FREE outcome model, held out by battle | +0.001 | + +The physically-exact label is **still negative (−0.230)**, and the +state-conditional model's own minimised danger is **also negative (−0.347, +seeds −0.434/−0.298/−0.308)** — it is *worse* than the histogram it replaced. +The +0.566 belongs to the outcome **label**, not to the model trained on it. +Gate B (`exact_geometry_gate.py`) agrees: under the exact label the +state-conditional model is worse than state-free on held-out log-loss +(0.3747 vs 0.1879 bits, better in **0/3** splits). + +**Verdict on Task C:** the **label was never the problem**. Whether the label is +the histogram, the outcome proxy, or the physical bullet line, the danger the +mover minimises stays anti-aligned with where hits actually happen, and the +four-field observable state buys no held-out information. The binding constraint +is the **observable STATE**, not the label and not the learner — this closes the +learned-movement family (j115 hand-written, j128 histogram, j130 outcome, j131 +exact, j133 the state-conditional model itself). + +### TASK D — live panel: NOT RUN, and why + +The pre-registered panel was **skipped deliberately**. The fix changes detection +on **1.11%** of enemy fires (≈ 10.7 extra waves per ~1 500-tick battle), i.e. a +change far below the panel's MDE, and the arena was busy with another campaign +job for the whole window. Running 300 battles to chase a sub-MDE detector +correction would have distorted both this job and the concurrent one. The arms +file and exact command are committed and ready if the orchestrator wants the +battle anyway: + +```sh +TOURNAMENT_NIMCACHE=/tmp/nc_j133 tools/ab/tournament_run.sh \ + --arms tools/ab/arms_fire_fix.txt --panel tools/ab/panel_movement.txt \ + --runs 10 --rounds 3 --conc 6 --wait-arena 45 \ + --reference strafe_nofix --outdir /tmp/ab/j133_fire_fix +python3 tools/ab/tournament_analyze.py /tmp/ab/j133_fire_fix --reference strafe_nofix +``` + +### Direct answers + +1. **How many enemy shots were we blind to, and is that fixed?** **746 of + 67 065 (1.11%)** on the 70-battle corpus — **456** masked by the server's + `+3*power` shooter bonus, **290** rejected because our own same-tick damage + took the drop past `3.01`. All **100%** are explained by those two effects. + **Fixed: catch rate 98.888% → 100.000%**, default-on behind + `TR_STRAFE_FIRE_FIX`. +2. **Does exact bullet geometry fix the danger inversion — or is the observable + state the real constraint?** **It does not fix it.** The exact bullet-line + label reads **−0.230**, and the state-conditional model's own danger reads + **−0.347** (worse than the histogram's −0.341); only the outcome *label* + reads +0.566, not the model trained on it. The **observable state is the + binding constraint.** + +### MEASURED vs INFERRED + +**MEASURED:** the catch-rate and latency tables on 67 065 true fires from 70 +recorded battles; the 100% attribution of every miss to the `+3*power` bonus or +to our own damage (both read from the corpus's `hit` events); the live scan +interval (3024/3024 readings `+1`); the four-way correlation table; the Gate B +log-loss; the unit tests (13/13) and env-report guard (25/25); the clean-archive +(`git archive HEAD | tar -x`) compile of `ModularBot` and the fire-fix tests. +**INFERRED:** that the correction transfers live with the same tick alignment as +the corpus — the corpus's event/row offset is a capture artifact (the live event +and the reading are delivered in the same turn), and this was **not** confirmed +in a live battle (Task D skipped). **NOT MEASURED:** the live movement effect of +the fix. + +**Status: the shipped movement default is UNCHANGED (`TR_MOVEMENT=strafe`); the +detector fix is ON by default behind `TR_STRAFE_FIRE_FIX` (revert with +`TR_STRAFE_FIRE_FIX=0`).**