j130 learned movement: outcome label (P(hit|state,g)) mode + Gate A; pre-registered outcome arms
This commit is contained in:
@@ -0,0 +1,38 @@
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# arms_movement_outcome.txt — the OUTCOME-LABELLED learned-movement arms
|
||||
# (job j130), on the FROZEN panel tools/ab/panel_movement.txt.
|
||||
#
|
||||
# Pre-registered in docs/movement_campaign.md, "Learned movement — outcome
|
||||
# label (P(hit))", BEFORE any battle. REFERENCE is `strafe` — the SHIPPED
|
||||
# champion. Every delta is (arm − strafe).
|
||||
#
|
||||
# j128 measured `corr( P(arrival bin), P(hit | arrival bin) ) = -0.342` over the
|
||||
# 31 bins: minimising the resolved-position histogram steers INTO the bullets.
|
||||
# j130 replaces the label with the dense outcome
|
||||
# hit(state, g) = hit and |g - b_our| <= window(wave)
|
||||
# and learns P(hit | state, candidate g) with a counted 2-class SBC.
|
||||
#
|
||||
# Gate A (common_libs/tests/outcome_label_gate.py, corpus /tmp/tfil_ab2/out):
|
||||
# the alignment correlation flips to +0.566 (histogram -0.341), so the veto
|
||||
# does NOT fire; the state-conditional outcome model however is NOT better than
|
||||
# the state-free one on held-out log-loss, and the open-loop decision
|
||||
# counterfactual barely moves (3.53% -> 3.33%). The live panel decides.
|
||||
#
|
||||
# Format: name | ENV=value ENV=value | label
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
# 1. THE CHAMPION — the arm a challenger has to beat (round wins + hit rate).
|
||||
strafe | TR_MOVEMENT=strafe | champion/reference — shipped strafe defaults
|
||||
|
||||
# 2. THE OLD LABEL — j128's state-conditional counted SBC (arrival-bin label),
|
||||
# so the new label is isolated against the old one on the same binary.
|
||||
learned | TR_MOVEMENT=learned | j128 histogram label (arrival bin)
|
||||
|
||||
# 3. THE NEW LABEL — the same mover, same geometry, same counted+decayed SBC
|
||||
# and penalties; only the training label changes (dense hit outcome).
|
||||
learned_outcome | TR_MOVEMENT=learned TR_LEARNED_LABEL=outcome | outcome label P(hit | state, g)
|
||||
|
||||
# 4. INFORMATION CONTROL — the outcome label with the state forced to one cell
|
||||
# (state-free outcome model). Isolates whether the state carries anything
|
||||
# under the new label (Gate A says it does not).
|
||||
learned_outcome_global | TR_MOVEMENT=learned TR_LEARNED_LABEL=outcome TR_LEARNED_GLOBAL=1 | outcome label, state conditioning OFF
|
||||
Reference in New Issue
Block a user