Files
SirRoboGarage/tools/ab/arms_movement_outcome.txt

39 lines
2.5 KiB
Plaintext
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ─────────────────────────────────────────────────────────────────────────────
# arms_movement_outcome.txt — the OUTCOME-LABELLED learned-movement arms
# (job j130), on the FROZEN panel tools/ab/panel_movement.txt.
#
# Pre-registered in docs/movement_campaign.md, "Learned movement — outcome
# label (P(hit))", BEFORE any battle. REFERENCE is `strafe` — the SHIPPED
# champion. Every delta is (arm − strafe).
#
# j128 measured `corr( P(arrival bin), P(hit | arrival bin) ) = -0.342` over the
# 31 bins: minimising the resolved-position histogram steers INTO the bullets.
# j130 replaces the label with the dense outcome
# hit(state, g) = hit and |g - b_our| <= window(wave)
# and learns P(hit | state, candidate g) with a counted 2-class SBC.
#
# Gate A (common_libs/tests/outcome_label_gate.py, corpus /tmp/tfil_ab2/out):
# the alignment correlation flips to +0.566 (histogram -0.341), so the veto
# does NOT fire; the state-conditional outcome model however is NOT better than
# the state-free one on held-out log-loss, and the open-loop decision
# counterfactual barely moves (3.53% -> 3.33%). The live panel decides.
#
# Format: name | ENV=value ENV=value | label
# ─────────────────────────────────────────────────────────────────────────────
# 1. THE CHAMPION — the arm a challenger has to beat (round wins + hit rate).
strafe | TR_MOVEMENT=strafe | champion/reference — shipped strafe defaults
# 2. THE OLD LABEL — j128's state-conditional counted SBC (arrival-bin label),
# so the new label is isolated against the old one on the same binary.
learned | TR_MOVEMENT=learned | j128 histogram label (arrival bin)
# 3. THE NEW LABEL — the same mover, same geometry, same counted+decayed SBC
# and penalties; only the training label changes (dense hit outcome).
learned_outcome | TR_MOVEMENT=learned TR_LEARNED_LABEL=outcome | outcome label P(hit | state, g)
# 4. INFORMATION CONTROL — the outcome label with the state forced to one cell
# (state-free outcome model). Isolates whether the state carries anything
# under the new label (Gate A says it does not).
learned_outcome_global | TR_MOVEMENT=learned TR_LEARNED_LABEL=outcome TR_LEARNED_GLOBAL=1 | outcome label, state conditioning OFF