254c7dc997
- Gun harness: virtual bullet tracker, rolling fitness, auto-selector - Guns: head-on, linear (extrapolation), circular (integrated formula), tsetlin machine (learning) - Movement: phantom meteor gravity engine (danger histograms, phantom bullets, fire detection) - Radar: harness + radar_lock adapter - Color-coded modules: turret/bullet color per gun, body per movement, scan per radar - Beats Target, SpinBot, Crazy, TrackFire in 10-round battles
4.5 KiB
4.5 KiB
BNNBot Research Brief
Problem
Predict enemy future position in Robocode Tank Royale to aim bullets accurately. The prediction must happen online (during battle), without pre-training.
Hard Constraints
- NO supervised learning (no labeled input→output training pairs)
- NO gradient descent (no derivatives, no surrogate gradients, no STE)
- Online learning only — must learn and improve during a single battle
- Computational budget: ~1ms per tick
- Binary-friendly (690-bit input encoding already exists)
Allowed
- Backpropagation of SIGNALS (non-gradient information flowing backward through layers)
- Reinforcement learning (reward signal available from wave hit system)
- Self-supervised learning
- Unsupervised learning
- Network structure modification during runtime
Current Architecture
Input Engineering (binary_encoding.nim)
- 690-bit binary vector: 10 frames × 69 bits
- Per frame: bearing sin/cos (16b), distance (7b), velocity (5b), heading sin/cos (16b), enemy X/Y position (14b), enemy energy (11b)
- Gray-coded for Hamming distance smoothness
- Temporal window: 10 most recent radar scans
Feedback System (wave system in BNNBot.nim)
- Every tick: 10 circular waves spawned at bot position
- Powers: 0.1 to 3.0 (10 levels), speeds: 19.7 to 11.0 px/tick
- When wave radius reaches enemy: records enemy state as 69-bit frame
- Provides ground truth: "if you fired at power X, the enemy would be HERE when the bullet arrives"
- ~95.6% hit rate in testing
Current Predictor (predictor.nim)
- Linear extrapolation: predicted = current_pos + velocity * ticks_to_arrival
- Hebbian residual table: 8 heading sectors × 3 distance bands = 24 cells
- Each cell stores (correction_x, correction_y), updated online with lr=0.2
- Backtest results: 14-21% MAE reduction over pure linear extrapolation
- Converges within one battle (MAE 14.85 → 2.62, first 50 vs last 50 rows)
Key Findings
Data Analysis (analysis/report.txt)
- Enemy movement is 97.8% constant-velocity straight lines
- Acceleration is negligible (std 0.25-0.49 px/tick²)
- Heading is very stable across 10-frame windows
- Linear extrapolation MAE: 15-27px (1.5-2.7% of arena)
- Distance to enemy is the main error driver
- Scalar velocity alone is weak predictor (r=0.15); directional velocity from frame deltas is strong
Backtest Results (analysis/backtest_report.txt)
- P1 (linear): MAE 8.1-18.7 encoded units
- P2 (weighted 4-frame): ~8% improvement, trivial cost
- P3 (linear + Hebbian residual): 14-21% improvement, converges fast
- Most residual table cells stay empty — only ~10/24 activate
Encoding Insights
- sin/cos angle encoding avoids wraparound discontinuity — worth the extra bits
- Enemy X/Y position partially redundant with bearing+distance (encodes absolute position)
- Wall distance → XY% compression saved 140 bits losslessly
- Self-state removed (not needed for aiming)
What We've Tried
- ✅ Input engineering with Gray coding and temporal window — works well
- ✅ Linear extrapolation — strong baseline, 15-27px error
- ✅ Hebbian residual table — learns online, 14-21% improvement
- ❌ Pure XOR layer stacking — collapses (associative, no non-linearity)
- ❌ XOR + AND layers — AND with fixed mask is still linear over GF(2)
- ✅ XOR + popcount + threshold = valid binary neuron (non-linear)
Open Questions
- Can we go deeper than the current shallow predictor while respecting the constraints?
- What non-gradient learning rules can train multi-layer binary networks?
- Can the temporal structure (10 frames) be exploited by the network architecture?
- Is there a way to do credit assignment through depth without gradients?
- Can the wave hit system provide richer learning signal than just miss distance?
Architecture Philosophy
- Input engineering IS the feature hierarchy (handcrafted, domain-informed)
- Current approach is essentially reservoir computing: rich fixed features → simple learnable readout
- Question: can we do better with a learnable feature extractor, or is the handcrafted one already near-optimal?
Files
src/BNNBot.nim— main bot, wave system, integrationsrc/binary_encoding.nim— 690-bit input encodingsrc/predictor.nim— linear extrapolation + Hebbian residual tableanalysis/correlations.py— data analysis scriptanalysis/backtest.py— predictor comparison scriptanalysis/report.txt— correlation analysis resultsanalysis/backtest_report.txt— predictor backtest resultsdata/— CSV battle logs (enabled via BNNBOT_CSV=1)