1ed7797cb6
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay) - New modules: minimum-risk melee movement, spinning melee radar - New test bots: PatternMover, RandomMover, WaveSurfer - Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning - Fixed: TM gun warmup gating + directional residuals - Fixed: circular gun integrated formula + multi-bin omega cache - Fixed: oscillator wall-bounce lockout - Fixed: phantom meteor perpendicular body orientation - 6/6 battle wins across all enemy types
391 lines
27 KiB
Markdown
391 lines
27 KiB
Markdown
# Biology Domain Report: Learning Mechanisms for BNNBot
|
||
|
||
**Context**: Binary neural network (690-bit input, online learning, no gradient descent, reward from wave-hit miss distance, ~1ms/tick budget).
|
||
|
||
---
|
||
|
||
## 1. Three-Factor Learning Rules (Neuromodulation)
|
||
|
||
### How it works biologically
|
||
In cortex, synaptic plasticity depends on three signals simultaneously: pre-synaptic activity, post-synaptic activity, and a neuromodulator (dopamine for reward, acetylcholine for attention, norepinephrine for arousal). The local Hebbian correlation (pre × post) creates an **eligibility trace** — a molecular tag that marks a synapse as "recently active." The neuromodulator gates whether that trace converts to actual weight change. This solves the temporal credit assignment problem: the trace persists for seconds, so a delayed reward can still modulate the right synapses.
|
||
|
||
### Core algorithmic principle
|
||
```
|
||
eligibility[w] += pre_activity * post_activity * decay
|
||
Δw = eligibility[w] * reward_signal(t)
|
||
```
|
||
The trace bridges the gap between action (synapse fires) and evaluation (reward arrives). Different layers can have different decay constants, giving the reward signal variable reach depth.
|
||
|
||
### Mapping to our system
|
||
- Wave hit system fires a reward signal at a delay (bullet travel time). This is exactly the temporal gap eligibility traces are built for.
|
||
- Each binary synapse (XOR input, threshold sum) computes pre × post locally at no extra cost — the product is 1 only if both sides were active.
|
||
- The reward scalar (miss distance inverted) becomes the neuromodulator: `Δw ∝ eligibility * (1 / miss_distance)`.
|
||
- Deeper layers get a weaker or differently-shaped neuromodulatory signal — natural depth-aware credit assignment without backprop.
|
||
|
||
### Update rule sketch
|
||
```
|
||
# Per synapse, per tick:
|
||
e[i,j] += pre[i] * post[j]
|
||
e[i,j] *= decay # e.g. 0.95/tick
|
||
|
||
# On wave hit:
|
||
reward = 1.0 - (miss_px / arena_width)
|
||
for each (i,j) in layer:
|
||
w[i,j] += lr * e[i,j] * reward
|
||
e[i,j] = 0 # consume trace
|
||
```
|
||
|
||
### Strengths / weaknesses
|
||
+ Handles delayed reward natively (critical for our wave system).
|
||
+ Works with binary activations — pre=0/1, post=0/1, product is cheap.
|
||
+ Scales well: O(weights), no global error computation needed.
|
||
+ Different layers can use different lr or different eligibility decay → implicit layer-wise credit assignment.
|
||
- Variance is high with sparse binary activations (most products are 0). Need enough co-activations to get signal.
|
||
- Reward signal is scalar; it's not instructive (doesn't say which direction to move). Works best when combined with exploration.
|
||
|
||
---
|
||
|
||
## 2. Cerebellar Learning (Supervised via Error Signal + Timing)
|
||
|
||
### How it works biologically
|
||
The cerebellum is the brain's timing engine. Granule cells provide a massive population code of context (parallel fibers). Purkinje cells are the output neurons. Climbing fibers from the inferior olive deliver a precise error signal ("something went wrong") that triggers long-term depression (LTD) at the active parallel fiber → Purkinje cell synapses. The critical detail: **only the synapses that were active just before the error are depressed** — this is temporal specificity without backpropagation.
|
||
|
||
Recent (2024) work shows granule cells ramp at different rates toward a reward moment, and the climbing fiber spike at reward time selectively strengthens the synapses whose ramp was timed correctly. This is literally: learn to predict when the bullet arrives.
|
||
|
||
### Core algorithmic principle
|
||
```
|
||
# Climbing fiber = error event at time T
|
||
# Depress all parallel fiber synapses that were active in window [T-Δ, T]
|
||
for synapse active in recent_window:
|
||
w -= lr * (was_active AND error_occurred)
|
||
```
|
||
It's a delayed anti-Hebbian rule gated by an error signal. No error = no change. Error = punish recently-active pathways.
|
||
|
||
### Mapping to our system
|
||
The wave system is essentially a climbing fiber: it fires an error/reward signal at a defined time (when the wave reaches the enemy). Our predictor was active with specific binary patterns when the shot was committed. We can depress the synapses that fired and caused a miss, or potentiate those that fired and caused a near-hit.
|
||
|
||
The granule cell population code idea maps directly to our 690-bit sparse input — we already have a rich, diverse temporal representation that can encode "which part of the trajectory context was active."
|
||
|
||
### Update rule sketch
|
||
```
|
||
# Store recent binary activations per layer (ring buffer, ~50 ticks)
|
||
on wave_hit(miss_distance, ticks_ago):
|
||
reward = clip(1.0 - miss_distance/threshold, -1, 1)
|
||
pattern = activation_history[ticks_ago]
|
||
for each active synapse in pattern:
|
||
w += lr * reward # LTP on near-hit, LTD on miss
|
||
```
|
||
|
||
### Strengths / weaknesses
|
||
+ Elegant credit assignment without any backward pass: just "what was active before the error?"
|
||
+ Temporal specificity is free — ring buffer of activations is cheap.
|
||
+ Directly analogous to our wave timing structure.
|
||
+ The granule cell / parallel fiber idea suggests we want MANY diverse binary features upstream — our 690-bit encoding already does this.
|
||
- Classic cerebellar model is essentially supervised (climbing fiber knows the correct output). Our reward is scalar, not a correction vector. We get "wrong" but not "which direction is right."
|
||
- Works better with many output cells (Purkinje cells) each tuned to a sub-task. With a single X/Y output we lose diversity.
|
||
|
||
---
|
||
|
||
## 3. Immune Clonal Selection + Somatic Hypermutation
|
||
|
||
### How it works biologically
|
||
When a pathogen appears, B-cells with partial affinity are selected and cloned. Clones undergo somatic hypermutation — rapid, local random changes to antibody genes (not genome-wide). Clones with better affinity bind more antigen and survive; others die. The result is rapid local optimization starting from a working seed. No central controller, no gradient — just: copy, mutate locally, select, repeat.
|
||
|
||
### Core algorithmic principle
|
||
```
|
||
population = [copy(best) for _ in range(n_clones)]
|
||
for clone in population:
|
||
mutate locally with rate ∝ 1/affinity # better clones mutate less
|
||
affinity = evaluate(clone)
|
||
keep top k
|
||
```
|
||
This is a hill-climber with adaptive mutation radius and population diversity. The key insight: **mutation rate inversely proportional to current performance** — near-optimal solutions mutate less, preventing regression.
|
||
|
||
### Mapping to our system
|
||
We can maintain a small population (e.g., 4–8) of weight vectors for the readout layer. Every N ticks, clone the best-performing weight vector, apply random bit-flips to the binary weights with probability ∝ miss_distance (worse performance = more aggressive mutation). Test clones against incoming wave hits. Keep survivors.
|
||
|
||
This sidesteps gradient computation entirely: evaluation is the reward signal itself.
|
||
|
||
### Update rule sketch
|
||
```
|
||
# Every 20 ticks:
|
||
clones = [flip_bits(best_weights, rate=miss_ema) for _ in range(4)]
|
||
# Over next 20 ticks, score each clone on wave hits
|
||
best_weights = clone with lowest average miss distance
|
||
```
|
||
|
||
### Strengths / weaknesses
|
||
+ Zero gradient computation. Works on discrete/binary weights directly.
|
||
+ Adaptive mutation rate naturally does simulated annealing.
|
||
+ Population diversity prevents local minima.
|
||
+ Compatible with any loss landscape shape.
|
||
- Requires parallel evaluation — need to run multiple weight hypotheses simultaneously and compare. Adds memory.
|
||
- Convergence is slower than gradient methods for smooth losses (but our loss is NOT smooth — binary weights make gradient methods useless here anyway).
|
||
- Population size ≥ 2 costs memory; with binary weights and ~100 readout weights this is negligible.
|
||
|
||
**This is underexplored for binary networks where gradients are meaningless. Strongest candidate here.**
|
||
|
||
---
|
||
|
||
## 4. Kauffman Boolean Networks / Edge-of-Chaos Self-Organization
|
||
|
||
### How it works biologically
|
||
Kauffman's NK Boolean networks model gene regulatory networks. Each gene node has K inputs from other nodes and a random Boolean function. With K=2, networks sit at a phase transition: ordered (frozen, no computation) vs. chaotic (noise floods signal). At K≈2 (the "edge of chaos"), networks show maximal information transmission, long transients before attractors, and sensitivity to initial conditions without instability.
|
||
|
||
Living cells appear to self-tune toward this critical point. Perturbations propagate far but don't explode exponentially.
|
||
|
||
### Core algorithmic principle
|
||
The edge-of-chaos is an architectural prior: design connectivity so each node has ≈ 2 inputs. This maximizes the repertoire of distinguishable states the network can encode, making the network a rich feature extractor with minimal parameters.
|
||
|
||
For learning: networks can be guided toward the edge by rewarding connectivity patterns that show intermediate sensitivity (not frozen, not chaotic), using damage spreading tests.
|
||
|
||
### Mapping to our system
|
||
Our XOR+popcount+threshold neurons already have arbitrary fan-in. If we constrain each neuron to K≈2 inputs (or K≈4 for richer functions), we could build a Boolean reservoir that sits at the edge of chaos and provides a rich nonlinear projection of the 690-bit input. The readout layer then does the learning.
|
||
|
||
This is essentially "design the hidden layer using Kauffman's criticality principle."
|
||
|
||
### Strengths / weaknesses
|
||
+ Principled architecture design without training hidden layers.
|
||
+ Maximizes representational capacity of fixed random layer.
|
||
+ Binary everywhere — exact fit.
|
||
- It's an architectural heuristic, not a learning algorithm. Still needs a supervised/RL readout.
|
||
- Self-tuning criticality during online battle is complex; better as initialization strategy.
|
||
- K=2 limits each neuron's discriminability; our current threshold neurons with high fan-in may already be better.
|
||
|
||
---
|
||
|
||
## 5. Octopus: Hierarchical Autonomy with Peripheral Pre-processing
|
||
|
||
### How it works biologically
|
||
The octopus has ~500M neurons, but only 36% are in the central brain. Each arm has its own mini-brain that executes local motor programs (grasp, explore, retract) without waiting for central commands. The central brain issues high-level goals ("reach there"); the arm negotiates locally with its environment. Crucially: **each peripheral layer learns what it needs for its task**, decoupled from the other layers.
|
||
|
||
### Core algorithmic principle
|
||
Hierarchical decomposition where each processing stage has its own local learning objective, not a globally supervised one. The central system trains on outcomes; the periphery trains on local sensory prediction (self-supervised). This creates natural credit assignment boundaries: you don't need to propagate error through the arm's local controller to train the central goal selector.
|
||
|
||
### Mapping to our system
|
||
This suggests a **two-stage architecture**:
|
||
- Stage 1 (peripheral, fixed or self-supervised): encode the 690-bit input into a compressed, locally predictive representation. Train to predict next frame from current frame (self-supervised). No reward needed here.
|
||
- Stage 2 (central, RL-trained): use the stage-1 representation as input. This layer receives the wave hit reward and trains with Hebbian/three-factor rules.
|
||
|
||
Stage 1 insulates stage 2 from input noise and does feature extraction. Stage 2 does the goal-directed aiming.
|
||
|
||
### Strengths / weaknesses
|
||
+ Clean credit assignment: stage 2 trains on reward, stage 1 trains on local prediction error.
|
||
+ Self-supervised stage 1 can learn from every tick (no reward needed), converges faster.
|
||
+ Modular: can replace either stage independently.
|
||
- Two-stage adds complexity. For our problem (movement is 97.8% constant-velocity), stage 1 may not be worth the cost.
|
||
- Self-supervised prediction of next frame is essentially "predict the enemy continues straight" — which linear extrapolation already does perfectly.
|
||
|
||
---
|
||
|
||
## 6. Slime Mold (Physarum polycephalum): Flow-Reinforced Network Topology
|
||
|
||
### How it works biologically
|
||
*Physarum* is a single-celled organism (no neurons) that solves shortest-path problems. Its body is a network of cytoplasmic tubes. Food sources cause oscillatory contractions that push fluid through tubes. Tubes that carry more flow grow thicker (lower resistance → more flow → more growth). Tubes that carry less flow shrink and disappear. The result: the network self-organizes to the minimum Steiner tree connecting food sources.
|
||
|
||
This is **Hebbian learning on a physical network**: use it more → strengthen it. The "signal" is flow (pressure differential), the "synapse" is tube cross-section.
|
||
|
||
### Core algorithmic principle
|
||
```
|
||
# For each edge (i,j) with flow Q[i,j]:
|
||
D[i,j] += alpha * (Q[i,j] - beta * D[i,j]) # grow with flow, decay otherwise
|
||
```
|
||
Where D is tube diameter (conductance). This is a local, flow-proportional reinforcement rule with decay. No central controller.
|
||
|
||
### Mapping to our system
|
||
Map "flow" to activation frequency: a binary weight that fires more often under successful predictions gets stronger. This is a frequency-weighted Hebbian rule:
|
||
```
|
||
Δw[i,j] += alpha * activation_freq[i,j] * reward_ema - beta * w[i,j]
|
||
```
|
||
The decay term (-beta * w) prevents runaway potentiation (Physarum tubes that aren't used shrink). This gives us **structural plasticity**: unused pathways die, active+rewarded pathways survive.
|
||
|
||
This is subtly different from standard Hebbian learning: the reinforcement is on **path flow** (consistent activation patterns over time), not just single-trial coincidence. It's an exponential moving average of "did this synapse contribute to successes?"
|
||
|
||
### Strengths / weaknesses
|
||
+ Extremely simple update rule. Naturally implements Occam's razor: unused connections prune themselves.
|
||
+ Decay term prevents runaway weights and provides implicit regularization.
|
||
+ The flow-reinforcement principle is perfect for our wave system: each wave hit tells us which connections were "on the path" to the prediction.
|
||
- Physarum solves shortest-path, which is a topological problem. Our problem is regression (predict X, Y). The flow metaphor requires careful translation.
|
||
- Pruning live connections during battle could cause instability if exploration is needed.
|
||
|
||
---
|
||
|
||
## 7. Predictive Processing / Free Energy Minimization
|
||
|
||
### How it works biologically
|
||
Karl Friston's free energy principle proposes the brain minimizes "surprise" by constantly predicting its sensory inputs and updating internal models when predictions fail. Prediction errors flow upward (bottom-up), predictions flow downward (top-down). Each layer only sees the prediction error from the layer below, not the raw input. Learning minimizes the sum of prediction errors across all layers simultaneously — this turns out to be mathematically equivalent to variational inference, and approximately equivalent to backpropagation under certain conditions.
|
||
|
||
### Core algorithmic principle
|
||
Each layer maintains a prediction of the layer below's state and updates based on the error:
|
||
```
|
||
error[l] = actual[l] - prediction[l]
|
||
prediction[l] = W[l+1] * activity[l+1] # top-down
|
||
Δw[l] = lr * error[l] * activity[l+1] # Hebbian on error
|
||
```
|
||
No global error required — each layer minimizes its local prediction error. This achieves approximate credit assignment through depth.
|
||
|
||
### Mapping to our system
|
||
Train the lower layers to predict the next input frame (self-supervised, every tick). Train the top layer to minimize aiming error (RL, every wave hit). The lower layers develop representations that are useful for prediction, which in turn aids the top layer.
|
||
|
||
For binary networks: the "prediction" is a reconstructed binary pattern, and the "error" is the XOR (bitwise difference) between predicted and actual pattern.
|
||
|
||
### Strengths / weaknesses
|
||
+ Every tick provides a learning signal for lower layers (no need to wait for wave hits).
|
||
+ Layer-local updates — no need to propagate anything through the full depth.
|
||
+ Mathematically principled; proven to approximate backprop in continuous case.
|
||
- In binary networks, the "prediction error" (XOR) is not differentiable and doesn't provide a directional update signal — just "wrong" or "right" per bit.
|
||
- The top-down prediction path requires an explicit generative model (decoder weights), doubling parameter count.
|
||
- For our near-linear enemy motion, prediction error will be near-zero almost always, starving the learning signal.
|
||
|
||
---
|
||
|
||
## 8. Reward-Modulated STDP (R-STDP) with Eligibility Traces
|
||
|
||
### How it works biologically
|
||
Spike-timing-dependent plasticity (STDP): if pre fires before post within ~20ms, potentiate; if post fires before pre, depress. R-STDP adds a third factor: the dopaminergic reward signal gates the eligibility trace that STDP creates. The eligibility trace is a slow-decaying shadow of the Hebbian correlation; only when reward arrives does it convert to permanent weight change.
|
||
|
||
### Core algorithmic principle (adapted for rate-coded binary neurons)
|
||
```
|
||
# At each tick:
|
||
e[i,j] = pre[i] * post[j] - (1-pre[i]) * post[j] * depression_factor
|
||
e[i,j] *= trace_decay
|
||
|
||
# At wave hit:
|
||
Δw[i,j] = lr * e[i,j] * reward
|
||
```
|
||
The anti-Hebbian term (inactive pre, active post) prevents post-synaptic neurons from firing without cause — this is the "STDP asymmetry" that drives selectivity.
|
||
|
||
### Mapping to our system
|
||
Direct application: run R-STDP on the readout layer with wave-hit reward as the neuromodulator. For binary neurons, "pre fires before post" translates to "this input bit was 1 when this output bit became 1." The eligibility trace naturally bridges the bullet travel delay.
|
||
|
||
This is a well-studied, published combination: R-STDP + binary/SNN networks + reward-based learning.
|
||
|
||
### Strengths / weaknesses
|
||
+ Well-studied, proven to train binary/spiking networks on classification tasks.
|
||
+ Eligibility trace handles delayed reward explicitly.
|
||
+ Anti-Hebbian depression term improves specificity over pure Hebbian.
|
||
- Standard R-STDP trains only the readout layer well; deep credit assignment remains hard.
|
||
- In rate-coded binary (not spike-timing) settings, the "timing" aspect is lost — reduces to three-factor Hebbian.
|
||
|
||
---
|
||
|
||
## 9. Lateral Inhibition / Winner-Take-All for Representation Learning (LESS KNOWN angle)
|
||
|
||
### How it works biologically
|
||
In cortex, excitatory neurons compete via inhibitory interneurons. Only the most-activated cell "wins" and suppresses its neighbors. This WTA dynamic creates **sparse, non-overlapping representations** where each input pattern activates only a few neurons. Used in the olfactory bulb, hippocampus (sparse place cells), and visual cortex.
|
||
|
||
The key learning rule is **SoftHebb**: Hebbian updates gated by the WTA competition. The winner potentiates its input weights; losers depress theirs. No labels, no reward — purely competitive.
|
||
|
||
### Core algorithmic principle
|
||
```
|
||
winner = argmax(activation_vector)
|
||
Δw[winner, :] = lr * (input - w[winner, :]) # move toward input
|
||
Δw[losers, :] = -small_lr * input # move away
|
||
```
|
||
This is online k-means / vector quantization, but implemented as neural competition.
|
||
|
||
### Mapping to our system
|
||
Use a WTA hidden layer between input (690 bits) and readout. The WTA layer creates a sparse binary code from the input — essentially a hash into a learned codebook of "enemy movement contexts." The readout then maps each context-hash to an (x, y) correction.
|
||
|
||
This extends our current Hebbian residual table (8 heading sectors × 3 distance bands = 24 cells) to a LEARNED partitioning with hundreds of cells, found automatically from data rather than hand-coded.
|
||
|
||
### Update rule sketch
|
||
```
|
||
# Hidden layer: WTA competition
|
||
scores = W_hidden @ input_bits # shape: [n_hidden]
|
||
k_active = top_k_indices(scores, k=5) # k-sparse activation
|
||
binary_hidden = sparse_binary(k_active, n_hidden)
|
||
|
||
# Hebbian update on hidden layer weights (unsupervised, every tick):
|
||
for i in k_active:
|
||
W_hidden[i] += lr_unsup * (input_bits - W_hidden[i])
|
||
|
||
# Readout update (on wave hit):
|
||
Δw_readout = lr_rl * binary_hidden * reward
|
||
```
|
||
|
||
### Strengths / weaknesses
|
||
+ Replaces hand-coded sector table with a learned, data-driven partitioning.
|
||
+ Unsupervised layer learns from every tick; RL layer learns from wave hits. Two timescales.
|
||
+ Sparse binary hidden code is efficient and interpretable.
|
||
+ Scales naturally to hundreds of "cells" vs. our current 24.
|
||
+ SoftHebb (2022) showed this achieves competitive image classification without backprop.
|
||
- The hidden layer learns the input distribution, not the reward structure — may learn irrelevant features.
|
||
- k-sparse WTA is sensitive to the choice of k and n_hidden. Needs tuning.
|
||
|
||
---
|
||
|
||
## 10. Plant Root Tropism: Gradient-Free Local Search with Memory (VERY UNDEREXPLORED)
|
||
|
||
### How it works biologically
|
||
Plant roots grow toward nutrients (chemotropism), water (hydrotropism), and away from toxins. Each root tip integrates local chemical gradients over its surface and biases growth direction. No central nervous system, no backprop. The mechanism: differential auxin concentration across the root tip causes asymmetric elongation. The root "bends" toward the side with less auxin (more elongation).
|
||
|
||
Crucially: **the root remembers where it came from** (gravitropism provides a reference frame), and successful growth (reaching nutrient) is consolidated by lateral root branching — the successful path is reinforced structurally.
|
||
|
||
### Core algorithmic principle
|
||
```
|
||
# Local gradient sensing:
|
||
direction_bias = integrate(local_signal, window=tip_width)
|
||
grow(direction_bias)
|
||
|
||
# Reinforcement on success:
|
||
if nutrient_found:
|
||
branch here # spawn lateral root from successful path
|
||
consolidate path weights # auxin redistribution
|
||
```
|
||
This is a stochastic local search with: (a) local gradient estimation from small perturbations, (b) structural reinforcement of successful paths, (c) no global controller.
|
||
|
||
### Mapping to our system
|
||
The "root tip" is the current weight vector. "Growing" is perturbing weights. "Nutrient gradient" is the miss distance improvement over recent ticks. The plant root insight is: **estimate gradient locally by trying slightly perturbed directions and measuring which direction improves reward** — this is exactly node/weight perturbation learning.
|
||
|
||
The "lateral branching on success" maps to: when performance improves significantly, save the current weight vector as a new candidate in a small ensemble.
|
||
|
||
### Strengths / weaknesses
|
||
+ Gradient-free by construction — works on any loss landscape.
|
||
+ "Branch on success" idea gives a cheap ensemble without parallel evaluation.
|
||
+ Local perturbation + reward correlation = unbiased gradient estimate (weight perturbation theorem).
|
||
- Perturbation-based gradient estimates have high variance with many weights.
|
||
- For binary weights, even a single bit flip is a discrete jump, not an infinitesimal perturbation. Must flip multiple bits to get meaningful reward signal difference.
|
||
|
||
---
|
||
|
||
## Cross-Cutting Synthesis
|
||
|
||
### The Four Most Promising Mechanisms for BNNBot
|
||
|
||
**Tier 1 (implement now):**
|
||
|
||
1. **Three-factor Hebbian + Eligibility Traces** — maps directly onto existing Hebbian residual table. Add eligibility trace to bridge the bullet travel delay. Extends the current architecture minimally.
|
||
|
||
2. **Clonal Selection (Immune)** — for binary weight vectors, gradient descent is meaningless. Maintain 2–4 clones of the readout weights, mutate with rate ∝ miss_distance, keep winner. This is the correct optimizer for discrete weight spaces. ~10 lines of code.
|
||
|
||
**Tier 2 (worth exploring):**
|
||
|
||
3. **WTA + SoftHebb hidden layer** — replaces the hand-coded 8×3 sector table with a learned ~256-cell partitioning. Every tick provides an unsupervised update; every wave hit provides a reward update. Two decoupled learning loops.
|
||
|
||
4. **Physarum flow-reinforcement (structural pruning)** — add a decay term to the current weight update. Connections that fire often under reward survive; others die. Implicit regularization.
|
||
|
||
**Tier 3 (architectural ideas, not ready for implementation):**
|
||
|
||
5. **Kauffman edge-of-chaos** — use as an initialization principle for any fixed random hidden layer: set K≈2 per neuron for maximal computational richness.
|
||
|
||
6. **Cerebellar granule cell ramp** — train binary hidden cells to ramp (increase activation rate) toward the expected bullet arrival time. Those that are active at hit time get potentiated. Requires temporal structure the current encoding lacks.
|
||
|
||
### Key Insight from Biology
|
||
|
||
Every mechanism above shares one property: **local information + delayed global signal = credit assignment**. The brain never backpropagates a gradient; it stores a **trace** of what was active, then modulates that trace when a reward arrives. The three-factor rule is the minimal mathematical expression of this principle. Everything else (cerebellar timing, immune cloning, Physarum flow) is a specialization of it to a particular domain.
|
||
|
||
For our problem: the wave system already provides the delayed global signal. The missing piece is the **trace** — recording what the network did when it committed to a prediction, then crediting or debiting those activations when the wave hit occurs.
|
||
|
||
---
|
||
|
||
## Appendix: Less-Known Mechanisms (Brief)
|
||
|
||
**Bacterial chemotaxis (run-and-tumble):** E. coli compares current receptor occupancy to a memory of receptor occupancy ~1 second ago. If occupancy improved, bias toward running; if worsened, tumble (random new direction). No neurons, no gradient. Pure temporal comparison. Maps to: compare current miss_distance to miss_distance 20 ticks ago; if improved, continue current weight direction; if worsened, randomize. The "memory" is a single exponential moving average.
|
||
|
||
**Voltage-gated calcium channels as eligibility trace substrate:** In real neurons, calcium transients triggered by coincident pre-post firing persist for hundreds of milliseconds and activate CaM-KII, which phosphorylates AMPA receptors. This is the molecular implementation of the eligibility trace. For us: any variable that accumulates during active periods and decays when inactive is a valid trace substrate. A simple integer counter per weight, incremented on pre×post=1 and multiplied by 0.95 each tick, is the computational equivalent.
|
||
|
||
**Neurogenesis under stress (hippocampus):** The hippocampus generates new neurons under novelty/stress. New neurons have lower activation thresholds and are more plastic. They either get incorporated (if they contribute) or die within weeks. For us: "spawn" new binary neurons (random weight vectors) when miss distance spikes (enemy starts dodging). Test them for 50 ticks. Keep those that improve prediction; discard others. This is a second form of clonal selection operating at the architectural level.
|
||
|
||
**Octopus chromatophore pattern learning:** Octopuses can learn to flash specific chromatophore patterns to match backgrounds they've never seen, using reinforcement from camouflage success. Each skin patch (papilla) has local autonomy but is coordinated by central pattern generators. The learning is: local patches learn to respond to local visual input; global coordination emerges from shared output constraints. Maps to: local hidden neurons learn from local input sub-regions; global readout imposes output coordination.
|