1ed7797cb6
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay) - New modules: minimum-risk melee movement, spinning melee radar - New test bots: PatternMover, RandomMover, WaveSurfer - Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning - Fixed: TM gun warmup gating + directional residuals - Fixed: circular gun integrated formula + multi-bin omega cache - Fixed: oscillator wall-bounce lockout - Fixed: phantom meteor perpendicular body orientation - 6/6 battle wins across all enemy types
765 lines
38 KiB
Markdown
765 lines
38 KiB
Markdown
# BNNBot Deep Learning: Gradient-Free Methods for Binary Networks
|
||
|
||
**Context**: 690-bit binary input, online battle learning, reward = miss distance from circular wave
|
||
feedback, binary weights/activations preferred. No gradient descent, no STE, no supervised labels.
|
||
Backprop of *signals* (not gradients) is OK.
|
||
|
||
---
|
||
|
||
## Evaluation Scorecard Legend
|
||
|
||
| Symbol | Meaning |
|
||
|--------|---------|
|
||
| ✅ | Full support |
|
||
| ⚠️ | Partial / with caveats |
|
||
| ❌ | Fails this criterion |
|
||
|
||
**Criteria for each method:**
|
||
|
||
1. **No GD** — truly avoids gradient descent (no derivatives, no STE, no surrogate gradients)
|
||
2. **Binary** — works with binary weights and/or activations
|
||
3. **RL** — can learn from reward signal instead of labels
|
||
4. **Online** — can update weights tick-by-tick, not batch-offline
|
||
5. **Depth** — practically trains 3+ layers
|
||
6. **Cost** — compute per weight update (relative: Low / Med / High)
|
||
7. **Finicky** — sensitivity to hyperparameters (Low / Med / High)
|
||
|
||
---
|
||
|
||
## Method 1: Contrastive Hebbian Learning (CHL) / Equilibrium Propagation
|
||
|
||
### How It Works
|
||
|
||
CHL runs two phases using an energy-based (Hopfield-like) network:
|
||
|
||
1. **Free phase**: Input clamped, hidden + output units relax to energy minimum → record activations
|
||
`s⁻`
|
||
2. **Clamped phase**: Input + output clamped to target, hidden units relax → record activations `s⁺`
|
||
3. **Weight update**: `ΔW_ij ∝ s⁺_i · s⁺_j − s⁻_i · s⁻_j`
|
||
|
||
This is purely local — each weight only needs the activations of its two endpoints in both phases.
|
||
No global error signal flows backward; activity differences drive learning.
|
||
|
||
**Equilibrium Propagation (Scellier & Bengio 2017)** is the modern formulation: the clamped phase is a
|
||
"nudged" version where the output is soft-clamped with a small `β` factor. In the limit β→0, EqProp
|
||
provably computes the same updates as backprop, but without passing gradients.
|
||
|
||
**Binary variant (Laydevant et al., CVPRW 2021)**: EqProp was directly applied to dynamical binary
|
||
neural networks. The free/clamped phases remain unchanged; binary activations are compatible because
|
||
the update rule only compares activation products, not derivatives.
|
||
|
||
### Can It Work Without Supervised Targets?
|
||
|
||
**Critical limitation**: The clamped phase requires knowing the *target output*. In standard CHL/EqProp
|
||
the output neurons are pushed toward a labeled target. To use reward instead of labels, you need a
|
||
"nudging" signal at the output layer.
|
||
|
||
**Viable workaround for RL**: Instead of nudging toward a ground-truth label, nudge the output in the
|
||
direction that increases reward. If reward is scalar, define a "desired output shift":
|
||
- Positive reward: nudge current output activations → reinforce current pattern
|
||
- Negative reward: nudge away from current output
|
||
|
||
This converts CHL into a reward-modulated version. It is not a standard use case but has been explored
|
||
in neo-Hebbian RL literature. The key: the "nudge" is a *signal* (direction to move), not a gradient.
|
||
It fits the constraint.
|
||
|
||
### Scorecard
|
||
|
||
| Criterion | Score | Notes |
|
||
|-----------|-------|-------|
|
||
| No GD | ✅ | Pure activity differences, no derivatives |
|
||
| Binary | ✅ | Demonstrated in CVPRW 2021 on dynamic BNNs |
|
||
| RL | ⚠️ | Needs reward→nudge translation; not native but viable |
|
||
| Online | ⚠️ | Requires two settling phases per step; adds latency |
|
||
| Depth | ✅ | EqProp scales to multiple layers; credit is local |
|
||
| Cost | Med | Two full network relaxations per update |
|
||
| Finicky | High | β nudge strength, settling iterations, energy landscape |
|
||
|
||
### Fit Assessment
|
||
|
||
Strong theoretical foundation, but the two-phase settling requirement means you need the network to
|
||
relax to equilibrium twice per firing tick. In Robocode's 1ms budget this is expensive. The reward→
|
||
nudge translation is non-trivial to implement correctly. **Viable but complex.**
|
||
|
||
---
|
||
|
||
## Method 2: Forward-Forward Algorithm (FF) / Self-Contrastive FF
|
||
|
||
### How It Works
|
||
|
||
Hinton (2022): Replace one forward + one backward pass with *two forward passes*:
|
||
|
||
1. **Positive pass**: real/good data through the layer → maximize "goodness" G = Σ(y²)
|
||
2. **Negative pass**: corrupted/bad data → minimize goodness
|
||
3. **Layer-local update** for each layer independently: push goodness above threshold θ for positive,
|
||
below for negative. No signal crosses layer boundaries.
|
||
|
||
The Self-Contrastive Forward-Forward (SCFF, 2024/2025, Nature Communications) eliminates the label
|
||
requirement entirely: positive sample = [x_k, x_k] (same input concatenated), negative sample =
|
||
[x_k, x_n] (two different inputs). The network learns to distinguish same-from-different without
|
||
any labels.
|
||
|
||
### Goodness Function
|
||
|
||
`G(l) = (1/M) Σ_m y²_m(l)` — sum of squared activations in layer l over M neurons.
|
||
|
||
Learning rule: `ΔW ∝ (σ(G - θ) - target) · ∂G/∂W`
|
||
|
||
**Wait — does this have gradients?** Yes, the standard FF uses a sigmoid-of-goodness loss with
|
||
gradient descent on the goodness. However: (a) the update is strictly *local* to each layer, and
|
||
(b) ∂G/∂W = 2y · ∂y/∂W. For binary activations with sign() this derivative is zero/undefined.
|
||
|
||
**Binary compatibility issue**: The goodness gradient requires ∂y/∂W. With hard binary activations
|
||
(sign function), this derivative is zero everywhere except at the threshold, breaking the goodness
|
||
update. The STE would be required to push gradients through — which is explicitly disallowed.
|
||
|
||
**Alternative**: Replace the goodness gradient with a Hebbian rule — if G > θ for positive data,
|
||
do Hebbian update on active neurons; if G < θ for negative data, do anti-Hebbian. This is the
|
||
"spirit" of FF without the gradient. It loses convergence guarantees but preserves the structure.
|
||
|
||
### Adaptation for RL
|
||
|
||
Since SCFF already generates its own positive/negative pairs without labels, it operates as a
|
||
*representation learning* system. For Robocode: the FF layers would learn good binary representations;
|
||
then a reward-modulated Hebbian output layer would learn to map representations to predictions.
|
||
|
||
### Scorecard
|
||
|
||
| Criterion | Score | Notes |
|
||
|-----------|-------|-------|
|
||
| No GD | ⚠️ | Standard FF uses local GD on goodness; binary requires Hebbian approximation |
|
||
| Binary | ⚠️ | Needs Hebbian approximation of goodness; SCFF uses ReLU throughout |
|
||
| RL | ⚠️ | SCFF is self-supervised (representation learning); needs separate RL output layer |
|
||
| Online | ✅ | Layer-local, greedy, update after each sample |
|
||
| Depth | ✅ | By design — each layer trains independently |
|
||
| Cost | Low | Two forward passes per update, strictly local |
|
||
| Finicky | Med | Threshold θ, concatenation design choices |
|
||
|
||
### Fit Assessment
|
||
|
||
FF is excellent for building deep binary representations without labels. It cannot directly predict
|
||
position with reward alone — you'd hybridize it: FF middle layers for feature extraction, reward-
|
||
modulated Hebbian output layer for prediction. The goodness-gradient problem with hard binary
|
||
activations is the main obstacle. **Good as a representation sub-system; needs hybridization.**
|
||
|
||
---
|
||
|
||
## Method 3: Predictive Coding Networks (PCN)
|
||
|
||
### How It Works
|
||
|
||
Each layer l predicts the activation of layer l-1 from above using top-down weights. The prediction
|
||
error e_l = x_l − μ_l (actual minus predicted) propagates locally upward and downward.
|
||
|
||
Weight update: `ΔW ∝ e_l · x_{l-1}^T` (Hebbian on error × lower-layer activation)
|
||
|
||
This is **self-supervised by construction**: the "target" of each layer is the actual activation
|
||
of the layer below it. No external labels needed. With temporal data, each layer can predict the
|
||
next time frame — perfectly suited to 10-frame temporal windows.
|
||
|
||
**Active Predictive Coding (APC, 2022)**: Combines predictive coding with RL for robotic sparse
|
||
reward problems. The prediction error serves as an intrinsic reward signal. Demonstrated on robotic
|
||
control tasks with sparse rewards.
|
||
|
||
**PCN-TA (2025)**: Temporal amortization preserves latent states across frames, yielding 50% fewer
|
||
inference steps and 10% fewer weight updates vs. backprop.
|
||
|
||
### Gradient Status
|
||
|
||
Standard PCN inference minimizes free energy F = Σ e_l² by gradient descent on *activations*
|
||
(not weights). This is allowed — you're adjusting activity, not weights, via derivatives. Weight
|
||
updates are then Hebbian: `ΔW ∝ e · x`. No gradient of the loss w.r.t. weights is computed.
|
||
|
||
**Technically pure**: the weight update is `ΔW_l = η · e_l · x_{l-1}^T`, which is purely local and
|
||
Hebbian-like. The inference phase uses a form of gradient descent on activations, which is a
|
||
signal-flow operation (activations are signals), consistent with "backprop of signals allowed."
|
||
|
||
### Binary Compatibility
|
||
|
||
PCN with continuous activations works cleanly. Binary activations are harder: the inference phase
|
||
requires iterative adjustment of continuous latent variables before thresholding. A common approach:
|
||
keep latent variables continuous during inference, threshold only for output/communication. This is
|
||
the "semi-binary" regime used in neuromorphic PCN implementations.
|
||
|
||
### Fit for Temporal Prediction
|
||
|
||
Our 10 temporal frames map directly to a hierarchical PCN structure:
|
||
- Layer 1: current-frame features
|
||
- Layer 2: short-term dynamics (frames 0-2)
|
||
- Layer 3: trajectory trend (frames 0-9)
|
||
- Each layer predicts the layer below across time
|
||
|
||
The reward signal (miss distance) can modulate the prediction error magnitude at the output layer,
|
||
biasing learning toward accurate predictions.
|
||
|
||
### Scorecard
|
||
|
||
| Criterion | Score | Notes |
|
||
|-----------|-------|-------|
|
||
| No GD | ✅ | Weight update is Hebbian `e·x`; inference GD is on activations (signals) |
|
||
| Binary | ⚠️ | Latent variables need continuous inference; output can be binary |
|
||
| RL | ✅ | APC demonstrated; error = intrinsic reward; reward modulates output layer |
|
||
| Online | ✅ | PCN-TA designed for online per-frame learning |
|
||
| Depth | ✅ | Hierarchical by design; 3+ layers well studied |
|
||
| Cost | Med | Multiple inference iterations per step (reduced 50% in PCN-TA) |
|
||
| Finicky | Med | Inference step count, precision/ratio of prediction error weights |
|
||
|
||
### Fit Assessment
|
||
|
||
**Excellent match for our temporal structure.** The self-supervised nature, temporal hierarchy,
|
||
and local Hebbian weight updates align with every constraint. Main cost: inference iterations.
|
||
PCN-TA reduces this significantly. The semi-binary regime (continuous latent, binary output) is a
|
||
realistic implementation path.
|
||
|
||
---
|
||
|
||
## Method 4: Target Propagation / Difference Target Propagation (DTP)
|
||
|
||
### How It Works
|
||
|
||
Instead of propagating gradients, TP learns an *inverse model* at each layer: a backward network
|
||
g_l that maps activations of layer l+1 back to layer l. The target for layer l is:
|
||
|
||
`h_l^target = g_l(h_{l+1}^target)`
|
||
|
||
**Difference Target Propagation (DTP)**: corrects for imperfect inverses by using:
|
||
|
||
`h_l^target = h_l + g_l(h_{l+1}^target) − g_l(h_{l+1})`
|
||
|
||
The target for the output layer *still requires a supervised target* — this is the fundamental
|
||
limitation. DTP cannot generate its own target; it only propagates given targets layer-by-layer.
|
||
|
||
### Does It Need Supervised Output?
|
||
|
||
**Yes, inherently.** The output-layer target must come from somewhere. In RL, you could define the
|
||
output target as "the output that would have hit the enemy" — but this requires knowing the correct
|
||
output, which is equivalent to supervised learning. The reward tells you *how much* you were off,
|
||
not *where* you should have aimed.
|
||
|
||
Additionally: "the target for the final hidden layer is determined by a formula which relies on
|
||
gradient descent" (from DTP theory paper). This means gradient descent leaks in even in DTP.
|
||
|
||
### Binary Compatibility
|
||
|
||
DTP can work with binary units (Lee et al. showed this: "it can be applied even when units exchange
|
||
stochastic bits rather than real numbers"). But non-invertible binary networks cause reconstruction
|
||
errors to interfere with target propagation updates.
|
||
|
||
### Scorecard
|
||
|
||
| Criterion | Score | Notes |
|
||
|-----------|-------|-------|
|
||
| No GD | ❌ | Output-layer target computation involves GD; inverse model also trained with GD |
|
||
| Binary | ⚠️ | Stochastic binary supported but reconstruction error problem |
|
||
| RL | ❌ | Requires supervised output target; reward alone insufficient |
|
||
| Online | ⚠️ | Possible but inverse models add computational overhead |
|
||
| Depth | ✅ | Designed for depth; that's the whole point |
|
||
| Cost | High | Two networks (forward + inverse) per layer |
|
||
| Finicky | High | Inverse model training, reconstruction accuracy |
|
||
|
||
### Fit Assessment
|
||
|
||
**Not suitable.** Despite claiming "no backprop," DTP requires gradient descent for the inverse
|
||
models and for the output target. Does not fit RL-only constraint. Skip.
|
||
|
||
---
|
||
|
||
## Method 5: Reward-Modulated STDP (R-STDP) / Three-Factor Learning / E-prop
|
||
|
||
### How It Works
|
||
|
||
**Three-factor rule**: `ΔW_ij = η · e_ij · r(t)` where:
|
||
- `e_ij` = eligibility trace: local pre×post spike correlation (STDP-like)
|
||
- `r(t)` = global reward signal (scalar, broadcast to all synapses)
|
||
|
||
Eligibility trace: `de_ij/dt = −e_ij/τ + STDP(t_pre, t_post)` — decaying memory of recent spike
|
||
correlations.
|
||
|
||
This handles temporal credit assignment: the trace remembers which synapses fired recently when
|
||
the (delayed) reward arrives.
|
||
|
||
**E-prop (Bellec et al., Nature Comms 2020)**: Extends three-factor learning to deep and recurrent
|
||
SNNs. Each layer l has an eligibility trace that accounts for the temporal dynamics of the LIF neuron.
|
||
E-prop with reward-based RL (r-e-prop) was demonstrated winning Atari games.
|
||
|
||
**Critical depth finding**: "e-prop implemented in a single-layer recurrent SNN consistently
|
||
outperforms a multi-layer variant." Adding more layers hurts because the eligibility trace at layer l
|
||
only captures local pre/post correlations — it cannot account for multi-layer credit assignment
|
||
without approximate gradients. E-prop for deep feedforward networks approximates backprop, not a
|
||
clean departure from it.
|
||
|
||
### Binary Compatibility
|
||
|
||
SNNs fire binary spikes by definition (0 or 1), so R-STDP is natively binary-compatible. Weights
|
||
can also be binarized. The eligibility trace is continuous (a real-valued running average), but that
|
||
is internal state, not a gradient propagated through the network.
|
||
|
||
### Credit Assignment Through Depth
|
||
|
||
This is the weak point. R-STDP solves temporal credit assignment (delay between action and reward)
|
||
but not *spatial* credit assignment (which layer contributed to the good/bad output). In a 3-layer
|
||
network with R-STDP applied uniformly, all layers update on the same reward signal regardless of
|
||
their contribution. This works empirically in shallow networks but degrades with depth.
|
||
|
||
### Scorecard
|
||
|
||
| Criterion | Score | Notes |
|
||
|-----------|-------|-------|
|
||
| No GD | ✅ | Local STDP trace × reward; no derivatives anywhere |
|
||
| Binary | ✅ | Native binary (spikes); weights can also be binary |
|
||
| RL | ✅ | Reward is the direct learning signal; no labels |
|
||
| Online | ✅ | Per-spike updates; inherently online |
|
||
| Depth | ⚠️ | Works but degrades with depth; single-layer LSNN beats multi-layer |
|
||
| Cost | Low | Eligibility trace update O(n_synapses) per tick |
|
||
| Finicky | Med | τ_eligibility, reward scaling, STDP window |
|
||
|
||
### Fit Assessment
|
||
|
||
**Best fit for shallow networks.** For a 2-layer binary SNN (one hidden layer), R-STDP is the
|
||
cleanest solution: no gradients, native binary, direct reward learning, very cheap. For 3+ layers,
|
||
credit assignment degrades. Consider using it for the *output layer* in combination with another
|
||
method for hidden layers. **Highly recommended for shallow or hybrid architectures.**
|
||
|
||
---
|
||
|
||
## Method 6: Node Perturbation / Weight Perturbation
|
||
|
||
### How It Works
|
||
|
||
**Node perturbation (NP)**: Add noise ξ to each layer's activations, measure reward R, update:
|
||
`ΔW_ij ∝ ξ_j · (R − R̄)` where R̄ is baseline reward.
|
||
|
||
**Weight perturbation (WP)**: Add noise directly to weights: `ΔW_ij ∝ ε_ij · (R − R̄)`
|
||
|
||
**Key insight**: This estimates the gradient without computing it. It is a Monte Carlo gradient
|
||
estimate — unbiased but high variance.
|
||
|
||
**Decorrelated NP (DNP, 2023)**: Applying input decorrelation at each layer "dramatically improves
|
||
convergence by orders of magnitude." Makes NP practical for multi-layer networks.
|
||
|
||
**Variance problem**: The SNR of the NP gradient estimate scales as `1/(N·σ_noise²)` where N =
|
||
number of parameters. For a 690-bit → 256-neuron hidden layer, N ≈ 176,640 weights. The variance
|
||
explosion makes learning essentially random without variance reduction.
|
||
|
||
**For binary networks specifically**: Weight perturbation on binary weights means randomly flipping
|
||
bits and measuring reward change. This is exactly the "simulated annealing on neural net weights"
|
||
approach. It is well-defined and requires no derivatives. The update becomes:
|
||
|
||
`if flip W_ij causes ΔR > 0: keep flip; else: revert (with probability based on ΔR)`
|
||
|
||
This is pure stochastic search — not gradient estimation. Works cleanly with binary weights.
|
||
|
||
### Scorecard
|
||
|
||
| Criterion | Score | Notes |
|
||
|-----------|-------|-------|
|
||
| No GD | ✅ | Statistical gradient estimation; no derivatives |
|
||
| Binary | ✅ | WP maps naturally to bit-flip search |
|
||
| RL | ✅ | Reward is the direct signal |
|
||
| Online | ✅ | Update per trial; eligible for online use |
|
||
| Depth | ⚠️ | DNP scales to 3-9 layers; variance still significant |
|
||
| Cost | High | Many perturbation samples needed for convergence |
|
||
| Finicky | Med | Perturbation magnitude σ, baseline estimator |
|
||
|
||
### Fit Assessment
|
||
|
||
**The cleanest in principle, but expensive.** Weight perturbation on a binary network is essentially
|
||
a guided random walk through {-1,+1}^N. With N ≈ 176K weights in a hidden layer, random perturbation
|
||
converges very slowly. Works best when: (a) network is small, (b) few samples needed per weight. For
|
||
the *output layer only* (small, ~64 weights to prediction neurons), this is practical. **Suitable as
|
||
output-layer fine-tuner; too slow as sole learning method for deep networks.**
|
||
|
||
---
|
||
|
||
## Method 7: InfoMax / Information Maximization
|
||
|
||
### How It Works
|
||
|
||
Linsker (1988), Bell & Sejnowski (1995): Maximize mutual information I(X; Y) between layer input X
|
||
and output Y, subject to noise constraints.
|
||
|
||
Local learning rule (continuous case): `ΔW_ij ∝ [φ'(y_i)/φ(y_i)] · x_j − W_ij^{-T}`
|
||
|
||
For the binary/ICA case (Bell-Sejnowski): involves a nonlinear function of output activation and
|
||
the input pattern. In practice, this maximizes entropy of the output distribution, preventing the
|
||
network from collapsing all outputs to zero or all-ones.
|
||
|
||
### Binary Compatibility
|
||
|
||
InfoMax with binary activations maximizes entropy H(Y) — binary outputs should have p(y=1) ≈ 0.5
|
||
across data. The update rule involves the "score function" of the output distribution, which for
|
||
binary units is well-defined without derivatives through the activation function itself.
|
||
|
||
### Self-Supervised Nature
|
||
|
||
InfoMax requires *no labels*. It maximizes information preserved from input to output. This makes
|
||
it useful for *representation learning in hidden layers* — building diverse, non-collapsed binary
|
||
features. It does not directly learn to predict enemy position.
|
||
|
||
### Limitations for RL
|
||
|
||
InfoMax is a representation learning method. It maximizes informativeness of intermediate features
|
||
but has no notion of task reward. You cannot directly optimize miss distance with InfoMax alone.
|
||
It would serve as a **pretraining / regularization layer** to prevent dead neurons and maintain
|
||
diverse binary representations.
|
||
|
||
### Scorecard
|
||
|
||
| Criterion | Score | Notes |
|
||
|-----------|-------|-------|
|
||
| No GD | ✅ | Local Hebbian-like rule with anti-Hebbian lateral term |
|
||
| Binary | ✅ | Entropy maximization is well-defined for binary units |
|
||
| RL | ❌ | Pure representation learning; no reward optimization |
|
||
| Online | ✅ | Fully online, local update |
|
||
| Depth | ✅ | Applied layer-by-layer independently |
|
||
| Cost | Low | Single forward pass per update |
|
||
| Finicky | Low | Few parameters; entropy maximization is self-stabilizing |
|
||
|
||
### Fit Assessment
|
||
|
||
**Excellent regularizer / auxiliary learning rule.** Run InfoMax as an auxiliary update on hidden
|
||
layers to prevent representational collapse while the output layer learns from reward. Very cheap.
|
||
**Use as a complement to primary method, not standalone.**
|
||
|
||
---
|
||
|
||
## Method 8: Evolutionary / Genetic Algorithms for Binary BNNs
|
||
|
||
### How It Works
|
||
|
||
Maintain a population of binary weight matrices. Each individual is a complete set of weights
|
||
`{W_l}` with values in {-1, +1}. Evolution:
|
||
|
||
1. Evaluate fitness (reward = hit rate, negative miss distance)
|
||
2. Select top-K individuals (tournament or truncation selection)
|
||
3. Crossover: combine weight sub-matrices from two parents
|
||
4. Mutation: randomly flip bits with probability p_mut
|
||
5. Repeat
|
||
|
||
Binary networks are particularly GA-friendly because:
|
||
- The search space is discrete and finite
|
||
- Crossover has a clean interpretation (taking different weight blocks)
|
||
- No gradient computation at all
|
||
- No issues with binary activations (trivially compatible)
|
||
|
||
### Population-Based Training During a Battle
|
||
|
||
A battle has ~1000+ ticks. With a population of P=10-20 individuals, each tick evaluates the
|
||
current "best" policy, and every N ticks (after a reward arrives from circular wave) the population
|
||
updates. This is a form of online evolution.
|
||
|
||
**Key problem**: Each individual needs to be evaluated on the *same input* to compare fitnesses.
|
||
During a battle, inputs change each tick — you cannot hold them constant while evaluating 20 models.
|
||
Solutions:
|
||
- Maintain an experience replay buffer; evaluate candidates on stored states
|
||
- Use the single "best" candidate in battle, explore with perturbations (= weight perturbation)
|
||
- Use island model: different individuals fight in different battles
|
||
|
||
### Practical Issues
|
||
|
||
- Population of 20 requires 20× memory for weights
|
||
- Fitness estimates from sequential battle experience are noisy (different opponents, different
|
||
positions each eval)
|
||
- Crossover between weights that serve different layers may break learned structure
|
||
- Convergence within a single 1000-tick battle is unlikely for deep networks
|
||
|
||
### Scorecard
|
||
|
||
| Criterion | Score | Notes |
|
||
|-----------|-------|-------|
|
||
| No GD | ✅ | Pure fitness-based search |
|
||
| Binary | ✅ | Perfect; binary is ideal for genetic operators |
|
||
| RL | ✅ | Fitness = battle reward; no labels |
|
||
| Online | ❌ | Population requires multiple evaluations; impossible in a single live battle |
|
||
| Depth | ✅ | Architecture-agnostic |
|
||
| Cost | High | O(P × network_size) per generation |
|
||
| Finicky | Med | Population size, mutation rate, selection pressure |
|
||
|
||
### Fit Assessment
|
||
|
||
**Not suitable for within-battle online learning.** The population-based evaluation requirement
|
||
breaks against the single-agent, real-time constraint. Best applied *between battles* (offline
|
||
evolution over many battles). If we're restricted to online learning, this is out. **Viable only
|
||
as inter-battle meta-learning.**
|
||
|
||
---
|
||
|
||
## Method 9: Feedback Alignment (FA) / Direct Feedback Alignment (DFA)
|
||
|
||
### How It Works
|
||
|
||
**Feedback Alignment (Lillicrap et al., 2016)**: Replace backprop's transposed weight matrices
|
||
W^T in the backward pass with fixed random matrices B. The claim: the forward weights align to
|
||
B during training, so the random backward signal still provides useful credit assignment.
|
||
|
||
**Direct Feedback Alignment (DFA)**: Each layer receives credit directly from the output error
|
||
multiplied by a random matrix B_l (layer-specific). No sequential backward pass needed.
|
||
|
||
### The Gradient Problem
|
||
|
||
FA and DFA fundamentally still compute a gradient of the loss at the *output layer* and propagate
|
||
that signal (or a random-projection of it) backward. The output gradient `∂L/∂y` is a derivative.
|
||
This violates the "no gradient descent" constraint. FA avoids the *weight transport* problem but
|
||
does not avoid derivatives entirely.
|
||
|
||
**Is the "signal" interpretation valid?** You could argue: the output error `(y - target)` is a
|
||
signal (not a derivative). If you replace this with a reward-modulated output unit error, you
|
||
avoid computing a derivative. But this is a stretch — in practice, FA implementations use `∂L/∂y`.
|
||
|
||
**Binary compatibility**: FA fails on binary networks because the backward signal through hard
|
||
thresholds is zero everywhere (the same STE problem as standard backprop). Papers report "FA
|
||
fails in deep networks, convolutional layers, or architectures with bottlenecks."
|
||
|
||
### Scorecard
|
||
|
||
| Criterion | Score | Notes |
|
||
|-----------|-------|-------|
|
||
| No GD | ❌ | Output gradient still computed; only backward weight symmetry is removed |
|
||
| Binary | ❌ | Fails with binary activations without STE |
|
||
| RL | ⚠️ | Could replace loss gradient with reward signal, but then it's R-STDP |
|
||
| Online | ✅ | Online update possible |
|
||
| Depth | ⚠️ | DFA works at depth but with performance degradation |
|
||
| Cost | Low | Same cost as forward pass + random projection |
|
||
| Finicky | Low | Random fixed matrices; no tuning needed |
|
||
|
||
### Fit Assessment
|
||
|
||
**Not suitable as stated.** FA/DFA require output-layer gradients and fail with binary activations.
|
||
If you replace the output gradient with a reward signal and the backward pass with fixed random
|
||
projections, you get something closer to random reward feedback — a variant of node perturbation,
|
||
which is covered above. **Skip in favor of R-STDP or node perturbation.**
|
||
|
||
---
|
||
|
||
## Method 10: Additional Methods Found
|
||
|
||
### 10a: Equilibrium Propagation + CHL as a Unified Framework
|
||
|
||
Recent work (arXiv 2206.02629) shows that CHL, EqProp, and predictive coding all converge to the
|
||
same limit at infinitesimal inference steps — they are the same algorithm in different dynamical
|
||
regimes. This means choosing among them is primarily about implementation tradeoffs (settling speed,
|
||
binary compatibility) not fundamental differences.
|
||
|
||
### 10b: Counter-Current Learning (CCL, 2024)
|
||
|
||
A dual-network approach: one network runs forward (inference), the other runs backward (teaching).
|
||
The backward network generates targets for the forward network without computing gradients. Used
|
||
as a biologically plausible alternative to backprop. Requires paired network architecture, doubling
|
||
memory. Interesting but adds complexity. **Skip for now.**
|
||
|
||
### 10c: E-prop with Reward (r-e-prop)
|
||
|
||
Specifically the reward-based variant of e-prop demonstrated on Atari. This is three-factor learning
|
||
with eligibility traces designed for temporal credit assignment, formally shown to approximate
|
||
policy gradients. Does use gradient approximations internally (via the eligibility trace derivation),
|
||
but the weight update itself is `ΔW = e_ij · r(t)` — a local multiplication.
|
||
|
||
**Whether this "counts" as gradient descent**: The eligibility trace `e_ij` is derived from the
|
||
neuron's dynamical equations and approximates `∂h_l/∂W`. This is a gradient of *activations*
|
||
w.r.t. weights (signal propagation), not a gradient of the *loss*. It satisfies "backprop of
|
||
signals" while avoiding "gradient descent on loss." **This is the cleanest theoretical fit.**
|
||
|
||
### 10d: Noise-Based Reward-Modulated Learning (2025, arXiv 2503.23972)
|
||
|
||
Explicitly designed for spiking/binary units: natural noise in binary neurons drives stochastic
|
||
exploration, reward signal modulates which noise patterns are reinforced. Weight update:
|
||
`ΔW ∝ ξ · r(t)` where ξ is inherent neural noise (spike timing jitter). Online, local, binary-native.
|
||
**Very relevant — this is weight perturbation where perturbation = natural spike noise.**
|
||
|
||
### 10e: Self-Contrastive Forward-Forward (SCFF, Nature Comms 2025)
|
||
|
||
Published result: MNIST 98.7%, CIFAR-10 80.75%, STL-10 77.3% without any labels. Layer-local
|
||
goodness function. Greedy layer-wise training. The best current result for label-free local learning
|
||
in deep networks. Needs Hebbian approximation for hard binary activations.
|
||
|
||
---
|
||
|
||
## Comparative Scorecard Summary
|
||
|
||
| Method | No GD | Binary | RL | Online | Depth | Cost | Finicky | **Score** |
|
||
|--------|-------|--------|-----|--------|-------|------|---------|-----------|
|
||
| CHL / EqProp | ✅ | ✅ | ⚠️ | ⚠️ | ✅ | Med | High | **6/9** |
|
||
| FF / SCFF | ⚠️ | ⚠️ | ⚠️ | ✅ | ✅ | Low | Med | **5.5/9** |
|
||
| Predictive Coding | ✅ | ⚠️ | ✅ | ✅ | ✅ | Med | Med | **7.5/9** |
|
||
| Target Propagation | ❌ | ⚠️ | ❌ | ⚠️ | ✅ | High | High | **2/9** |
|
||
| R-STDP / 3-factor | ✅ | ✅ | ✅ | ✅ | ⚠️ | Low | Med | **7.5/9** |
|
||
| Node/Weight Perturb | ✅ | ✅ | ✅ | ✅ | ⚠️ | High | Med | **6.5/9** |
|
||
| InfoMax | ✅ | ✅ | ❌ | ✅ | ✅ | Low | Low | **5/9** (aux only) |
|
||
| Genetic / EA | ✅ | ✅ | ✅ | ❌ | ✅ | High | Med | **4/9** |
|
||
| Feedback Alignment | ❌ | ❌ | ⚠️ | ✅ | ⚠️ | Low | Low | **2/9** |
|
||
| r-e-prop | ✅ | ✅ | ✅ | ✅ | ⚠️ | Low | Med | **7.5/9** |
|
||
|
||
---
|
||
|
||
## TOP 3 RECOMMENDATIONS
|
||
|
||
### Rank 1: R-STDP / Three-Factor Learning (Shallow + Output Layer)
|
||
|
||
**Why #1**: Every criterion met cleanly. No derivatives anywhere. Binary spikes are native. Reward
|
||
is the direct learning signal. Per-tick online updates. Cheap (O(n_synapses) per tick).
|
||
|
||
**The constraint**: poor depth scaling. Solution: limit depth or use shallow BNN with this method.
|
||
|
||
**Architecture sketch** (2-layer BNN):
|
||
|
||
```
|
||
Input (690 bits)
|
||
│
|
||
▼ W1 ∈ {-1,+1}^{690×128}
|
||
Hidden Layer (128 binary neurons)
|
||
│ ↑ InfoMax auxiliary update (prevents collapse)
|
||
▼ W2 ∈ {-1,+1}^{128×64}
|
||
Output Layer (64 binary neurons → [predX, predY] via dot-product readout)
|
||
│
|
||
▼
|
||
Circular wave feedback → miss distance → reward r(t)
|
||
│
|
||
└──→ Eligibility trace e_ij = e_ij · τ_decay + STDP(pre_j, post_i)
|
||
Weight update: ΔW_ij = η · e_ij · r(t) [for all layers uniformly]
|
||
```
|
||
|
||
**Signal flow for depth**: Use a **reward-gated reverse signal** — after reward arrives, compute
|
||
output layer error signal (not gradient, just "output was wrong by X direction") and multiply by a
|
||
random fixed matrix B to project to hidden layer. This gives hidden layer a noisy credit signal.
|
||
This is the "feedback alignment without gradients" version where the output error is a reward signal
|
||
(binary direction: aim left or right), not a loss derivative.
|
||
|
||
**Hypers**: τ_eligibility ≈ 10 ticks (covers bullet travel time), η ≈ 0.01, STDP window ≈ 3 ticks.
|
||
|
||
---
|
||
|
||
### Rank 2: Predictive Coding with Reward-Gated Output
|
||
|
||
**Why #2**: Best theoretical fit to our temporal 10-frame structure. Self-supervised hidden layers
|
||
(no labels needed), reward modulates only the output layer. Proven online on edge robots (IROS 2025).
|
||
|
||
**Architecture sketch** (3-layer PCN with temporal hierarchy):
|
||
|
||
```
|
||
Frame buffer: [f0..f9] (10 temporal frames × 69 fields = 690 bits)
|
||
|
||
Layer 3 (top): trajectory-level representation, 64 neurons
|
||
↕ predicts/corrects ↕ W3 = Hebbian update: e3 · h2^T
|
||
Layer 2: motion dynamics, 128 neurons
|
||
↕ predicts/corrects ↕ W2 = e2 · h1^T
|
||
Layer 1: per-frame features, 256 neurons
|
||
↕ corrects ↕ W1 = e1 · x^T
|
||
Input: 690-bit frame vector
|
||
|
||
Signal flow:
|
||
Forward: h_l = σ(W_l · h_{l-1}) [inference]
|
||
Backward: μ_l = W_{l+1}^T · h_{l+1} [top-down prediction, W^T of SAME weights]
|
||
Error: e_l = h_l − μ_l [local prediction error]
|
||
Weight: ΔW_l = η_rep · e_l · h_{l-1}^T [Hebbian on error × input]
|
||
|
||
Reward integration:
|
||
Output layer e_out += η_reward · r(t) · sign(h_out − target_direction)
|
||
where target_direction is estimated from miss distance (left/right of predicted pos)
|
||
```
|
||
|
||
**Key**: the weight update `ΔW = e · h^T` is Hebbian, not gradient descent. The inference phase
|
||
(adjusting activations to minimize prediction error) uses signal flow, not backprop. Binarize outputs
|
||
only; keep latent variables as integers 0-255 (8-bit) for inference, threshold for communication.
|
||
|
||
**Hypers**: inference iterations ≈ 3-5 per tick (PCN-TA reduces this), η_rep ≈ 0.001, η_reward ≈ 0.05.
|
||
|
||
---
|
||
|
||
### Rank 3: Self-Contrastive FF (SCFF) for Deep Representations + R-STDP Output
|
||
|
||
**Why #3**: Best combination for a deeper (3+ layer) network when richer representations are needed.
|
||
SCFF trains hidden layers purely from self-supervised data (no labels, no reward). R-STDP on the
|
||
output layer uses reward directly. Two independent, specialized learning mechanisms.
|
||
|
||
**Architecture sketch**:
|
||
|
||
```
|
||
Input: 690-bit × 2 (concatenated for SCFF contrast)
|
||
|
||
Layer 1: 512 binary neurons
|
||
SCFF update: Hebbian-approx goodness rule
|
||
positive pair: [x_k, x_k] → increase goodness
|
||
negative pair: [x_k, x_n] → decrease goodness
|
||
ΔW ∝ (target_goodness − actual_goodness) · pre_activity [no STE — approximate]
|
||
|
||
Layer 2: 256 binary neurons
|
||
SCFF update: same as Layer 1
|
||
|
||
Layer 3: 128 binary neurons
|
||
SCFF update: same
|
||
|
||
Output: 32 binary neurons → [predX, predY] decoding
|
||
R-STDP update: ΔW ∝ e_ij · r(t) [three-factor, reward = −miss_distance]
|
||
Eligibility trace: STDP on output spikes × layer-3 spikes
|
||
```
|
||
|
||
**Why this works**: Hidden layers learn to preserve temporal motion patterns from the input buffer
|
||
(self-supervised: same-frame vs different-frame contrast captures motion coherence). Output layer
|
||
learns which patterns correlate with correct predictions (reward). The two learning rules are
|
||
orthogonal and can run simultaneously without interference.
|
||
|
||
**The SCFF-binary approximation**: Replace `∂goodness/∂W` with Hebbian on active neurons when
|
||
goodness exceeds/falls below threshold θ. Loses convergence guarantees but works empirically for
|
||
binary feature learning. θ ≈ 0.5 × expected_activation_rate.
|
||
|
||
**Hypers**: θ ≈ 0.4, τ_eligibility ≈ 10, η_scff ≈ 0.001 (slow), η_rstdp ≈ 0.01 (fast).
|
||
|
||
---
|
||
|
||
## Final Notes and Red Flags
|
||
|
||
### Things That Look Gradient-Free But Aren't
|
||
|
||
1. **Target Propagation**: Output target still needs gradient descent. ❌
|
||
2. **Feedback Alignment**: Random backward weights but still computes `∂L/∂y`. ❌
|
||
3. **Standard Forward-Forward**: Goodness gradient is a real derivative through relu. Needs approximation for binary. ⚠️
|
||
4. **E-prop (standard version)**: Eligibility trace approximates `∂h/∂W` — it is a gradient of
|
||
activations, which satisfies "backprop of signals" but is worth flagging. ✅ (barely)
|
||
|
||
### The Depth-vs-Purity Trade-off
|
||
|
||
There is a fundamental tension: pure local learning rules (InfoMax, R-STDP) have no depth credit
|
||
assignment. Any method that propagates credit through depth (CHL, EqProp, PCN, e-prop) uses some
|
||
form of signal backpropagation. The question is whether those signals are *gradients of a loss*
|
||
(forbidden) or *prediction errors / activity differences* (allowed as signals). All three
|
||
recommendations above fall in the "allowed" category by that interpretation.
|
||
|
||
### Practical Starting Point
|
||
|
||
Start with **Rank 1** (R-STDP, 2 layers). It is the fastest to implement, cheapest to run, and
|
||
cleanest theoretically. Add **InfoMax** as a free auxiliary update on the hidden layer to prevent
|
||
dead neurons. Measure performance. If representation quality is bottlenecking predictions, move to
|
||
**Rank 2** (PCN temporal hierarchy). Only move to **Rank 3** (SCFF+R-STDP) if deeper representations
|
||
are demonstrably needed.
|
||
|
||
---
|
||
|
||
## Sources
|
||
|
||
- [Contrastive Hebbian Learning with Random Feedback Weights](https://arxiv.org/pdf/1806.07406)
|
||
- [Two Tales of Single-Phase Contrastive Hebbian Learning](https://arxiv.org/pdf/2402.08573)
|
||
- [Equilibrium Propagation: Bridging the Gap between Energy-Based Models and Backpropagation](https://arxiv.org/html/1602.05179v5)
|
||
- [Training Dynamical Binary Neural Networks with Equilibrium Propagation (CVPRW 2021)](https://openaccess.thecvf.com/content/CVPR2021W/BiVision/papers/Laydevant_Training_Dynamical_Binary_Neural_Networks_With_Equilibrium_Propagation_CVPRW_2021_paper.pdf)
|
||
- [The Forward-Forward Algorithm: Some Preliminary Investigations (Hinton 2022)](https://www.cs.toronto.edu/~hinton/FFA13.pdf)
|
||
- [Self-Contrastive Forward-Forward Algorithm (Nature Comms 2025)](https://www.nature.com/articles/s41467-025-61037-0)
|
||
- [Self-Contrastive Forward-Forward Algorithm (arXiv 2409.11593)](https://arxiv.org/abs/2409.11593)
|
||
- [Efficient Online Learning with Predictive Coding Networks: Exploiting Temporal Correlations (PCN-TA)](https://arxiv.org/abs/2510.25993)
|
||
- [Introduction to Predictive Coding Networks for Machine Learning](https://arxiv.org/pdf/2506.06332)
|
||
- [Active Predicting Coding: Brain-Inspired RL for Sparse Reward Robotic Control](https://arxiv.org/pdf/2209.09174)
|
||
- [A Theoretical Framework for Target Propagation](https://arxiv.org/pdf/2006.14331)
|
||
- [Towards Scaling Difference Target Propagation by Learning Backprop Targets](https://proceedings.mlr.press/v162/ernoult22a/ernoult22a.pdf)
|
||
- [Three-factor learning in spiking neural networks (PMC 2025)](https://pmc.ncbi.nlm.nih.gov/articles/PMC12745983/)
|
||
- [A solution to the learning dilemma for recurrent networks of spiking neurons (e-prop, Nature Comms 2020)](https://www.nature.com/articles/s41467-020-17236-y)
|
||
- [Including STDP to eligibility propagation in multi-layer recurrent SNNs](https://arxiv.org/pdf/2201.07602)
|
||
- [BioLCNet: Reward-Modulated Locally Connected SNNs](https://arxiv.org/pdf/2109.05539)
|
||
- [On the Stability and Scalability of Node Perturbation Learning (NeurIPS 2022)](https://proceedings.neurips.cc/paper_files/paper/2022/file/cf38eb1549024cce4b3d2c1bb87a6c27-Paper-Conference.pdf)
|
||
- [Effective Learning with Node Perturbation in Deep Neural Networks](https://arxiv.org/html/2310.00965v3)
|
||
- [Noise-based reward-modulated learning (arXiv 2503.23972)](https://arxiv.org/pdf/2503.23972)
|
||
- [Weight versus Node Perturbation Learning (Phys. Rev. X 2023)](https://link.aps.org/doi/10.1103/PhysRevX.13.021006)
|
||
- [Infomax (Wikipedia)](https://en.wikipedia.org/wiki/Infomax)
|
||
- [Local Synaptic Learning Rules Suffice to Maximize Mutual Information](https://scite.ai/reports/local-synaptic-learning-rules-suffice-4kGY1R)
|
||
- [Are alternatives to backpropagation useful for training Binary Neural Networks? (ACM SAC 2023)](https://dl.acm.org/doi/10.1145/3555776.3577674)
|
||
- [Learning Without Feedback: Fixed Random Learning Signals for Deep Networks (PMC)](https://pmc.ncbi.nlm.nih.gov/articles/PMC7902857/)
|
||
- [Counter-Current Learning: A Biologically Plausible Dual Network Approach](https://arxiv.org/pdf/2409.19841)
|
||
- [Towards Biologically Plausible Computing: A Comprehensive Comparison](https://arxiv.org/pdf/2406.16062)
|