Files
SirRoboGarage/SNNBot_garage/research/binary-snn-learning.md
T

31 KiB
Raw Blame History

Binary SNN Learning Mechanisms: Research Survey

A systematic review of learning methods compatible with binary spiking neural networks and real-time robotic control. Focus: mechanisms without expensive backpropagation, suitability for neuromorphic hardware.

Date: 2026-09-13
Sources: Primary papers, arXiv surveys, official documentation


1. Hyperdimensional Computing (HDC) / Vector Symbolic Architectures (VSA)

What It Is

HDC is a computational framework using high-dimensional distributed representations (typically 10,000+ dimensions) where information is encoded as binary hypervectors. Operations rely on algebraic properties that exploit high-dimensional geometry.

Key Models:

  • Binary Spatter Codes
  • Holographic Reduced Representations (HRR)
  • Tensor Product Representations
  • Sparse Binary Distributed Representations
  • Multiply-Add-Permute (MAP)

Core Operations

  1. Binding (Multiplicative): Combine two hypervectors via XOR or element-wise operations to create a new vector orthogonal to both parents. v_combined = v1 ⊕ v2

  2. Bundling (Additive): Sum/average hypervectors to create superpositions. Preserves overlapping bit patterns for similarity retrieval.

  3. Permutation: Rotate/shift dimensions to encode sequences and order information. Can be random or structured.

Similarity Measure: Hamming distance or cosine similarity of binary vectors. Two vectors are considered "similar" if overlap ≥ threshold (typically 15-30% of bits).

How It Learns

  • Single-pass learning: Process each sample once; accumulate patterns in holographic memory through bundling
  • Classification: Encode input → bind with class-specific keys → measure similarity to learned class prototypes
  • No backpropagation required
  • Bidirectional retrieval: Can recall from partial/noisy inputs (content-addressable memory)

Binary Operations & Efficiency

All core operations use binary logic (XOR, AND, OR) or bit counting. No floating-point arithmetic. Amenable to:

  • FPGA implementation
  • In-memory computing (memristor arrays)
  • Neuromorphic chips with binary spike events

Computational Cost

  • Training: O(d) per sample (d = dimensionality, typically 10K)
  • Inference: O(d) per query
  • Memory: O(classes × d) bits
  • Latency: Single-pass; no iteration needed

Real-Time Control Suitability

Strong fit: Single-pass operation, fixed computational budget, sparse binary operations. Example: encode sensor state → bind with action → retrieve best matching action. No weight update overhead between timesteps.

Limitation: Large dimensionality (10K bits) requires efficient implementation. Good for high-level perception/decision; not ideal for pixel-level processing without preprocessing.

References


2. Liquid State Machines (LSM) / Echo State Networks (ESN)

What It Is

Reservoir computing model: fixed random recurrent network (the "liquid" or "reservoir") + trainable linear readout layer. The untrained reservoir performs rich temporal filtering; only the readout weights learn.

LSM: Spiking neural networks (biological realism, event-driven)
ESN: Rate-coded neurons (simpler math, similar principles)

How It Learns

  1. Initialization: Create random recurrent SNN with fixed weights (no learning rule here)
  2. Reservoir dynamics: Present input spike train; dynamics evolve, creating rich temporal signatures
  3. Readout training: Collect reservoir activations over time; train output layer via linear regression or simple Hebbian rule (one-pass or few-pass)

No backpropagation through reservoir. Temporal memory emerges from dynamics alone.

Binary Spikes & Efficiency

  • Input: spike train (binary events, sparse in time)
  • Reservoir: binary spike emissions (integrate-and-fire neurons)
  • Readout training: can use binary weights with thresholding or continuous approximations

For hardware: spike events are sparse, reducing energy. Training cost is low (linear regression on collected traces).

Computational Cost

  • Inference: O(N × T) where N = reservoir size, T = timesteps (simulate forward)
  • Training: O(N × T) data collection + O(N³) or O(N² × T) for readout fit (linear algebra)
  • Memory: O(N²) for recurrent weights + O(N_out × N) for readout

Reservoir size typically 100–10K neurons.

Real-Time Control Suitability

Strong fit for temporal tasks: Sequential decision-making, trajectory following, filtering noisy sensor data. Inherent memory without learning overhead.

Limitation: High online inference cost (must simulate reservoir forward for each timestep). Not ideal for ultra-low-latency single-decision tasks. Readout training requires data collection phase.

References


3. Random Weight Perturbation

What It Is

Gradient-free optimization: perturb weight randomly, measure effect on loss, update in direction of improvement. No backprop, no explicit gradient needed.

How It Learns

  1. Forward pass 1: Evaluate network with current weights, measure loss L₀
  2. Forward pass 2: Add small random noise to weights, re-evaluate, measure loss L₁
  3. Update: If L₁ < L₀, move weights in direction of noise with step size η; otherwise move opposite

Repeat for each weight or layer.

Binary Operations & Efficiency

  • Can work with binary weights: noise is small perturbation around quantization point; decision based on loss direction
  • Stochastic nature provides implicit regularization
  • No matrix ops (matrix multiplies still needed for forward passes)

Computational Cost

  • Training: 2 forward passes per update cycle; ~2× inference cost
  • Convergence: Slow compared to gradient-based methods (noisy gradient estimates); requires more iterations
  • Variance: High (noise-based updates); recent work on decorrelated perturbations improves this

Real-Time Control Suitability

Moderate fit: Online learning capability (can update weights during operation). No gradient computation overhead. Training inefficient but suitable for continual learning on robotic platforms where compute budget allows 2 forward passes per learning step.

Limitation: Slow convergence, high variance. Better for adjusting pre-trained weights than learning from scratch.

References


4. STDP with Binary Spikes

What It Is

Spike-Timing Dependent Plasticity: synaptic strength changes based on precise timing between pre- and post-neuron spikes. Biologically validated, event-driven (suitable for neuromorphic hardware).

Core Rule

  • Pre-before-post (causal): Pre-neuron fires, then post-neuron fires → weight increases (LTP)
  • Post-before-pre (acausal): Post-neuron fires, then pre-neuron fires → weight decreases (LTD)
  • Time window: Potentiation/depression peaks near ~20 ms, decays after

Mathematical form: ΔW = A₊ exp(-Δt/τ₊) if Δt > 0 (pre before post), or -A₋ exp(Δt/τ₋) if Δt < 0

Binary Spikes & Challenges

Classic STDP works with graded synaptic weights (continuous [0,1] or [-1,1]). With binary weights, the challenge arises: discrete jumps between high and low states lose memory stability.

Solution in literature: Use stochastic binary synapses

  • Synaptic strength = transition probability between binary states
  • Cumulative distribution function (CDF) of weight probability evolves sigmodally with LTP/LTD trials
  • Can be realized with paired memristive devices

How It Learns

  1. Initialize: Binary weights, probabilistic state
  2. Each spike pair: Update probability CDF based on timing
  3. Plasticity window: Exponential decay of learning signal with time
  4. Stabilization: Hebbian learning balances growth; homeostasis prevents runaway potentiation

No explicit "training phase"; learning occurs online during task execution.

Computational Cost

  • Inference: O(1) per spike event (check timing, update state)
  • Learning: O(1) per spike pair (update probability)
  • Memory: O(N²) for synaptic weights + small overhead for stochastic state

Extremely efficient for neuromorphic platforms where spikes are hardware events.

Real-Time Control Suitability

Excellent fit: True online learning during closed-loop control. No batch processing. Sparse spike events → low power. Time constants tuned to behavioral timescales (100s of ms to seconds).

Limitation: Complex parameter tuning (time constants, learning rates). Requires stable initial random synapses. Convergence is slow; better for fine-tuning than bootstrap learning.

References


5. Evolutionary Strategies for Neural Networks

What It Is

Population-based black-box optimization: maintain population of candidate weight vectors, perturb each, evaluate fitness (e.g., task reward), select/recombine best performers. No gradients needed; rewards only feedback.

OpenAI ES (2017): Scaled to train vision + control networks with distributed evolution on thousands of cores.

How It Learns

  1. Initialize: Population of N weight vectors (e.g., N=100–10K)
  2. Perturbation: Add Gaussian noise to each candidate: w_i = w_base + σ × noise_i
  3. Evaluation: Run task with each w_i, collect scalar reward R_i
  4. Selection: Estimate gradient ∝ E[R_i × noise_i]; update base weights
  5. Repeat: Next generation of population

No explicit backprop; reward signal is scalar (e.g., task score, survival time).

Binary Operations & Efficiency

  • Works with any weight representation (continuous, binary, mixed)
  • For binary: perturbations flip bits stochastically; keep if reward improves
  • Natural fit with binary SNNs: reward = task completion, no gradient flow needed

Computational Cost

  • Training: N forward simulations per generation (highly parallelizable)
  • Convergence: Slower than gradient-based (fewer bits of gradient info per eval), but parallelizable
  • Memory: O(N × W) for population (W = total weights); population size trades off diversity vs. cost

Typical: 100–1000 population members, 1000s of generations.

Real-Time Control Suitability

Good fit for:

  • Sim-to-real transfer (evolve in sim, deploy on robot)
  • Evolving network topology + weights (neuroevolution)
  • Multi-objective optimization (Pareto evolution for speed + accuracy)
  • Sparse rewards (evolution is robust to noise)

Limitation: Inherent lag (must wait for population evaluation before update). Not suited for online single-step learning during deployment. Better for offline training.

References


6. BCM Theory (Bienenstock-Cooper-Munro)

What It Is

Sliding-threshold Hebbian learning rule: potentiation and depression depend on whether postsynaptic activity exceeds a dynamically adapting threshold. Biologically validated; explains selectivity in visual cortex.

Core Rule

ΔW = η × y × (y - θ) × x

Where:

  • y = postsynaptic activity (firing rate)
  • x = presynaptic activity
  • θ = sliding threshold (adapts based on recent y statistics)
  • η = learning rate

Interpretation:

  • If y > θ: Hebbian potentiation (ΔW > 0)
  • If y < θ: Anti-Hebbian depression (ΔW < 0)
  • θ adjusts so that roughly half of postsynaptic events are above/below threshold

How It Learns

  1. Feedforward input: Afferent spike trains x
  2. Postsynaptic response: Integrate-and-fire or rate-coded y
  3. Threshold estimation: θ = E[y²]/E[y] (second moment / first moment) or moving average
  4. Weight update: Apply BCM rule based on current timing
  5. Homeostasis: Threshold self-adjusts; network finds balanced state

No explicit error signal; unsupervised. Learns feature selectivity (neurons develop preference for specific input patterns).

Binary Spikes & Efficiency

  • Works with spike counts (integrate over small window) rather than instantaneous spikes
  • Threshold can be binary decision: is spike rate above/below running average?
  • Simple to implement on neuromorphic hardware (local computation, homeostatic negative feedback)

Computational Cost

  • Inference: O(1) per spike (increment counter)
  • Learning: O(1) per spike (update weight based on threshold comparison)
  • Memory: O(N²) weights + O(N) threshold estimates

Minimal overhead; suitable for online learning.

Real-Time Control Suitability

Good fit: Self-organizing layers for feature extraction. No labeled data required. Scales to high-dimensional inputs. Natural fit with recurrent SNNs.

Limitation: Unsupervised (doesn't directly optimize task performance). Requires careful tuning of θ dynamics to avoid instability. Often used as unsupervised preprocessor, not end-to-end control.

References


7. Competitive Learning / Winner-Take-All Networks

What It Is

Unsupervised clustering: neurons compete to respond to input. Only "winner" (neuron with strongest response) activates strongly; losers silenced via lateral inhibition. Weights updated only for winner.

Algorithms: Self-Organizing Maps (Kohonen), Learning Vector Quantization (LVQ), Neural Gas, Adaptive Resonance Theory (ART)

How It Learns

  1. Input presentation: Sensory x presented to all neurons
  2. Competition: Each neuron computes activation a_i = sim(w_i, x) (e.g., dot product, Euclidean)
  3. Winner selection: i* = argmax(a_i)
  4. Lateral inhibition: Winner fires strongly; others suppressed via inhibitory connections
  5. Learning: Update winner weights toward input: w_i* ← w_i* + η(x - w_i*); others unchanged
  6. Repeat: Next input, new winner possibly emerges

Result: neurons self-organize to cluster input space. Similar inputs activate same winner (topological map).

Binary Operations & Efficiency

  • Similarity metric can be Hamming distance (for binary vectors) or binary dot product
  • Winner selection: simple argmax (can use spiking threshold)
  • Weight updates: Hebbian (increment on coincidence) or anti-Hebbian (decrement on mismatch)

Computational Cost

  • Inference: O(N) per input (compute similarity to all N prototypes)
  • Learning: O(1) per winner update (only update winner, not full network)
  • Memory: O(N × D) for prototype weights (N clusters, D dimensions)

Scales linearly with cluster count; sparse updates (only winner).

Real-Time Control Suitability

Good fit:

  • Online clustering of sensor inputs (e.g., ball position discretization for aiming)
  • Basis function learning (prototypes become features for downstream layer)
  • Low-latency inference (single argmax query)

Limitation: Cluster centers drift if input distribution non-stationary. Sensitive to initial conditions and learning rate. Requires rebalancing to prevent dead neurons. Better for stable environments than adaptive/adversarial settings.

References


8. Sparse Distributed Representations (SDR) — Numenta HTM

What It Is

Binary encoding scheme inspired by cortex: information encoded as sparse binary vector (e.g., 2048 bits, ~40 active). Similarity = overlap; sparse codes enable simultaneous representation of multiple items without interference.

Core principle: Learned associations are stored implicitly in sparse overlaps, not explicit weights.

How It Learns

HTM Spatial Pooler (online unsupervised):

  1. Input encoding: Raw data (e.g., sensor reading) → SDR (sparse binary vector)
  2. Competitive Hebbian: Columns compete; active columns increment weight to active input bits, inhibited columns decrement
  3. Homeostasis: Learning rates self-adjust to maintain target sparsity (e.g., 2% active)
  4. Result: Learns distributed sparse codes that compress input space

HTM Temporal Memory (sequential learning):

  • Adds temporal context: cells within column compete; prediction reinforces expected active cells
  • Learns state machine implicitly; transitions are sparse activations

No backprop; purely local rules.

Binary Operations & Efficiency

  • All operations on binary vectors: overlap (bit AND), population coding (multiple bits per concept)
  • Similarity metric: Hamming distance / Tanimoto coefficient
  • No floating-point; bit counting operations

Computational Cost

  • Inference: O(bits) per input encoding + O(columns × bits) for pooling
  • Training: Online, O(active_bits) updates per input
  • Memory: O(columns × input_bits) for connection matrix; sparse (only active connections stored)

HTM systems typically 2048–65K bit vectors; 10s of ms per inference on CPU.

Real-Time Control Suitability

Good fit:

  • Hierarchical temporal prediction (anticipate ball trajectory)
  • Anomaly detection (identify novel states)
  • Online learning from streaming data
  • Energy efficiency (sparse bit operations, no backprop)

Limitation: Hyperparameter tuning (sparsity target, learning rates, column/cell counts). Performance depends on input encoding quality. Less suited to function approximation (direct state→action mapping) than state representation.

References


9. Kanerva's Sparse Distributed Memory (SDM)

What It Is

Early model (1988) of associative memory using sparse high-dimensional space. Similar to HDC but predates modern formulations. Binary address space; sparse activation pattern; content-addressable retrieval.

How It Works

  1. Hard locations: Randomly sample N addresses in D-dimensional binary space (e.g., D=1000, N=1M)
  2. Hamming radius selection: For input x, activate all hard locations within Hamming distance k (e.g., k=100)
  3. Write: Increment counters at active locations for each bit of data
  4. Read: Average activated counters to reconstruct data

Result: associative memory with graceful degradation. Partial/noisy queries retrieve best match.

Binary Operations & Efficiency

  • Hamming distance computation: O(D) bit comparisons
  • Memory allocation: one counter per location per bit (can be binary: increment/decrement)
  • Distributed storage: each datum written to multiple locations; retrieval robust to damage

Computational Cost

  • Write: O(N_active × D) where N_active = number of hard locations within radius
  • Read: O(N_active × D)
  • Typical: N_active ∝ D (depends on D and radius threshold)

Sparse activation keeps practical cost low.

Real-Time Control Suitability

Moderate fit: Good for stored recall tasks (memorize state-action pairs). Less suited to generalization or online learning (no weight update mechanism, only counter increment).

Limitation: Essentially a lookup table with fuzzy matching; doesn't extrapolate beyond learned examples. Better as auxiliary memory (recall previous strategies) than primary controller.

References


Summary Table: Learning Methods Comparison

Method Binary Ops Real-Time Online Convergence Memory Suitability for Bot Control
HDC/VSA Excellent (XOR, Hamming) Single-pass Fast (1-pass train) High (10K+ bits) Good for discrete decisions, perception layers
LSM/ESN Good (spike events) Per-timestep Slow (data collection + solve) Moderate (N²) Excellent for temporal sequences
Random Perturbation Good (weight noise) Per-update Slow (noisy gradient) Moderate Moderate; online fine-tuning only
STDP Binary Excellent (event-driven) Per-spike Slow (biological timescale) Moderate (N²) Excellent if tuned; online, spiking-native
Evolution Strategies Fair (works with any) Batch (population eval) Moderate (population search) High (N × W) Good for offline training, topology search
BCM Good (rate-based) Per-spike-window Slow (self-organizing) Moderate (N² + thresholds) Good for feature layers; unsupervised
Winner-Take-All Excellent (argmax + Hamming) Per-sample Fast (local update) Moderate (N × D) Good for clustering, prototypes
SDR (HTM) Excellent (binary operations) Per-input Fast (online) Moderate (sparse matrix) Good for sequential prediction, anomaly detection
Kanerva SDM Excellent (Hamming distance) Per-query Instant (lookup) Very high (sparse matrix huge) Moderate; auxiliary memory only

Hybrid Approaches & Practical Recommendations

For SirRoboGarage Real-Time Aiming Task

Best candidates:

  1. STDP + LSM (spiking pipeline):

    • Reservoir learns temporal dynamics (lead prediction, ball tracking)
    • Output layer trained via STDP during deployment for fine-tuning
    • Fully neuromorphic; event-driven; online learning
  2. HDC for state discretization + LSM readout:

    • Encode sensor input (ball position, velocity) → binary HDC vector (single-pass)
    • Use as input to small LSM (~100 neurons)
    • LSM output trained with simple Hebbian rule
    • Hybrid: discrete perception, continuous temporal dynamics
  3. Competitive learning + Winner-take-all basis:

    • Learn clusters of ball positions / velocities (proto-strategy space)
    • Map each proto-state → action via local Hebbian learning
    • Fast inference; supports online cluster drift
  4. STDP + random weight perturbation:

    • STDP for online synaptic plasticity (slow, stable)
    • Perturbation for rapid adaptation to environment shifts (fast, noisy)
    • Dual timescale learning

Computational Footprint Estimate

  • Neuromorphic chip (SpiNNaker, Loihi): Full STDP + LSM (thousands of neurons) feasible
  • Embedded CPU (Jetson Nano): Small LSM (100–500 neurons) or HDC classifiers, ~10 ms latency per decision
  • Microcontroller (Arduino, ESP32): Competitive learning (few neurons) or small HDC lookup; no LSM (reservoir simulation too slow)

Open Questions for SirRoboGarage

  1. Latency vs. Accuracy: Does a 50 ms decision cycle allow LSM + STDP, or must we use single-pass HDC?
  2. Training data availability: Can we pre-collect battle logs for offline ES/LSM training, then fine-tune with STDP online?
  3. Hardware target: Is neuromorphic chip available, or must we use standard CPU/GPU? (Affects batch size, parallelism)
  4. Behavioral complexity: Is aiming task best solved by memorized prototypes (WTA + lookup) or by learned dynamics (LSM)?

References (Complete List)

Hyperdimensional Computing

Liquid State Machines & Reservoir Computing

Random Weight Perturbation

STDP & Binary Synapses

Evolution Strategies

BCM Theory

Competitive Learning & Winner-Take-All

Sparse Distributed Representations (HTM)

Sparse Distributed Memory (Kanerva)


Document Status: Research complete. All claims cited to primary sources (papers, official docs, peer-reviewed). Ready for implementation roadmap.