Network architecture #9

Closed
opened 2026-08-16 11:51:22 +02:00 by SirStone · 1 comment
Owner

Question

What MLP shape for the PPO actor-critic?

  • Shared trunk with separate actor/critic heads, or fully separate networks?
  • How many hidden layers, what sizes?
  • Activation functions (ReLU, tanh, etc.)?
  • Output layer activations (actor already partly decided in #5 — tanh/sigmoid per action dim)
  • Critic output: single linear value node

Context: 42-float state input (#4), 5-float action output (#5), PPO with mini-batch 64 (#8). Pure Nim + Arraymancer.

## Question What MLP shape for the PPO actor-critic? - Shared trunk with separate actor/critic heads, or fully separate networks? - How many hidden layers, what sizes? - Activation functions (ReLU, tanh, etc.)? - Output layer activations (actor already partly decided in #5 — tanh/sigmoid per action dim) - Critic output: single linear value node Context: 42-float state input (#4), 5-float action output (#5), PPO with mini-batch 64 (#8). Pure Nim + Arraymancer.
SirStone added the wayfinder:grilling label 2026-08-16 11:51:22 +02:00
Author
Owner

Resolution

Structure: Separate actor and critic networks (no shared trunk). ~8K parameters total + 5 log-std floats.

Actor:

  • Input: 42 floats (state vector from #4)
  • Hidden: 64 → tanh → 64 → tanh
  • Output: 5 linear logits, per-dimension activation applied post-forward-pass:
    • targetSpeed: tanh × 8
    • turnRate: tanh × speed-dependent max
    • gunTurnRate: tanh × 20
    • fireDecision: tanh, threshold at 0 (masked when gun hot)
    • firePower: sigmoid × 2.9 + 0.1
  • Stochastic policy: separate learnable log-std parameter vector (5 floats, initialized to 0 → std=1.0). Actions sampled from Normal(mean=network_output, std=exp(log_std)). Log-std is state-independent, updated by the optimizer alongside network weights.

Critic:

  • Input: 42 floats (same state vector)
  • Hidden: 64 → tanh → 64 → tanh
  • Output: 1 linear node, no activation (unbounded value estimate)

Weight initialization: Arraymancer defaults. Orthogonal init is the upgrade path if convergence fails.

Upgrade paths:

  • 128→64 layers if 64→64 plateaus early
  • Orthogonal initialization if training fails to converge
  • Shared trunk if training is too slow
  • State-dependent std (network output) if state-independent std is too coarse
## Resolution **Structure:** Separate actor and critic networks (no shared trunk). ~8K parameters total + 5 log-std floats. **Actor:** - Input: 42 floats (state vector from #4) - Hidden: 64 → tanh → 64 → tanh - Output: 5 linear logits, per-dimension activation applied post-forward-pass: - targetSpeed: tanh × 8 - turnRate: tanh × speed-dependent max - gunTurnRate: tanh × 20 - fireDecision: tanh, threshold at 0 (masked when gun hot) - firePower: sigmoid × 2.9 + 0.1 - Stochastic policy: separate learnable log-std parameter vector (5 floats, initialized to 0 → std=1.0). Actions sampled from Normal(mean=network_output, std=exp(log_std)). Log-std is state-independent, updated by the optimizer alongside network weights. **Critic:** - Input: 42 floats (same state vector) - Hidden: 64 → tanh → 64 → tanh - Output: 1 linear node, no activation (unbounded value estimate) **Weight initialization:** Arraymancer defaults. Orthogonal init is the upgrade path if convergence fails. **Upgrade paths:** - 128→64 layers if 64→64 plateaus early - Orthogonal initialization if training fails to converge - Shared trunk if training is too slow - State-dependent std (network output) if state-independent std is too coarse
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#9