Learning rate and optimizer hyperparameters #12

Closed
opened 2026-08-16 12:56:47 +02:00 by SirStone · 1 comment
Owner

Question

What optimizer, learning rate, and related hyperparameters for the PPO actor and critic networks?

  • Optimizer choice: Adam (standard PPO), SGD, or other?
  • Learning rate: fixed or decayed? Same for actor and critic, or separate?
  • Adam-specific: epsilon, beta1, beta2?
  • Gradient clipping: max norm?

Context: separate actor/critic networks (#9), 64→64 tanh, ~8K params total. 4 PPO epochs, mini-batch 64 (#8). Entropy coefficient 0.01 (#10).

Blocked by: nothing (all dependencies resolved)

## Question What optimizer, learning rate, and related hyperparameters for the PPO actor and critic networks? - Optimizer choice: Adam (standard PPO), SGD, or other? - Learning rate: fixed or decayed? Same for actor and critic, or separate? - Adam-specific: epsilon, beta1, beta2? - Gradient clipping: max norm? Context: separate actor/critic networks (#9), 64→64 tanh, ~8K params total. 4 PPO epochs, mini-batch 64 (#8). Entropy coefficient 0.01 (#10). Blocked by: nothing (all dependencies resolved)
SirStone added the wayfinder:grilling label 2026-08-16 12:56:47 +02:00
Author
Owner

Resolution

Optimizer: Adam. Two instances — one for actor, one for critic. Arraymancer has Adam built in (confirmed working in spike).

Learning rate: 3e-4, same for actor and critic. Fixed — no decay (the bot trains indefinitely across battles, no endpoint to decay toward).

Adam hyperparameters: Defaults — beta1=0.9, beta2=0.999, epsilon=1e-8.

Gradient clipping: Max global norm = 0.5. Standard PPO practice, prevents catastrophic weight updates from noisy advantages. May need hand-rolling if Arraymancer doesn't support it natively (a few lines).

Upgrade paths:

  • Separate actor/critic learning rates if critic converges too fast/slow relative to actor
  • Cosine annealing with warm restarts if fine-tuning against a specific opponent
  • Remove clipping if training is stable without it
## Resolution **Optimizer:** Adam. Two instances — one for actor, one for critic. Arraymancer has Adam built in (confirmed working in spike). **Learning rate:** 3e-4, same for actor and critic. Fixed — no decay (the bot trains indefinitely across battles, no endpoint to decay toward). **Adam hyperparameters:** Defaults — beta1=0.9, beta2=0.999, epsilon=1e-8. **Gradient clipping:** Max global norm = 0.5. Standard PPO practice, prevents catastrophic weight updates from noisy advantages. May need hand-rolling if Arraymancer doesn't support it natively (a few lines). **Upgrade paths:** - Separate actor/critic learning rates if critic converges too fast/slow relative to actor - Cosine annealing with warm restarts if fine-tuning against a specific opponent - Remove clipping if training is stable without it
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#12