Weights are learned. Hyperparameters are chosen. Learning rate, batch size, number of layers, dropout rate, and regularization strength all fall into this category. They control how training proceeds and how the model is structured.
Getting them wrong wastes time and compute. A learning rate too high diverges. Too low crawls. Batch size affects gradient noise and memory use. Depth and width determine capacity, which must match the complexity of the task and the size of the data.
Common tuning methods
- Grid search over a fixed set
- Random search for broader coverage
- Bayesian optimization for efficiency
- Population-based training
- Manual tuning guided by intuition
Tuning is expensive. Each configuration requires a full training run, and the search space grows quickly. Experienced practitioners narrow the space with domain knowledge before automating, which often beats brute-force search.
Comments (2)
Leave a comment