Imagine walking down a foggy hill, feeling the slope under your feet, and stepping in the direction that goes down. Gradient descent does the same with a loss function. It computes the slope, takes a step, and repeats until the loss stops improving.
The learning rate controls step size. Too large and you overshoot the minimum. Too small and training takes forever. Momentum, Adam, and other optimizers adapt the rate during training to speed convergence.
Common variants
- Batch gradient descent using all data
- Stochastic gradient descent using one sample
- Mini-batch gradient descent using small groups
- Momentum and adaptive methods like Adam
Non-convex loss surfaces make the journey tricky. Saddle points, local minima, and plateaus all slow progress. In practice, stochastic methods escape shallow traps and find good solutions even without guarantees of global optimality.
Comments
No comments yet. Be the first to share a thought.
Leave a comment