Training a network requires knowing how to adjust each weight. Backpropagation computes that. It applies the chain rule of calculus to propagate the error from the output layer backward through the network, assigning a gradient to every weight along the way.
The algorithm works in two passes. A forward pass computes predictions and loss. A backward pass calculates gradients with respect to each parameter. An optimizer like stochastic gradient descent then updates the weights in the direction that reduces loss.
Key steps
- Forward pass through all layers
- Compute loss against targets
- Backward pass applying the chain rule
- Update weights using gradients
The technique was popularized in the 1980s and made deep learning practical when paired with faster hardware and better activation functions. Without it, training multi-layer networks would require intractable search over weight space.
Comments (2)
Leave a comment