Stack linear operations and you still get a line. Add an activation function between layers and the network bends. That bend is what lets a neural network model curves, edges, and everything else that straight arithmetic cannot capture.
Rectified linear units, or ReLUs, dominate modern practice. They output zero for negative inputs and pass positive values unchanged. Cheap to compute, and they train fast. Sigmoid and tanh came first, squashing outputs into bounded ranges, but they saturate and slow learning in deep networks.
Common activation functions
- ReLU, the default for most deep networks
- Leaky ReLU, allowing a small gradient for negatives
- Sigmoid, used for binary outputs
- Tanh, zero-centered squashing
- Softmax, converting scores to probabilities
Choosing wrong matters. A network of only linear activations collapses into a single linear transformation, no matter how many layers you stack. Nonlinearity is not decoration. It is the reason depth helps at all.
Comments
No comments yet. Be the first to share a thought.
Leave a comment