Take a simple function, stack many copies, and let each one learn. A neural network is a composition of layers, where each layer applies weights, a bias, and a nonlinearity. Training adjusts those weights to fit data.
The design is loosely inspired by biological neurons, though the resemblance is superficial. What matters is that the architecture can approximate complex functions given enough capacity and data. Universal approximation theorems support this, though they say nothing about whether training will find a good solution.
Common network types
- Feedforward networks
- Convolutional networks
- Recurrent networks
- Transformers
- Graph networks
Depth helps. Shallow networks can approximate many functions, but deep ones learn hierarchical features more efficiently. That efficiency is why deep learning took off once training became practical.
Comments
No comments yet. Be the first to share a thought.
Leave a comment