Convolutional networks struggle with orientation. Rotate a face and a CNN may still detect a nose, but it loses the spatial relationship between nose and mouth. Capsule networks attempt to preserve that information by grouping neurons into capsules that output vectors rather than scalars.
Each capsule encodes both the presence of a feature and its properties, such as pose or orientation. Dynamic routing lets lower-level capsules send their output to higher-level capsules that agree with them, building part-whole relationships.
Potential advantages
- Better handling of rotation and viewpoint
- Preservation of spatial hierarchies
- Less need for huge training sets
- More interpretable internal representations
Adoption has been limited. Capsule networks train slowly and have not matched CNNs on large-scale benchmarks. They remain an active research direction rather than a production standard.
Comments
No comments yet. Be the first to share a thought.
Leave a comment