Nobody tells a child which sounds belong to which language before they start sorting them. Unsupervised learning works the same way: the algorithm receives data without labels and must find structure on its own. There is no correct answer to check against, only patterns that emerge from the data's own shape.
Clustering is the classic example. A retailer might feed a year of purchase records into a clustering algorithm and discover five distinct customer groups it never defined. Dimensionality reduction is another: compressing hundreds of measurements into a two-dimensional map that preserves meaningful relationships.
Where it earns its keep
- Customer segmentation for marketing
- Anomaly detection in network traffic or manufacturing
- Topic discovery in large document collections
- Pretraining representations for later supervised tasks
The absence of labels is both freedom and burden. Without ground truth, evaluating results is subjective and often requires a human to interpret whether the clusters make sense. Still, when labels are impossible or too costly, unsupervised methods are the only way to see what the data is hiding.
Comments (2)
Leave a comment