Labels are expensive. Raw data is cheap. Semi-supervised learning combines a small labeled set with a large unlabeled set, using the structure of unlabeled data to improve generalization beyond what labels alone could achieve.
Approaches vary. Self-training labels unlabeled examples with a model, then retrains on the combined set. Consistency regularization forces the model to produce similar outputs for perturbed versions of the same input. Graph-based methods propagate labels through similarity networks.
Common semi-supervised techniques
- Self-training and pseudo-labeling
- Consistency regularization
- Entropy minimization
- Graph-based label propagation
- Contrastive pretraining followed by fine-tuning
Results depend heavily on data. If unlabeled examples come from a different distribution than labeled ones, performance can degrade. Careful validation and domain matching matter as much as the algorithm.
Comments
No comments yet. Be the first to share a thought.
Leave a comment