A model trained to recognize cats and dogs already knows something about edges, textures, and fur. Transfer learning exploits that: instead of starting from random weights, you begin with a network that has learned from a large dataset and fine-tune it on your smaller, specific problem. The knowledge gained on the first task transfers to the second.
The approach became standard because labeled data is scarce and compute is expensive. Training a vision model from scratch on 5,000 medical images rarely works. Taking a network pretrained on ImageNet and retraining its final layers often does. The early layers detect generic features; the later layers specialize.
Common strategies
- Feature extraction: freeze the pretrained layers and train only a new classifier head
- Fine-tuning: unfreeze some or all layers and continue training with a small learning rate
- Domain adaptation: adjust a model from one data distribution to another, such as synthetic to real
Large language models have pushed this further. A single pretrained model can be adapted to legal writing, medical coding, or customer support with modest additional training. The catch is negative transfer: if the source and target tasks are too different, the pretrained weights can hurt rather than help.
Comments
No comments yet. Be the first to share a thought.
Leave a comment