A child who has never seen a zebra can still recognize one after being told it looks like a horse with stripes. Zero-shot learning asks machines to do something similar: handle classes or tasks they were never trained on, using auxiliary information such as attribute descriptions or semantic relationships.
In image classification, a zero-shot model might be trained on cats and dogs but asked to identify a wolf. If the model has learned attributes like "has fur," "has pointed ears," and "lives in packs," it can compose those into a representation of wolf without ever seeing one. The bridge between seen and unseen classes is usually a shared semantic space, often built from text descriptions or knowledge graphs.
Why it matters
- New categories appear constantly and labeling them is slow
- Some classes are rare or impossible to collect at scale
- Large language models now perform many tasks zero-shot from a prompt alone
- It reduces the cost of deploying a model to new domains
Zero-shot performance still trails supervised learning on most benchmarks. The gap narrows when the auxiliary descriptions are rich and the unseen classes are close to the training distribution. When they are far apart, the model guesses poorly and often confidently.
Comments (2)
Leave a comment