Change a few pixels in an image and a classifier sees something else entirely. Add imperceptible noise to a stop sign and a self-driving system may read it as a speed limit. These crafted inputs are adversarial examples, and they expose how differently machines and humans perceive the world.
Researchers generate them by computing the gradient of the model's loss with respect to the input, then nudging the input in the direction that maximizes error. The change is often invisible to the eye. Fast gradient sign method and projected gradient descent are two common techniques.
Why they matter
- Security systems can be fooled deliberately
- Autonomous vehicles may misread signs
- Medical imaging models can be manipulated
- Speech systems can be attacked with hidden audio
Defenses exist, including adversarial training and input sanitization, but none is fully reliable. Each new defense tends to be broken by a new attack, which keeps the field moving.
Comments (3)
Leave a comment