Destroy an image with noise, then learn to reverse the process. Diffusion models train by gradually adding noise to data, then learning to denoise step by step. Once trained, they can start from pure noise and generate new samples that look like the training distribution.
The approach powers many of today's image generators. Quality is high, and the training objective is stable compared to generative adversarial networks, which can suffer from mode collapse and unstable dynamics.
Key stages
- Forward process, adding noise over many steps
- Reverse process, learning to denoise
- Sampling, generating from random noise
- Conditioning, guiding output with text or other input
The cost is speed. Generating a single image may require dozens or hundreds of network passes. Distillation and faster samplers reduce the step count, but diffusion remains more compute-intensive than some alternatives.
Comments
No comments yet. Be the first to share a thought.
Leave a comment