ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI Encyclopedia
What Is a Diffusion Model? Why Image Generation First Scrambles an Image and Then Slowly Restores It

What Is a Diffusion Model? Why Image Generation First Scrambles an Image and Then Slowly Restores It

AI Encyclopedia • Admin • • 8 views

A diffusion model is a generative model that creates new content by first gradually adding noise to data and then learning how to gradually remove that noise, and it is one of the mainstream methods for image generation today. It does not draw an image directly. Instead, it starts from pure noise and clears the noise away step by step, letting the picture slowly take shape.

First Add Noise, Then Remove It: Two Opposite Processes

The forward process adds noise: it takes a real image and mixes in random noise layer by layer in fixed steps, until after many steps the image becomes noise in which the original is completely unrecognizable. This process requires no learning; it simply scrambles the image according to a rule.

The reverse process removes noise, and it is the direction used during generation. The model starts from a mass of random noise, predicts and removes part of the noise to obtain a slightly clearer result, and then keeps iterating until it produces a complete new image. Generation starts from random noise, not from any existing image.

What Training Actually Learns Is to Predict Noise

Training does not make the model memorize images. Instead, it makes the model practice predicting noise: given a noised image and the noise step, it must accurately guess what noise was mixed in. The same image is noised to different degrees so the model can practice repeatedly.

Once it can predict noise, denoising has a basis. At each step, subtracting part of the predicted noise brings the image a little closer to how real data looks. What the model learns is the distribution of the data itself, so it can generate new combinations that never appeared in the training set, rather than copying an original image.

The Number of Steps Decides Speed and Also Affects Quality

Generation with a diffusion model is iterative: each step changes only a little, and dozens of steps are usually needed to complete one image. With more steps, each change is smaller, and the result is often more stable with cleaner details, but it takes longer. With too few steps, each change spans too much, and distorted structures and blurred details are more likely.

This is also its clear weakness: compared with methods that produce a result in one pass, it is naturally slower. Later accelerated sampling and distillation techniques all aim, at their core, to reduce the number of steps needed while losing as little quality as possible.

How It Differs from GAN and Autoregressive Generation

A GAN relies on adversarial training between a generator and a discriminator. It outputs a complete result in one pass and is fast, but training is unstable and diversity is sometimes limited. Autoregressive generation, by contrast, works like writing a sentence: it generates one pixel or one small block after another in sequence, which is logically clear but prone to accumulating errors.

A diffusion model is relatively stable to train, and its results are diverse with controllable details. It trades multiple iterative steps for quality and stability, at the cost of inference speed. The three types of methods do not absolutely replace one another; they mainly involve different trade-offs.

Two Most Common Misconceptions

The first misconception is thinking that a diffusion model blurs an existing image and then restores it. Adding noise and then removing it is only a way to construct a learning task during training. During actual generation, the inputs are only random noise and conditions such as a text prompt; there is no original image to restore, and the output is a brand-new image.

The second misconception is thinking that it modifies some base image during generation. Only in a dedicated mode such as image-to-image does the user actively provide a base image and control how much it changes, and that is not the same as the default text-to-image generation process.

Stable Diffusion is one of the most representative applications of diffusion models, and it uses latent diffusion: it first compresses an image into a smaller latent space, performs the noising and denoising there, and finally decodes it back into an image. This greatly reduces the amount of computation, so it can also run on ordinary graphics cards.

Recommended Tools

More