…
Skip to content
Topics
On this page

Types of Generative AI Models

Types of generative AI models are the main architecture families used to create new content: transformers, diffusion models, generative adversarial networks (GANs), variational autoencoders (VAEs) and flow-based models. Each family learns the distribution of its training data differently, which determines its output quality, generation speed and the kind of data it suits most.

  • Transformers: They generate sequences one token at a time and power most text and code models.
  • Diffusion models: They create images, audio or video by removing noise from a random starting point over many steps.
  • GANs: They train two competing networks, a generator and a discriminator, until generated samples look realistic.
  • VAEs: They compress data into a compact latent space and generate new samples by decoding points from that space.
  • Flow-based models: They use reversible transformations, so they can calculate the exact probability of any sample.
  • Combined designs: Many production systems join several families, such as a VAE, a transformer and a diffusion model in one image pipeline.
Four types of generative AI models and how each generates outputTransformer: earlier tokens go in and the next token comes out, repeated. Diffusion model: a noisy image is denoised step by step into a clear image. GAN: a generator creates a sample and a discriminator judges whether it is real or fake. VAE: an encoder compresses the input into a latent space and a decoder builds a new sample from it.Transformerdef parse(Next tokenDiffusion modelNoise to image, step by stepGANGeneratorDiscriminatorJudges generated vs real samplesVAEEncode, latent space, decode
Four types of generative AI models and how each generates output

For example, a developer product might use a transformer to write API docs and test code, and a diffusion model to create illustrations for the same documentation.

Overview of Types of Generative AI Models

TypeHow it worksTypical use
TransformerPredicts the next token using attention over earlier tokensText, code, chat and summarisation
Diffusion modelLearns to reverse a gradual noising processImages, video and audio
GANA generator competes against a discriminatorImage synthesis, super-resolution and style transfer
VAEEncodes data into a latent space, then decodes samples from itCompression, anomaly detection and latent spaces for diffusion
Flow-based modelApplies reversible mappings between data and a simple distributionDensity estimation and some image and audio generation

Transformer Models

  • Mechanism: The transformer architecture uses attention, which lets each token weigh every earlier token when predicting the next one.
  • Autoregressive generation: Output is produced token by token, so long responses take proportionally longer to generate.
  • Scale: Large transformer models trained on text and code are known as large language models.

Example: a code assistant completes def parse_invoice( with arguments and a docstring.

Diffusion Models

  • Mechanism: During training, noise is added to images step by step, and the network learns to predict and remove that noise.
  • Generation: Sampling starts from pure noise and applies many denoising steps, guided by an encoded text prompt.
  • Trade-off: Output quality is high, but the repeated steps make generation slower than a single forward pass.

Example: a docs team generates a clean architecture illustration from the prompt "three services behind an API gateway".

Generative Adversarial Networks (GANs)

  • Mechanism: The generator creates samples, and the discriminator learns to separate generated samples from real ones.
  • Strength: A trained generator produces an output in one forward pass, which makes inference fast.
  • Weakness: Training can be unstable, and mode collapse, where the generator produces only a few similar outputs, is common.

Example: a design team upscales low-resolution product screenshots for high-density displays.

Variational Autoencoders (VAEs)

  • Mechanism: An encoder maps each input to a probability distribution in a latent space, and a decoder reconstructs data from sampled points.
  • Smooth latent space: Nearby latent points decode to similar outputs, which allows controlled variation and interpolation.
  • Role today: VAEs often act as the compression stage inside latent diffusion systems rather than as standalone generators.

Example: a monitoring team flags unusual telemetry traces that the VAE reconstructs poorly.

Flow-Based Models

  • Mechanism: A sequence of invertible functions maps simple random noise to data and back again.
  • Exact likelihood: Because each step is reversible, the model can compute the precise probability of any input.
  • Current use: Related techniques, such as flow matching, appear in some recent image and audio generators.

Example: a research team estimates how likely a new audio sample is under the training distribution.

Combining Model Types

  • Latent diffusion: A VAE compresses images, a diffusion network denoises in that compressed space, and a text encoder conditions the result on the prompt.
  • Diffusion transformers: Some newer image and video models replace the older convolutional denoiser with a transformer.
  • Multimodal models: Transformers that accept text and images together can describe screenshots or answer questions about diagrams.

Objective Functions of GANs, VAEs and Diffusion Models

Each model family optimises a different training objective, the quantity that training either maximises or minimises. The objectives of GANs, VAEs and diffusion models explain much of their practical behaviour, from unstable adversarial training to comparatively slow diffusion sampling.

GAN Loss Function

The GAN loss function describes a two-player minimax game, in which the discriminator attempts to maximise the value and the generator attempts to minimise it:

  • : the discriminator's estimated probability, between 0 and 1, that a sample is genuine.
  • : a synthetic sample that the generator creates from random input noise .
  • : the probability distributions of the real training data and of the input noise.
  • : the expected value, which is an average calculated over many samples.

In other words, the discriminator wants scores near 1 for genuine samples and near 0 for generated ones, while the generator wants its samples classified as genuine. For one real UI screenshot scored 0.9 and one generated screenshot scored 0.2, the value is . When the generator becomes good enough that the discriminator outputs 0.5 for both samples, the value falls to , which equals , the theoretical value of the game at equilibrium.

VAE Evidence Lower Bound

A variational autoencoder maximises the evidence lower bound (ELBO), a quantity that can never exceed the log-probability of the training data:

  • : the encoder, which produces a probability distribution over latent codes for a given input.
  • : the decoder, which measures how accurately the input is reconstructed from a latent code.
  • : the prior distribution over latent codes, usually a standard normal distribution.
  • : the Kullback-Leibler divergence, which measures how different one probability distribution is from another.

The first term rewards accurate reconstruction of the input. The second term penalises latent codes that drift away from the prior, which keeps the latent space smooth and continuous enough for sampling new outputs.

Diffusion Model Math

Diffusion model math begins with a fixed forward process that adds a small amount of Gaussian noise at every step :

Because consecutive Gaussian steps combine into a single Gaussian, any step can be calculated directly from the clean data :

  • : the original clean data, such as an image, and its noisy version after steps.
  • : the variance of the noise added at step , determined in advance by a noise schedule.
  • : the proportion of the original signal's variance that survives after steps.
  • : standard Gaussian noise, where denotes the identity matrix.

The neural network is trained to predict the added noise from the noisy input and the step number, and generation then reverses the process one denoising step at a time.

Worked example. With a constant noise variance of 0.1, the cumulative products are , and . After three steps, , so the original pixel still contributes about 85 percent of its value. The program below calculates both GAN values and a complete 1,000-step linear schedule, with the variance rising from 0.0001 to 0.02, for one pixel value of 1.0 and one fixed noise value of -0.8.

Python
import math

# GAN value V(D, G) for one real sample and one generated sample.
def gan_value(d_real, d_fake):
    return math.log(d_real) + math.log(1 - d_fake)

print(f"confident D (0.9 real, 0.2 fake): V = {gan_value(0.9, 0.2):.3f}")
print(f"fooled D    (0.5 real, 0.5 fake): V = {gan_value(0.5, 0.5):.3f}")

# Diffusion forward process with the linear schedule from the DDPM paper:
# beta rises from 0.0001 to 0.02 over T = 1000 steps.
T = 1000
betas = [1e-4 + (0.02 - 1e-4) * (t - 1) / (T - 1) for t in range(1, T + 1)]
x0, eps = 1.0, -0.8  # one pixel value and one fixed noise draw, so output repeats

alpha_bar = 1.0
print(f"\n{'t':>5}{'alpha_bar':>12}{'signal':>8}{'noise':>7}{'x_t':>8}")
for t, beta in enumerate(betas, start=1):
    alpha_bar *= 1 - beta
    if t in (1, 100, 250, 500, 750, 1000):
        signal, noise = math.sqrt(alpha_bar), math.sqrt(1 - alpha_bar)
        x_t = signal * x0 + noise * eps
        print(f"{t:>5}{alpha_bar:>12.5f}{signal:>8.3f}{noise:>7.3f}{x_t:>8.3f}")
Output
confident D (0.9 real, 0.2 fake): V = -0.329
fooled D    (0.5 real, 0.5 fake): V = -1.386

    t   alpha_bar  signal  noise     x_t
    1     0.99990   1.000  0.010   0.992
  100     0.89702   0.947  0.321   0.690
  250     0.52409   0.724  0.690   0.172
  500     0.07859   0.280  0.960  -0.488
  750     0.00335   0.058  0.998  -0.741
 1000     0.00004   0.006  1.000  -0.794
In the diffusion forward process the image signal fades as noise grows over 1,000 stepsA line chart over steps t from 0 to 1000. The signal coefficient, the square root of alpha bar t, starts at 1 and falls to 0.724 at step 250, 0.280 at step 500 and 0.006 at step 1000. The noise coefficient, the square root of one minus alpha bar t, starts at 0 and rises to about 1. The two curves cross near step 260, where signal and noise contribute equally. Values use a linear beta schedule from 0.0001 to 0.02 over 1000 steps and are computed by the worked example.1.0002505007501000coefficientstep tsignal √ᾱnoise √(1 - ᾱ)equal near t = 260
In the diffusion forward process the image signal fades as noise grows over 1,000 steps
  • Signal fades gradually: Signal and noise contribute equally near step 260, and by step 1,000 the signal coefficient has fallen to 0.006.
  • Sampling starts from noise: Because the final step is almost pure noise, generation can begin from random noise and reverse the process.
  • GANs aim for equilibrium: The ideal outcome is a discriminator that can no longer distinguish genuine from generated samples, which is the case.

In practice, each reverse diffusion step requires one call to the network, so the number of steps determines the generation cost, which is why faster samplers use far fewer steps than the 1,000 used during training.

How to Choose

  • Text or code output: Choose a transformer model, usually through an existing LLM API or an open-weight model.
  • High-quality images or video: Choose a diffusion model, accepting slower generation in exchange for quality.
  • Low-latency image generation: Consider a GAN when real-time output matters more than variety.
  • Compact representations or anomaly detection: Choose a VAE, which provides a structured latent space.
  • Exact probabilities: Choose a flow-based model when likelihood values are required.

The training process shared by these families is described in how generative AI works, and industry uses are listed in applications of generative AI. For the definition, see what is generative AI.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. Which model type generates images by removing noise over many steps?

Frequently Asked Questions

What are the main types of generative AI models?

The main families are transformers, diffusion models, generative adversarial networks, variational autoencoders and flow-based models. Transformers dominate text and code, while diffusion models dominate image and video generation.

How do GAN vs VAE vs diffusion models compare for images?

A GAN trains a generator against a discriminator and creates an image in one pass, which is fast but can be unstable to train. A diffusion model removes noise over many steps, which is slower but usually more stable and higher in quality. A VAE alone tends to produce blurrier images, so it mostly serves as the compression stage inside diffusion systems.

Is ChatGPT a GAN or a transformer?

ChatGPT is built on a transformer-based large language model. It generates text token by token and does not use a generator and discriminator pair.

What is a VAE used for?

A variational autoencoder compresses data into a smooth latent space and decodes new samples from it. It is used for anomaly detection, data compression and as a component inside image diffusion systems.