Types of Generative AI Models
Types of generative AI models are the main architecture families used to create new content: transformers, diffusion models, generative adversarial networks (GANs), variational autoencoders (VAEs) and flow-based models. Each family learns the distribution of its training data differently, which determines its output quality, generation speed and the kind of data it suits most.
- Transformers: They generate sequences one token at a time and power most text and code models.
- Diffusion models: They create images, audio or video by removing noise from a random starting point over many steps.
- GANs: They train two competing networks, a generator and a discriminator, until generated samples look realistic.
- VAEs: They compress data into a compact latent space and generate new samples by decoding points from that space.
- Flow-based models: They use reversible transformations, so they can calculate the exact probability of any sample.
- Combined designs: Many production systems join several families, such as a VAE, a transformer and a diffusion model in one image pipeline.
For example, a developer product might use a transformer to write API docs and test code, and a diffusion model to create illustrations for the same documentation.
Overview of Types of Generative AI Models
| Type | How it works | Typical use |
|---|---|---|
| Transformer | Predicts the next token using attention over earlier tokens | Text, code, chat and summarisation |
| Diffusion model | Learns to reverse a gradual noising process | Images, video and audio |
| GAN | A generator competes against a discriminator | Image synthesis, super-resolution and style transfer |
| VAE | Encodes data into a latent space, then decodes samples from it | Compression, anomaly detection and latent spaces for diffusion |
| Flow-based model | Applies reversible mappings between data and a simple distribution | Density estimation and some image and audio generation |
Transformer Models
- Mechanism: The transformer architecture uses attention, which lets each token weigh every earlier token when predicting the next one.
- Autoregressive generation: Output is produced token by token, so long responses take proportionally longer to generate.
- Scale: Large transformer models trained on text and code are known as large language models.
Example: a code assistant completes def parse_invoice( with arguments and a docstring.
Diffusion Models
- Mechanism: During training, noise is added to images step by step, and the network learns to predict and remove that noise.
- Generation: Sampling starts from pure noise and applies many denoising steps, guided by an encoded text prompt.
- Trade-off: Output quality is high, but the repeated steps make generation slower than a single forward pass.
Example: a docs team generates a clean architecture illustration from the prompt "three services behind an API gateway".
Generative Adversarial Networks (GANs)
- Mechanism: The generator creates samples, and the discriminator learns to separate generated samples from real ones.
- Strength: A trained generator produces an output in one forward pass, which makes inference fast.
- Weakness: Training can be unstable, and mode collapse, where the generator produces only a few similar outputs, is common.
Example: a design team upscales low-resolution product screenshots for high-density displays.
Variational Autoencoders (VAEs)
- Mechanism: An encoder maps each input to a probability distribution in a latent space, and a decoder reconstructs data from sampled points.
- Smooth latent space: Nearby latent points decode to similar outputs, which allows controlled variation and interpolation.
- Role today: VAEs often act as the compression stage inside latent diffusion systems rather than as standalone generators.
Example: a monitoring team flags unusual telemetry traces that the VAE reconstructs poorly.
Flow-Based Models
- Mechanism: A sequence of invertible functions maps simple random noise to data and back again.
- Exact likelihood: Because each step is reversible, the model can compute the precise probability of any input.
- Current use: Related techniques, such as flow matching, appear in some recent image and audio generators.
Example: a research team estimates how likely a new audio sample is under the training distribution.
Combining Model Types
- Latent diffusion: A VAE compresses images, a diffusion network denoises in that compressed space, and a text encoder conditions the result on the prompt.
- Diffusion transformers: Some newer image and video models replace the older convolutional denoiser with a transformer.
- Multimodal models: Transformers that accept text and images together can describe screenshots or answer questions about diagrams.
Objective Functions of GANs, VAEs and Diffusion Models
Each model family optimises a different training objective, the quantity that training either maximises or minimises. The objectives of GANs, VAEs and diffusion models explain much of their practical behaviour, from unstable adversarial training to comparatively slow diffusion sampling.
GAN Loss Function
The GAN loss function describes a two-player minimax game, in which the discriminator attempts to maximise the value and the generator attempts to minimise it:
- : the discriminator's estimated probability, between 0 and 1, that a sample is genuine.
- : a synthetic sample that the generator creates from random input noise .
- : the probability distributions of the real training data and of the input noise.
- : the expected value, which is an average calculated over many samples.
In other words, the discriminator wants scores near 1 for genuine samples and near 0 for generated ones, while the generator wants its samples classified as genuine. For one real UI screenshot scored 0.9 and one generated screenshot scored 0.2, the value is . When the generator becomes good enough that the discriminator outputs 0.5 for both samples, the value falls to , which equals , the theoretical value of the game at equilibrium.
VAE Evidence Lower Bound
A variational autoencoder maximises the evidence lower bound (ELBO), a quantity that can never exceed the log-probability of the training data:
- : the encoder, which produces a probability distribution over latent codes for a given input.
- : the decoder, which measures how accurately the input is reconstructed from a latent code.
- : the prior distribution over latent codes, usually a standard normal distribution.
- : the Kullback-Leibler divergence, which measures how different one probability distribution is from another.
The first term rewards accurate reconstruction of the input. The second term penalises latent codes that drift away from the prior, which keeps the latent space smooth and continuous enough for sampling new outputs.
Diffusion Model Math
Diffusion model math begins with a fixed forward process that adds a small amount of Gaussian noise at every step :
Because consecutive Gaussian steps combine into a single Gaussian, any step can be calculated directly from the clean data :
- : the original clean data, such as an image, and its noisy version after steps.
- : the variance of the noise added at step , determined in advance by a noise schedule.
- : the proportion of the original signal's variance that survives after steps.
- : standard Gaussian noise, where denotes the identity matrix.
The neural network is trained to predict the added noise from the noisy input and the step number, and generation then reverses the process one denoising step at a time.
Worked example. With a constant noise variance of 0.1, the cumulative products are , and . After three steps, , so the original pixel still contributes about 85 percent of its value. The program below calculates both GAN values and a complete 1,000-step linear schedule, with the variance rising from 0.0001 to 0.02, for one pixel value of 1.0 and one fixed noise value of -0.8.
import math
# GAN value V(D, G) for one real sample and one generated sample.
def gan_value(d_real, d_fake):
return math.log(d_real) + math.log(1 - d_fake)
print(f"confident D (0.9 real, 0.2 fake): V = {gan_value(0.9, 0.2):.3f}")
print(f"fooled D (0.5 real, 0.5 fake): V = {gan_value(0.5, 0.5):.3f}")
# Diffusion forward process with the linear schedule from the DDPM paper:
# beta rises from 0.0001 to 0.02 over T = 1000 steps.
T = 1000
betas = [1e-4 + (0.02 - 1e-4) * (t - 1) / (T - 1) for t in range(1, T + 1)]
x0, eps = 1.0, -0.8 # one pixel value and one fixed noise draw, so output repeats
alpha_bar = 1.0
print(f"\n{'t':>5}{'alpha_bar':>12}{'signal':>8}{'noise':>7}{'x_t':>8}")
for t, beta in enumerate(betas, start=1):
alpha_bar *= 1 - beta
if t in (1, 100, 250, 500, 750, 1000):
signal, noise = math.sqrt(alpha_bar), math.sqrt(1 - alpha_bar)
x_t = signal * x0 + noise * eps
print(f"{t:>5}{alpha_bar:>12.5f}{signal:>8.3f}{noise:>7.3f}{x_t:>8.3f}")confident D (0.9 real, 0.2 fake): V = -0.329
fooled D (0.5 real, 0.5 fake): V = -1.386
t alpha_bar signal noise x_t
1 0.99990 1.000 0.010 0.992
100 0.89702 0.947 0.321 0.690
250 0.52409 0.724 0.690 0.172
500 0.07859 0.280 0.960 -0.488
750 0.00335 0.058 0.998 -0.741
1000 0.00004 0.006 1.000 -0.794- Signal fades gradually: Signal and noise contribute equally near step 260, and by step 1,000 the signal coefficient has fallen to 0.006.
- Sampling starts from noise: Because the final step is almost pure noise, generation can begin from random noise and reverse the process.
- GANs aim for equilibrium: The ideal outcome is a discriminator that can no longer distinguish genuine from generated samples, which is the case.
In practice, each reverse diffusion step requires one call to the network, so the number of steps determines the generation cost, which is why faster samplers use far fewer steps than the 1,000 used during training.
How to Choose
- Text or code output: Choose a transformer model, usually through an existing LLM API or an open-weight model.
- High-quality images or video: Choose a diffusion model, accepting slower generation in exchange for quality.
- Low-latency image generation: Consider a GAN when real-time output matters more than variety.
- Compact representations or anomaly detection: Choose a VAE, which provides a structured latent space.
- Exact probabilities: Choose a flow-based model when likelihood values are required.
The training process shared by these families is described in how generative AI works, and industry uses are listed in applications of generative AI. For the definition, see what is generative AI.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. Which model type generates images by removing noise over many steps?
Frequently Asked Questions
What are the main types of generative AI models?
The main families are transformers, diffusion models, generative adversarial networks, variational autoencoders and flow-based models. Transformers dominate text and code, while diffusion models dominate image and video generation.
How do GAN vs VAE vs diffusion models compare for images?
A GAN trains a generator against a discriminator and creates an image in one pass, which is fast but can be unstable to train. A diffusion model removes noise over many steps, which is slower but usually more stable and higher in quality. A VAE alone tends to produce blurrier images, so it mostly serves as the compression stage inside diffusion systems.
Is ChatGPT a GAN or a transformer?
ChatGPT is built on a transformer-based large language model. It generates text token by token and does not use a generator and discriminator pair.
What is a VAE used for?
A variational autoencoder compresses data into a smooth latent space and decodes new samples from it. It is used for anomaly detection, data compression and as a component inside image diffusion systems.
Related Articles
- What is Generative AILearn what generative AI is, how it creates text, code and images from learned patterns, its traits, uses and limits, with a Python example and diagram.
- How Generative AI WorksLearn how generative AI works, from data preparation and pre-training to sampling and inference, with a Python example of scores, softmax and temperature.
- Transformer Architecture ExplainedTransformer architecture explained: self-attention, query, key and value vectors, multi-head attention and stacked layers, with Python attention code.
- What is a Large Language Model (LLM)Learn what a large language model is, how an LLM predicts the next token, its key characteristics, uses and limits, with a Python example and diagram.
- Top Applications of Generative AI in IndiaApplications of generative AI in India across IT services, banking, healthcare, public services and education, with engineering use cases and build steps.
- Generative AI vs Traditional AICompare generative AI vs traditional AI: output, training data, models, evaluation and cost, with a comparison table and one bug report handled both ways.