iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Autoregressive models factor a distribution into ordered conditionals; variational autoencoders (VAEs) introduce latent variables and approximate inference; normalizing flows transform a simple density through invertible mappings; and generative adversarial networks (GANs) train a generator against a discriminator. These are four different ways to make complex data distributions computationally tractable—and they differ in what they optimize and what they let you calculate.
What makes a complex distribution learnable?
A generative model aims to capture patterns in observed data well enough to represent or produce plausible examples. The challenge is that a joint distribution over many interdependent values can be difficult to describe or compute directly. Each model family changes the structure of that problem: it may break the joint distribution into smaller pieces, add hidden variables, reshape a tractable density, or use a learned adversarial signal.
A useful introductory distinction is between likelihood-based approaches and likelihood-free ones. Autoregressive models, VAEs, and normalizing flows are commonly grouped with likelihood-based methods because their formulations support probability or density calculations, though their objectives and mechanics differ. The original GAN formulation instead centers training on an adversarial signal, not explicit per-example likelihood evaluation. This is a broad teaching distinction, not an exhaustive taxonomy of every variant.
How do autoregressive models represent a distribution?
They apply the probability chain rule to factor a joint distribution into a product of conditional distributions. For variables ordered as x₁ through xₙ, the factorization is p(x) = ∏ᵢ p(xᵢ | x₁, …, xᵢ₋₁). This identity is exact; it does not depend on a neural network or on how well a model has been trained. In practice, a model learns the conditional distributions, and the chosen ordering affects how it behaves.
#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
Because the model predicts each value conditioned on preceding values, it can evaluate the corresponding conditional probabilities and optimize likelihood. Generation follows the same dependency: sample the first value, then the next given the first, and continue. That sequential process can make generation slow. PixelRNN, for example, predicts image pixels sequentially. Training-time parallelism depends on the architecture and factorization; it should not be confused with the sequential dependencies involved in sampling. Pixel Recurrent Neural Networks (2016)
How do variational autoencoders use latent variables?
A latent variable is an unobserved value that helps explain an observation. A VAE models observations x through a latent variable z, using a prior p(z) and a generative model p(x|z), often called the decoder. To learn about z from an observed x, it also uses an approximate inference model, commonly called an encoder or recognition model, to represent q(z|x).
The exact posterior p(z|x) can be intractable to calculate. Rather than treating the encoder as that exact posterior, a VAE uses q(z|x) as an approximation and optimizes a variational lower bound, or ELBO, on the data log-likelihood. This gives a tractable training objective while making the quality of the posterior approximation part of the modeling story. The foundational paper, Auto-Encoding Variational Bayes, was submitted in 2013 and revised in 2022.
Recommended Free Tools
How do normalizing flows make density calculations tractable?
A normalizing flow starts with a simple distribution whose density is known, then applies a sequence of invertible transformations to map it into a more complex distribution. Since each transformation can be reversed, the change in density can be accounted for along the mapping. This structure can support explicit density calculations while allowing the transformed distribution to represent richer patterns.
Rank #3
Invertibility is both the enabling idea and a constraint: the transformations must be designed so they can be inverted and their density changes handled. Different flow designs can have different computational costs; invertibility alone does not imply identical efficiency. Rezende and Mohamed’s Variational Inference with Normalizing Flows presents flows for variational inference; it should not be read as a claim that every flow design has the same properties.
How do GANs learn without centering explicit likelihood?
A GAN trains two models in an adversarial process. The generator G produces samples, while the discriminator D learns to distinguish samples from the training data from samples produced by G. The generator is trained in response to the discriminator’s signal, creating a minimax game rather than making per-example likelihood the central training quantity.
Rank #4
Goodfellow and coauthors described the framework as one that “simultaneously train[s] two models”: a generative model that captures the data distribution and a discriminative model that estimates whether a sample came from training data rather than the generator. The discriminator supplies a learning signal; it is not simply a direct density estimator. The original formulation is described in Generative Adversarial Networks (2014).
Why does likelihood training connect KL divergence to negative log-likelihood?
For an explicit-likelihood model, minimizing DKL(pdata || pmodel) with respect to model parameters is equivalent to minimizing cross-entropy: the data entropy is constant with respect to those parameters. Because the true data distribution is not known in closed form, training estimates the expectation using examples, yielding negative log-likelihood minimization.
Best Value
This connection explains the objective for likelihood-based approaches; it does not describe the original GAN minimax objective. A GAN’s adversarial training signal is structurally different.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How should you compare the four model families?
| Family | How it structures the problem | Training and density perspective | Key trade-off |
|---|---|---|---|
| Autoregressive | Exact chain-rule factorization into ordered conditionals. | Fits conditional probabilities and can optimize likelihood. | Sequential dependencies can slow generation; training parallelism varies with architecture and factorization. |
| VAE | Models observations through latent variables and an approximate inference model. | Optimizes an ELBO when exact posterior inference is intractable. | The latent representation and inference model provide structure, while the posterior approximation and objective shape what is learned. |
| Normalizing flow | Maps a simple density through a sequence of invertible transformations. | Transformation structure can support explicit density calculations. | Invertibility constrains available transformations, and computational costs vary by design. |
| GAN | Trains a generator and discriminator in an adversarial minimax process. | The original formulation uses the discriminator’s learning signal rather than explicit per-example likelihood as its central objective. | Learning depends on the interaction between two models rather than a direct likelihood objective. |
There is no universal ranking implied by these differences. The useful questions are whether you need explicit density access, what dependencies generation must follow, how inference is structured, and which training objective suits the goal. These are architectural and objective distinctions, not a head-to-head performance result: the cited sources do not establish a single comparison across all four families under one dataset or compute budget.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

