Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A conditional GAN (cGAN) learns to generate an output that matches a supplied condition. The condition must reach both networks: the generator uses it to shape a sample, and the discriminator judges whether a sample is real given that same condition. Start with one task—such as generating labeled digits or translating paired images—then build and train the two networks around that task.

What makes a GAN conditional?

An unconditional GAN generates samples from noise. A cGAN also receives information that specifies what the output should represent. In the original 2014 formulation, Mehdi Mirza and Simon Osindero fed the conditioning data to both the generator and discriminator. Their paper illustrates class-conditioned MNIST digits and preliminary image-tagging examples. Read the original cGAN paper.

For a class-conditional image generator, the condition might be a digit label. For paired image translation, it might be a source image that the model should transform. Both are conditional generation, but they are different tasks: their data, architectures, and ways of representing the condition can differ.

Choose the task and condition before coding

For a first implementation, use a narrow task and ensure every training example has the correct condition attached. Decide whether you want to generate examples from class labels or map an input image to its paired target. The original cGAN paper provides the class-label case; TensorFlow’s pix2pix tutorial demonstrates paired image-to-image translation. See TensorFlow’s pix2pix tutorial.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Deep Learning (Adaptive Computation and Machine Learning series)
  • Language Published: English
  • Binding: hardcover
  • It ensures you get the best usage for a longer period

How you encode the condition depends on the task and architecture. A label must be represented in a form the networks can use; a source-image condition is itself image data. The defining requirement is not a particular embedding or concatenation technique, but that the generator and discriminator both receive the condition.

Build the two networks around the condition

Generator: noise plus condition to output

Give the generator a noise input and the condition, then have it produce a sample in the same format as the target training data. For class-conditioned generation, the output should reflect the requested class. For paired translation, it should correspond to the supplied source image. The original cGAN formulation establishes the condition’s role, but does not prescribe one universal way to combine it with noise.

Discriminator: sample plus condition to real or generated

Train the discriminator on pairs: real data with its correct condition, and generated data with the condition used to create it. It should learn whether the sample is real in the context of that condition. Supplying the condition only to the generator leaves out a core part of the original conditional adversarial setup.

Prepare data and output ranges consistently

Preprocessing and the generator’s final activation need to agree. For example, the PyTorch DCGAN tutorial scales images to the range [-1, 1] and uses tanh at the generator output. That is a documented DCGAN example, not a universal rule for cGAN datasets or output layers. Review the PyTorch DCGAN tutorial.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep each training example paired with the correct label or input condition throughout batching and augmentation. If those associations are wrong, the discriminator is trained on misleading real pairs and the generator receives inconsistent targets.

Train with alternating adversarial updates

A practical training loop alternates between updating the discriminator and updating the generator. The PyTorch tutorial’s DCGAN example computes discriminator losses for real and generated samples, updates the discriminator, then updates the generator to make generated samples receive the real target. It uses separate optimizers for the two networks.

  1. Prepare a batch: load real samples and their conditions, and draw noise for the generator.
  2. Update the discriminator: evaluate real sample-condition pairs against the real target, generate fake samples, evaluate those with their conditions against the fake target, combine the losses, and update the discriminator.
  3. Update the generator: generate samples from noise and conditions, evaluate them with the discriminator, and train the generator using the real target so it learns to make the discriminator classify its outputs as real.
  4. Track progress: record both losses and generate samples from a fixed set of noise inputs and conditions so outputs can be compared during training.

These steps describe the adversarial structure, not a guarantee of convergence. The PyTorch tutorial notes that practical GANs do not always reach the ideal theoretical equilibrium; the balance between the two networks can require experimentation.

Use documented settings as a starting point, not a prescription

The PyTorch DCGAN tutorial uses binary cross-entropy, real targets of 1 and fake targets of 0, and two Adam optimizers. Its documented example settings are a learning rate of 0.0002 and beta1 = 0.5; the tutorial was last updated on January 19, 2024, and last verified on November 5, 2024. These values belong to that DCGAN example and are not established as optimal for a different cGAN task, dataset, architecture, or training scale.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Deep Learning: A Visual Approach
  • Deep Learning: A Visual Approach
  • No Starch Press
  • ABIS BOOK

The tutorial also describes the original minimax objective and the common practical choice of maximizing log(D(G(z))) for the generator rather than minimizing log(1-D(G(z))). The alternative gives a stronger gradient early in training. It is a training choice, not a promise that adversarial training will converge.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose an architecture that fits the task

Use case Condition Documented architectural direction Important fit
Class-conditioned small images A class label The original cGAN paper establishes feeding labels to both networks; a convolutional GAN is a practical image baseline, with label embeddings or other combinations left to the implementation. Use examples with reliable labels and an output format matching the image data.
Paired image-to-image translation A source image TensorFlow’s pix2pix tutorial uses a U-Net-based generator and a convolutional PatchGAN discriminator. Use aligned source-target pairs for the intended mapping.

Neither source establishes a universal architecture winner across condition types, output tasks, data alignment, resolution, or compute needs. Pix2pix’s U-Net and PatchGAN are choices for paired translation, not requirements for every cGAN.

Evaluate outputs by condition and plan for iteration

Use fixed noise inputs and vary the conditions to see whether outputs change in the intended way; also compare generated samples over training. This makes it easier to notice, for example, whether a class label appears to influence the output. Visual inspection is useful for monitoring, but by itself does not establish model quality. Consider the losses alongside the samples, and expect to adjust the model and training setup.

A GPU, or two, can help with the PyTorch tutorial’s example, but that does not establish a minimum hardware requirement for a small cGAN project. Compute needs and runtime depend on dataset size, resolution, model, and acceptable training time; no fixed hardware minimum or training-time estimate is established here.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

SaleBestseller No. 1
Deep Learning (Adaptive Computation and Machine Learning series)
Deep Learning (Adaptive Computation and Machine Learning series)
Language Published: English; Binding: hardcover; It ensures you get the best usage for a longer period
$51.51
SaleBestseller No. 2
SaleBestseller No. 5
Deep Learning: A Visual Approach
Deep Learning: A Visual Approach
Deep Learning: A Visual Approach; No Starch Press; ABIS BOOK
$73.40

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.