Poor generative samples can signal several different problems: weak fidelity, missing variety, unstable training, or a mismatch between the model and its evaluation metric. For GANs, repeated outputs may indicate mode collapse; when later models are trained on earlier models’ synthetic outputs, a separate recursive-collapse risk arises. Diagnose the observable failure first, then use measures suited to that failure rather than treating one score as a complete verdict.
What “poor samples” can mean
A model’s outputs can be unconvincing for different reasons, and the visible symptom does not by itself establish the cause. Separate these questions:
- Fidelity: Do individual outputs look plausible or meet the task’s requirements?
- Coverage and diversity: Does the model represent the range of examples in the target data, or does it repeatedly produce a narrow subset?
- Training stability: Do training dynamics fluctuate or fail to settle?
- Memorization: Are outputs reproducing training examples rather than generalizing? Available metrics may not reliably distinguish this from other failures.
These are useful diagnostic categories, not a universal failure taxonomy shared identically by GANs, diffusion models, language models, and other model families.
Why GANs can produce repetitive or weak outputs
Generator–discriminator imbalance
In a GAN, the generator learns against a discriminator. If the discriminator becomes too strong, the generator may receive too little useful gradient information to improve. Google for Developers lists vanishing gradients among common GAN problems and describes proposed approaches such as Wasserstein or modified minimax losses, unrolled GANs, input noise, and discriminator weight penalties. These are attempts, not guaranteed fixes; the overview notes that the problems remain active research. Google for Developers’ GAN guide was updated August 25, 2025.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
Mode collapse
Mode collapse is a GAN training failure in which the generator produces the same output or a small set of output types instead of covering the variety in the target distribution. One possible training dynamic is that the generator over-optimizes against a particular discriminator while the discriminator fails to adapt out of a local trap. A few striking samples do not rule out collapse: the relevant question is whether the outputs collectively cover the range of the data.
Failure to converge
GAN losses can be unstable, and training may fail to converge. Loss curves are evidence about training behavior, but a single loss value should not be treated as a direct measure of sample quality or diversity. Inspect training behavior alongside generated outputs.
Why a single evaluation score can mislead
Some evaluation metrics compress performance into one number. That can obscure whether a model produces convincing samples while missing parts of the target distribution, or covers more of the distribution while producing weaker samples. Sajjadi and colleagues’ precision-and-recall framework separates sample quality from target-distribution coverage, addressing a limitation of one-dimensional scores such as FID. Google Research’s publication record describes the work, published at NeurIPS 2018.
Even a metric suited to a particular question has limits. Stein and colleagues’ NeurIPS 2023 study reported that, in its experimental setup, none of the evaluated metrics strongly correlated with human evaluations; human judgments of diffusion-model realism were not reflected by commonly reported metrics such as FID. The paper also found that feature-extractor choice and training procedure affect evaluation, and that current metrics did not reliably distinguish memorization from underfitting or mode shrinkage. These findings are reasons to interpret scores cautiously—not evidence that metrics are useless in every setting. Read the NeurIPS 2023 paper.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsA practical diagnostic sequence
- Name the visible symptom. Implausible content or artifacts point first to a fidelity question; repeated outputs or absent categories point to diversity or coverage; oscillating or worsening training behavior points to instability. This is a practical way to organize observations, not a validated universal decision tree.
- Review representative samples. Examine a broad, systematically selected set rather than only the best-looking examples. Compare outputs across relevant categories or groups, and record what the model fails to generate.
- Separate quality from coverage. Where suitable, report precision and recall alongside any overall score. Do not infer a unique cause from one scalar result.
- Check low-density and underrepresented cases. Look for groups or examples that are consistently poor or absent even when common cases look convincing. Lee, Kim, Hong, and Chung’s “Self-Diagnosing GAN” proposes using per-instance discrepancy statistics to identify and emphasize underrepresented data during GAN training; the authors reported improvements for minor groups in their experiments. This is a proposed GAN technique, not a general-purpose remedy. See the NeurIPS 2021 paper.
- For GANs, inspect training dynamics. Review discriminator and generator behavior, losses, and whether training appears to converge. Consider GAN-specific approaches to imbalance only in light of the training setup; no listed technique guarantees a fix.
- Trace the origin of training data. If synthetic outputs from earlier model generations have entered the training set, investigate that pipeline separately from a GAN’s training-time mode collapse.
Inspect what is missing, not just the average score
Visual audits can reveal absent content that a single aggregate score does not explain. Bau and colleagues’ ICCV 2019 work, “Seeing What a GAN Cannot Generate,” studies visual content that GANs fail to synthesize and frames missing-content diagnosis as a complement to scalar evaluation. In practice, look for omitted categories, attributes, or visual modes in addition to reviewing plausible examples. Read the ICCV 2019 paper.
Representation matters too: image metrics rely on feature extractors, and the NeurIPS 2023 study found evaluation results sensitive to encoder choice and training procedure. When possible, combine scores with representative human review and checks tailored to the task rather than assuming one feature space captures every meaningful failure.
Do not confuse GAN mode collapse with recursive model collapse
GAN mode collapse happens within an adversarial generator’s training dynamic: it loses output diversity. Recursive collapse is a different data-pipeline problem: successive generations of models are trained on data produced by earlier models, potentially losing parts of the original distribution over rounds. Shumailov and colleagues reported this phenomenon across language models, variational autoencoders, and Gaussian mixture models in a 2024 Nature paper. If synthetic data is being reused, audit its provenance and assess distribution loss across generations instead of treating the issue as ordinary GAN mode collapse. Read the 2024 Nature paper.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

