Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

The bias–variance tradeoff describes how a model’s complexity can affect its predictions on new data: a model that is too restricted may miss important patterns, while one that is too flexible may fit quirks of its training examples. The practical goal is not to minimize training error, but to choose a model that performs well on data it did not use to learn.

What do bias and variance mean?

Bias is systematic error caused by assumptions or a model class that cannot represent the relevant pattern. A highly restricted model may make similar mistakes even when trained on different samples.

Variance describes how much a model’s predictions or fitted decision boundary change when the training sample changes. High variance means the learned result is sensitive to the particular examples it saw. It does not, by itself, say whether the predictions are correct. Stanford’s Information Retrieval text makes this distinction in classification and notes that high-variance methods can learn noise.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How underfitting and overfitting differ

Underfitting and overfitting are different ways a model can fail to generalize. They are useful diagnoses, not simple labels for “bad” and “good” complexity.

Underfitting: the model misses meaningful structure

Underfitting commonly occurs when the model is too restricted to capture a real pattern. For example, if the underlying relationship is curved, a straight line may systematically miss that curvature. This illustrates high bias; it is not a report of an experiment.

Overfitting: the model learns sample-specific detail

Overfitting occurs when a model fits details specific to its training examples—including noise—in a way that can harm predictions on new examples. A highly flexible curve might pass close to individual noisy observations, achieving low training error while changing substantially if the sample changes.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Why complexity can help, then hurt

In the classical teaching picture, increasing flexibility can initially reduce underfitting and improve performance on unseen data. Beyond some point, the model may become so sensitive to the training sample that generalization gets worse. Andrew Ng’s archived Stanford CS229 lecture transcript explains this familiar curve, with underfitting and high bias at one end and overfitting and high variance at the other.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That U-shaped curve is a useful baseline, not a law. A model’s parameter count, flexibility, or zero training error alone does not prove that it overfits. The balance depends on the task and data; Stanford’s Information Retrieval text cautions against a universal ranking of algorithms.

What the bias–variance formula does—and does not—say

For the familiar squared-error regression setup, expected prediction error can be decomposed into squared bias, variance, and irreducible noise, often written as:

Expected prediction error = bias² + variance + σ²

This decomposition belongs to that squared-error regression setting. It should not be treated as an identical formula for every loss function, classifier, or modern learning setup. Stanford’s MSE 125 chapter on validation and the bias–variance tradeoff covers the decomposition and evaluation roles.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to diagnose the tradeoff without fooling yourself

Compare candidate models on data kept outside the fitting process. Training performance describes how well a model fits examples it learned from; by itself, it cannot establish how well the model generalizes.

  • Poor training and validation performance: may indicate underfitting, but data quality, measurement noise, or a mismatch between evaluation data and the intended use can also explain poor results.
  • Much better training than validation performance: a large gap can be a warning sign of overfitting.
  • Unstable results across folds or samples: variation in performance can provide a clue that the fitted model is sensitive to which examples it receives.

Use validation data or cross-validation to choose among model complexities and settings. Keep a separate test set out of both fitting and model selection; use it for the final assessment. The Stanford MSE 125 chapter distinguishes training, validation, and test roles and discusses cross-validation.

A practical comparison checklist

When comparing real candidates, check the following together rather than relying on complexity alone:

  • Training and validation performance, including the size of their gap.
  • Performance across folds or repeated samples, as a clue to instability.
  • Model flexibility and regularization choices.
  • Validation performance on the metric that matters for the task.
  • Whether evaluation examples resemble the population where the model will be used.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why modern models complicate the classical picture

Some high-capacity models and datasets exhibit double descent: test risk may rise near the interpolation threshold and then fall again as capacity increases further. In their 2019 paper, “Reconciling modern machine learning practice and the bias-variance trade-off”, Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal describe this behavior as an extension of the textbook U-shaped account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

As the authors put it, “The classical thinking is concerned with finding the ‘sweet spot’ between under-fitting and over-fitting.” Double descent qualifies that picture: it challenges the claim that generalization must always worsen after a single optimum as capacity grows. It does not make validation unnecessary or make overfitting impossible.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.