Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

An AI loss function is a mathematical rule that assigns a numerical penalty to a model’s prediction based on how it differs from the target—or, more generally, how well it meets a training objective. Training uses an optimization algorithm to adjust the model’s parameters and reduce that loss. The loss defines what the model is being encouraged to do; it does not, by itself, establish whether the model is useful in the real world.

How a loss function scores a prediction

For a supervised learning example, the model produces a prediction and the training data supplies a target value or label. The loss function compares them and returns a score. A model that makes predictions favored by the chosen objective receives a lower loss than one that makes predictions the objective penalizes.

There can be a loss for each example. Training commonly combines losses across a batch or dataset, often by averaging them, to form the objective the optimizer works to reduce. Google for Developers describes this minimization goal in its Machine Learning Glossary. The loss function defines the score; an optimization algorithm, such as a gradient-based method, updates the model parameters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A simple regression example

Suppose a model predicts a house price of 210,000 when the target is 200,000. With squared error, the difference is squared, so this example contributes 100,000,000 squared currency units before any averaging or scaling. A miss of 1,000 would contribute 1,000,000. The larger miss therefore has 100 times the squared-error contribution, even though it is only 10 times as large in absolute terms.

This illustrates why the loss is a design choice: it determines which kinds of errors matter more during training. The numerical loss is not necessarily expressed in an intuitive unit, and its scale alone does not tell you how valuable the model is.

Common loss functions and when they differ

Loss and typical task What it measures Practical distinction
Mean squared error (MSE), also called squared error or L2 loss; regression The average of squared differences between predicted and target numeric values. Squaring gives large errors disproportionate influence. It may suit objectives where large misses deserve especially strong penalties, but outliers can dominate. See Google’s explanation of linear-regression loss and scikit-learn’s MSE definition.
Mean absolute error (MAE), also called absolute error or L1 loss; regression The average absolute difference between predicted and target numeric values. It is less sensitive to outliers than MSE and corresponds more directly to average error magnitude in the target’s units. It does not give unusually large errors the same extra weight that squaring does. See Google’s comparison of MAE and MSE.
Cross-entropy; commonly used for classification A penalty based on predicted class probabilities and the target labels. The label format expected by the implementation matters, as do settings that determine whether losses are summed, averaged, or otherwise reduced. See the PyTorch 2.14 CrossEntropyLoss documentation and OpenStax’s backpropagation discussion.

How to choose and interpret a loss

There is no single loss function that is best for every AI task. Start with the task and the consequences of different mistakes, then check that the loss matches the data and the labels your training code can provide.

  • For numeric prediction, decide whether large errors should receive extra weight. MSE does this; MAE treats errors in proportion to their absolute size.
  • For classification, a probability-based objective such as cross-entropy is common, but the framework’s input and target requirements must match your data.
  • Check scale and meaning. A loss value is useful for tracking the stated objective, but its magnitude may not correspond to an intuitive real-world error.
  • Evaluate separately. Training loss and evaluation metrics are related but need not be the same. Use measures suited to the task—such as error in meaningful units or classification performance—and assess them on data not used to fit the model. A falling training loss alone does not prove that the model generalizes or is useful.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What a loss function does not tell you

A loss function is not the optimizer, a universal measure of model quality, or a guarantee that the model’s predictions are correct. It formalizes a chosen training objective. Whether that objective captures the real costs of errors—and whether the trained model performs well outside its training data—requires separate evaluation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.