Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a regression metric according to the cost and pattern of prediction errors, the target’s scale, and the decision you need to make. There is no universally best metric: MAE describes typical absolute error, RMSE gives larger misses extra influence, and R² compares a model with a mean-prediction baseline. Because these metrics measure different things, they can rank the same models differently.

How to choose a regression metric

Start by asking what kinds of mistakes matter. If stakeholders need to understand the typical miss in familiar units, use an absolute-error metric. If unusually large misses are especially costly, consider a squared-error metric. If proportional error matters more than the number of target units, a relative-error metric may help—but only if the target values make its denominator meaningful.

For a compact comparison, report one interpretable error measure in the target’s units alongside R², making clear that R² is relative to a baseline and evaluation dataset. Add another metric when it answers a distinct question, rather than presenting a collection of scores without explaining their purpose.

Why models can rank differently

Metrics weight errors differently. For example, suppose one model makes errors of 1 and 9, while another makes errors of 5 and 5. Both have an MAE of 5. The first model has a larger RMSE because its error of 9 is squared before averaging. A metric’s ranking is therefore a consequence of what it emphasizes, not necessarily a contradiction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

MAE, MSE, and RMSE: absolute versus squared error

Metric What it summarizes Units Best suited to Main caveat
MAE Mean absolute prediction error Same units as the target Communicating typical absolute miss Large errors do not receive the extra emphasis they get under squared error.
MSE Mean squared prediction error Squared target units Giving large errors disproportionate weight, or evaluating a squared-error objective Squared units are less intuitive to interpret.
RMSE Square root of MSE Same units as the target Keeping squared-error sensitivity while reporting on the target’s scale Large errors still influence it more than they influence MAE.

MAE: average absolute miss

Mean absolute error (MAE) averages the absolute differences between predictions and true values. It is reported in the target’s units, so a forecast with MAE of 3 means an average absolute error of 3 target units on the evaluated samples. MAE does not square errors, making it a straightforward choice when the typical size of a miss is the main question.

MSE and RMSE: greater weight on large misses

Mean squared error (MSE) squares each error before averaging. As a result, a large miss contributes much more than a small one; this is useful when that pattern matches the cost of mistakes or the model’s squared-error objective. Its units are squared, however, which can make the score difficult to explain directly.

Root mean squared error (RMSE) is the square root of MSE, returning the result to the target’s units while retaining the influence of squared errors. The scikit-learn regression metrics guide describes RMSE as a common measure in the same units as the target variable.

R²: a comparison with a mean-prediction baseline

R² summarizes residual squared error relative to variation in the target on the evaluation data. In scikit-learn’s framing, an R² of 0 corresponds to predicting the evaluation target’s mean as a constant. A negative R² means the model performs worse than that reference under the R² calculation. It is not a percentage-accuracy score, and it is tied to the dataset used to calculate it; R² values from different datasets may not be comparable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use R² when the mean-prediction baseline is a useful reference, and state what data produced the score. Pair it with an error metric in target units when readers also need to know how far predictions are off in practical terms.

MAPE: relative error, with a denominator warning

Mean absolute percentage error (MAPE) expresses absolute errors relative to the true values, which can make it useful when proportional misses matter more than misses of a fixed size. In concept, it is invariant to scaling the target globally. But because actual values are in the denominator, zero and near-zero targets can make the score unstable or hard to interpret.

In scikit-learn, MAPE is returned as a relative fraction, not as a value from 0 to 100. Thus, 0.2 corresponds to 20% when expressed as a conventional percentage. The implementation uses a small positive epsilon to protect against division by zero; that protection does not make percentage-based interpretation reliable for zero or near-zero actuals. See the scikit-learn MAPE guidance before interpreting such scores.

When to consider MedAE or MSLE

MedAE when outliers distort an average

Median absolute error (MedAE) takes the median of absolute errors rather than their mean. It is less affected by unusually large errors than MAE, so it can describe the central or typical miss when outliers would distort a mean-based summary. It does not describe tail risk: a good MedAE alone does not show how severe the largest errors are.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

MSLE for suitable nonnegative, growing targets

Mean squared logarithmic error (MSLE) measures squared differences in log(1 + target) space. It may fit nonnegative targets that grow across orders of magnitude, when relative growth is more relevant than absolute distance in the original scale. Its penalties are asymmetric: the scikit-learn guide notes that MSLE penalizes under-prediction more than over-prediction. Confirm that both the target domain and that asymmetry make sense for the application before using it.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Specialized losses depend on the task

Scikit-learn also provides losses such as Poisson, Gamma, and Tweedie deviance, as well as pinball loss. Their availability does not make them generally preferable: consider them when the target distribution or objective—such as estimating a quantile—calls for that kind of loss. The regression metrics guide and metrics API reference describe the available options; they do not establish which is appropriate for a particular dataset.

Make scores interpretable and comparable

State how the evaluation data was chosen

A metric is a score on particular samples, not a context-free property of a model. Say whether results come from a held-out test set or cross-validation, and apply the same evaluation protocol when comparing models. The scikit-learn guide discusses metrics and scoring parameters in the context of evaluation and model-selection tools.

Inspect each output when predicting multiple targets

For multiple target variables, inspect per-target scores when scales or business importance differ. Many supported metrics use a uniform average across outputs by default; that average can conceal a poor result on one target or give an unimportant target equal influence. Use explicit weights when they reflect a justified priority, or report raw per-target values. See the scikit-learn multioutput regression guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use this checklist before reporting

  • Identify whether absolute miss, large-error risk, or relative error matters most.
  • Choose a metric whose units and scale readers can interpret.
  • Check whether outliers or zero and near-zero targets affect the metric you selected.
  • For R², identify the evaluation set and mean-prediction baseline.
  • For multiple targets, report per-target results or explain the weights used.
  • State whether scores come from a holdout set or cross-validation.
  • When useful, present complementary views—for example, MAE in target units alongside R²—without treating any score as a universal accuracy threshold.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.