R² (R-squared) is the share of variation in an outcome variable that a fitted regression model accounts for. In simple linear regression, it equals the squared correlation, r². A value of 0.44, for example, means the model accounts for about 44% of the observed variation in y in that dataset—not that the predictor causes 44% of the outcome or that future predictions will be 44% accurate.
See R² in one picture
The clearest picture compares two ways to predict every observed y value. The mean-only baseline predicts the same average for everyone. The regression line uses x to produce a different fitted value for each case. R² measures how much the regression reduces squared error relative to that baseline.
For each observation, the total deviation from the mean can be split into an explained component (the fitted value’s departure from the mean) and a residual component (the observed value’s departure from its fitted value). Squaring and summing those distances gives the quantities in the formula.
The formula and its parts
| Quantity | Definition | Meaning |
|---|---|---|
| SST | Σ(yi − ȳ)² | Total squared variation around the observed mean |
| SSE | Σ(yi − ŷi)² | Squared error left after fitting the model |
| SSR | Σ(ŷi − ȳ)² | Squared variation accounted for by fitted values |
| R² | SSR/SST = 1 − SSE/SST | Fraction of total variation accounted for by the model |
These identities apply to ordinary least-squares regression with an intercept. R² has no units and is often reported as a percentage. In simple linear regression, R² = r², where r is the Pearson correlation.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
Worked example: why 0.4397 becomes 44%
OpenStax’s example uses grades for 11 students. The correlation between third-exam and final-exam grades is r = 0.6631, so r² = 0.4397. Rounded for communication, that is 44%.
The precise interpretation is: In this dataset and model, about 44% of the variation in final-exam grades is accounted for by third-exam grades. The remaining 56% is variation that this one-predictor regression does not account for. It may reflect other variables, measurement noise, random variation, or model limitations.
Rank #2
- This guide is a perfect overview for the topics covered in introductory statistics courses.
What R² does—and does not—tell you
It describes fit for a particular response, dataset and model
Always name the response variable. R² concerns variation in y, not a generic measure of “accuracy.” The same numerical value can be useful in one field and inadequate in another, depending on the purpose and the noise level that subject matter makes plausible.
It is not a causal percentage
“Explained” is associational language. A high R² does not show that changing x will cause a change in y. Causal claims require an appropriate design, assumptions and supporting subject-matter evidence.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #3
It is not a guarantee for new cases
Ordinary R² is calculated on the data used to fit the model. It can look strong while predictions deteriorate on new observations. Use a held-out test set or cross-validation when prediction is the goal.
Checks to make before trusting an R² value
- Plot the data. Look for curvature, clusters and unequal spread that a straight-line summary hides.
- Inspect residuals. A pattern, funnel shape or changing variance signals that the model form may be inadequate.
- Check unusual observations. A single influential or high-leverage point can materially change both r and R²; compare results with and without justified sensitivity checks.
- State the sample and scope. Report which observations, response variable and predictors produced the value.
R² in multiple regression
With several predictors, adding a variable cannot decrease in-sample R², even when the new variable contributes little practical value. Therefore compare models on the same response and dataset using more than one number:
Rank #4
| Question | Evidence to examine |
|---|---|
| How well does it fit the observed data? | R² plus residual plots and outlier diagnostics |
| Is the model needlessly complex? | Number of predictors, interpretability and subject-matter justification |
| Will it generalize? | Held-out or cross-validated performance when available |
| What is the model for? | Explanation, prediction and causal inference require different evidence |
Adjusted R² can be useful for comparing models with different numbers of predictors because it penalizes unnecessary additions, but it still does not replace diagnostics or out-of-sample evaluation.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.A practical interpretation checklist
- Identify the response variable y.
- State the dataset and fitted model.
- Convert R² to a percentage only after preserving its decimal value for calculations.
- Say “accounts for” or “is associated with,” not “causes.”
- Pair the statistic with a scatterplot, residual checks and, for prediction, validation results.
Used this way, R² is a compact description of how much better a fitted regression is than predicting the observed mean for everyone—valuable context, but not a complete assessment of a model.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

