Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Simple linear regression fits a straight line to describe the average relationship between one quantitative predictor and one quantitative response. The fitted line, ŷ = b₀ + b₁x, gives a predicted response for a value of x; residuals show how far the observed responses fall from those predictions. It summarizes an association—it does not, by itself, prove that changing x causes y to change.

What is simple linear regression?

Simple linear regression models the relationship between two quantitative variables. The predictor, also called the explanatory variable, is x; the response is y. “Simple” means the model uses one predictor. Penn State’s STAT 501 introduction to simple linear regression presents it as a method for studying relationships between continuous quantitative variables.

The fitted sample line is:

ŷ = b₀ + b₁x

Here, ŷ (“y-hat”) is the fitted or predicted response, b₀ is the intercept, and b₁ is the slope. The hat matters: ŷ is the model’s prediction, while y is the response actually observed. A population model is often written as a linear mean plus an error term; the fitted line estimates that mean relationship from sample data.

How does the fitted line get chosen?

Ordinary least squares chooses the intercept and slope to minimize the sum of squared residuals:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Σ(yᵢ − ŷᵢ)²

For each observation, the residual is eᵢ = yᵢ − ŷᵢ. Squaring each residual prevents positive and negative discrepancies from canceling in the total, and gives larger discrepancies more influence. The standard one-predictor line with an intercept has coefficient formulas:

  • b₁ = Σ[(xᵢ − x̄)(yᵢ − ȳ)] / Σ[(xᵢ − x̄)²]
  • b₀ = ȳ − b₁x̄

With an intercept, the fitted line passes through the sample means, (x̄, ȳ). These formulas describe the usual least-squares fit, rather than a guarantee that a straight-line model is appropriate for every dataset. Penn State explains the fitting criterion and coefficient calculations in its STAT 501 lesson and STAT 200 lesson on correlation and simple linear regression.

How do you interpret the slope and intercept?

Slope: fitted response change per predictor unit

The slope b₁ is the model’s predicted change in response for a one-unit increase in the predictor. Its units are response units divided by predictor units. For example, if x is measured in hours and y in dollars, the slope is in dollars per hour. State both the units and the data context when interpreting it: the slope describes an average fitted change over the context represented by the data, not a guaranteed change for every individual case.

Intercept: fitted response when the predictor is zero

The intercept b₀ is the fitted response at x = 0. That value may have little practical meaning if zero is impossible, irrelevant, or far outside the predictor values in the data. The intercept still helps define the mathematical line, but it should not be given a real-world interpretation the data cannot support.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Predictions and the data range

A fitted value is the line’s prediction for a chosen predictor value. Predictions within the range of observed x values are generally more grounded in the data than predictions far outside it. Using the line beyond that range is extrapolation: the observed fit alone does not establish that the same relationship continues there.

What is a residual?

A residual is the observed response minus its fitted value: eᵢ = yᵢ − ŷᵢ. Penn State STAT 200 defines it as an individual’s observed y value minus the corresponding predicted y value in its lesson on correlation and simple linear regression.

  • A positive residual means the observed response is above the line’s fitted value.
  • A negative residual means the observed response is below the fitted value.
  • The residual’s magnitude is the vertical distance between the observed point and the fitted line.

Residuals make the model’s misses visible. A line can summarize the overall direction of a relationship while still missing curvature, changing spread, or other structure in individual observations.

How do you check whether a linear regression model is reasonable?

For the usual introductory model and related inference, check four conditions often summarized as LINE: linearity, independence, normality of errors, and equal error variance. The checks are evidence about whether the model is a reasonable summary for the data; plots cannot prove that assumptions hold. Penn State describes these conditions and graphical diagnostics in its STAT 501 overview and lesson on simple linear regression assumptions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Check the scatterplot of x against y. Look for an approximately straight-line pattern. Clear curvature suggests that a straight line may not capture the relationship.
  2. Plot residuals against fitted values. Also consider residuals against x or observation order where relevant. Random-looking scatter around zero is more consistent with a suitable linear summary than a systematic shape.
  3. Look for changing spread. A widening or narrowing, fan-shaped pattern in residuals can signal unequal error variance.
  4. Check independence in context. Patterns by observation order, time, location, or another meaningful sequence can point to dependent errors. Whether observations should be independent depends in part on how the data were collected.
  5. Assess normality when inference calls for it. A normal probability plot or residual histogram can show whether residuals are approximately normal. This check is distinct from asking whether the scatterplot is linear.

Describe the pattern you see before deciding what to do about it. A curved residual pattern, changing spread, or ordered sequence can have different causes; the appropriate next step depends on the data collection and whether the goal is description, prediction, or inference.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Does regression show that one variable causes another?

No. A fitted slope and useful predictions describe an association in the data; they do not establish that changing the predictor causes the response to change. A causal conclusion needs an appropriate study design and assumptions beyond the fitted line. The instructional sources cited here explain relationships and model conditions, not a causal design.

When is a simple linear model useful?

A one-predictor straight-line fit can be a compact, interpretable baseline when the variables are quantitative and the scatter and residual checks support a roughly linear summary. A richer or more flexible model may be worth considering if the data show structure the line misses, or if the question requires more predictors. The choice depends on the data, the assumptions needed for the intended analysis, and whether the aim is explanation or prediction; one model cannot be called superior without a dataset and a stated objective.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.