Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use regression in Python by separating three goals: predicting a numeric value, estimating relationships, or testing statistical hypotheses. A practical workflow is to prepare data in a leakage-safe pipeline, fit an ordinary least squares (OLS) baseline, validate it on unseen data, inspect residual and collinearity diagnostics, and then compare regularized or nonlinear models when the baseline is inadequate. Use scikit-learn for predictive pipelines and cross-validation; use statsmodels when coefficient estimates, standard errors, hypothesis tests, covariance structures, and diagnostic output matter.

What regression analysis does

Regression models a numeric outcome from one or more predictors. For an observation with features x, a linear model has the form ŷ = β₀ + β₁x₁ + … + βₚxₚ. The same equation can serve different purposes:

  • Prediction: estimate accurate values for future or otherwise unseen cases.
  • Explanation: describe how an outcome changes with predictors, conditional on the model and data.
  • Inference: quantify uncertainty and test hypotheses about coefficients or error structures.

These goals are not interchangeable. A model with excellent test-set error may be difficult to interpret, while a statistically interpretable model may predict poorly outside the data used to fit it.

Prepare the data before fitting a model

Inspect the outcome and predictors

  • Confirm that the target is numeric and that each row represents the intended observational unit.
  • Check data types, impossible values, duplicate records, missingness patterns, and the scale of each numeric feature.
  • Identify categorical columns and choose an encoding strategy rather than converting category labels to arbitrary numbers.
  • Investigate extreme observations. An outlier can be a valid rare case, a data error, or a point with disproportionate influence.

Prevent leakage

Any transformation that learns from data—imputation values, category levels, scaling parameters, feature selection, or dimensionality reduction—must be fitted only on the training portion of each split. A scikit-learn pipeline keeps those operations together and applies the same learned transformations to validation and test data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.

A reproducible preprocessing pipeline

from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

numeric = ["age", "income"]
categorical = ["region", "plan"]

numeric_pipe = Pipeline([
    ("impute", SimpleImputer(strategy="median")),
    ("scale", StandardScaler()),
])
categorical_pipe = Pipeline([
    ("impute", SimpleImputer(strategy="most_frequent")),
    ("encode", OneHotEncoder(handle_unknown="ignore")),
])

preprocess = ColumnTransformer([
    ("num", numeric_pipe, numeric),
    ("cat", categorical_pipe, categorical),
])

Scaling is essential for penalized models because ridge and lasso penalties depend on coefficient size. It is not required for ordinary least squares to be mathematically valid, but it can improve numerical conditioning and makes coefficient magnitudes easier to compare when predictors use different units.

Fit an OLS baseline

scikit-learn for a predictive baseline

scikit-learn’s LinearRegression fits coefficients by minimizing the residual sum of squares between observed targets and targets predicted by a linear approximation. The estimator includes an intercept by default.

from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
import numpy as np

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

model = Pipeline([
    ("preprocess", preprocess),
    ("regressor", LinearRegression()),
])
model.fit(X_train, y_train)
pred = model.predict(X_test)

print("MAE:", mean_absolute_error(y_test, pred))
print("RMSE:", np.sqrt(mean_squared_error(y_test, pred)))
print("R²:", r2_score(y_test, pred))

Keep the test set untouched until model selection is complete. If observations are time-ordered, use a time-based split rather than a random split. If rows are grouped by customer, patient, device, or location, use a group-aware split so related records cannot appear on both sides.

Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

statsmodels for coefficient inference

statsmodels expresses the model as Y = Xβ + ε, with an assumed error distribution and covariance structure. Its fitted results object provides a statistical summary, coefficient estimates, standard errors, confidence intervals, tests, and fit statistics.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import statsmodels.api as sm

# X_stat must contain the numeric/encoded columns used in the model.
X2 = sm.add_constant(X_stat)
result = sm.OLS(y, X2).fit()
print(result.summary())
print(result.conf_int())

Unlike a scikit-learn pipeline, statsmodels.OLS does not automatically encode categories, impute missing values, or add an intercept. Prepare those elements explicitly. statsmodels also documents weighted least squares (WLS), generalized least squares (GLS), and GLS with autoregressive errors (GLSAR) when the error variance or covariance is not adequately represented by ordinary least squares.

Validate predictions out of sample

Choose a split that matches deployment

A single holdout gives a straightforward final check; cross-validation uses several training/validation partitions and is usually more stable for model selection. Do not use the test set repeatedly to tune features or hyperparameters.

Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
from sklearn.model_selection import KFold, cross_validate
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge

ridge = make_pipeline(StandardScaler(), Ridge(alpha=1.0))
cv = KFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_validate(
    ridge, X, y, cv=cv,
    scoring={"mae": "neg_mean_absolute_error",
             "rmse": "neg_root_mean_squared_error",
             "r2": "r2"}
)
print(-scores["test_mae"].mean())
print(-scores["test_rmse"].mean())
print(scores["test_r2"].mean())

For time series, replace random K-fold with a forward-looking split. For grouped data, use a group-aware cross-validator and pass the group labels.

Pick a metric that reflects the decision

Metric Interpretation Useful when Important limitation
MAE Average absolute error in target units Each error should contribute linearly and the result must be easy to explain Does not emphasize large errors
RMSE Square-root average of squared errors, in target units Large mistakes are especially costly More sensitive to outliers
R² Relative improvement over predicting the training mean, evaluated on the chosen data Comparing explained variation under a common target and split Can be negative out of sample and is not an error magnitude
Median absolute error Median absolute prediction error Robust reporting when a few extreme errors should not dominate Can conceal tail risk
MAPE Percentage error relative to the actual value Positive targets where relative error is the business quantity Unstable or undefined near zero

Report the primary metric in the units and cost structure that matter, and add a secondary metric to expose a different failure mode. A high R² alone does not establish useful accuracy.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Diagnose assumptions before interpreting coefficients

Residual patterns and nonlinearity

Plot residuals against fitted values and important predictors. A curve, funnel, clusters, or changing spread indicates that a straight-line, constant-variance specification may be inadequate. Consider transformations, interactions, splines, or a nonlinear estimator; do not infer causality from a residual plot.

Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft

Heteroscedasticity

If residual variance changes with the fitted value or a predictor, ordinary standard errors can be misleading. Re-express the outcome, model the variance, use heteroscedasticity-robust covariance estimates where appropriate, or use WLS when the variance structure is known or defensibly estimated.

Autocorrelation

Sequential observations can have correlated errors. Randomly splitting them can also leak future information. Use time-aware validation and a covariance-aware model such as GLS or GLSAR when the dependence is part of the data-generating process.

Influential observations

Inspect leverage, studentized residuals, and Cook’s distance. Verify influential rows against the source data, run a sensitivity analysis with and without them, and document any exclusion. Removing a point solely because it changes the answer is not a valid diagnostic procedure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

Multicollinearity

Highly correlated predictors make least-squares coefficient estimates sensitive and can produce high variance, even when predictions remain reasonable. Examine a correlation matrix, condition measures, or variance inflation factors alongside domain knowledge. Do not interpret an unstable individual coefficient as a reliable isolated effect.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Compare OLS, ridge, lasso, and nonlinear models

Model How it differs Strengths Trade-offs
OLS Minimizes residual sum of squares without a coefficient penalty Simple, fast, and natural for statistical inference when assumptions are reasonable Sensitive to collinearity, influential points, and misspecification
Ridge Adds an L2 penalty; larger alpha shrinks coefficients more Stabilizes correlated predictors and often improves generalization Does not generally set coefficients exactly to zero; penalty complicates classical inference
Lasso Adds an L1 penalty Can produce sparse coefficients and a compact feature set Selection can be unstable with correlated features; scaling and regularization choice matter
Polynomial or spline regression Adds nonlinear basis terms while retaining a regression framework Captures smooth curvature and can remain interpretable More features increase overfitting risk and require validation
Tree-based regression Partitions feature space rather than fitting one global line Captures interactions and nonlinearities with little feature transformation Less direct coefficient interpretation; tuning and extrapolation require care
from sklearn.linear_model import Ridge, Lasso
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

ridge = make_pipeline(StandardScaler(), Ridge(alpha=1.0))
lasso = make_pipeline(StandardScaler(), Lasso(alpha=0.01, max_iter=10000))

Select alpha with cross-validation inside the training data. Compare models using the same splits, preprocessing rules, target definition, and primary metric.

A defensible end-to-end workflow

  1. Define the decision: specify whether the goal is prediction, explanation, inference, or a combination.
  2. Audit the data: document the target, observation unit, time boundaries, missingness, categories, outliers, and potential leakage.
  3. Create an appropriate split: hold out future, grouped, or otherwise representative data before tuning.
  4. Build preprocessing in a pipeline: fit imputers, encoders, scalers, and selectors only on training folds.
  5. Fit OLS: use scikit-learn for a predictive baseline and statsmodels when inferential output is required.
  6. Measure out of sample: report a decision-relevant metric and uncertainty across validation folds where possible.
  7. Diagnose: inspect residuals, variance, dependence, influence, and collinearity before interpreting coefficients.
  8. Compare alternatives: tune ridge, lasso, polynomial, tree-based, or other models on the same evaluation design.
  9. Finalize once: evaluate the selected procedure on the untouched test set, then retrain on permitted data for deployment.
  10. Monitor: track input drift, missingness, residual behavior, and production error after release.

Common failure modes

  • Encoding categories as 1, 2, 3: this imposes a false numeric distance; use one-hot or an explicitly justified ordinal encoding.
  • Scaling before cross-validation: this lets validation information influence training; put scaling inside the pipeline.
  • Reading causation from coefficients: regression associations can reflect confounding, selection, or measurement error.
  • Using training R² as proof of performance: training fit is optimistic; use held-out predictions.
  • Dropping outliers automatically: validate the observation and perform sensitivity analysis first.
  • Interpreting penalized coefficients like OLS estimates: regularization intentionally introduces bias to control variance, so classical standard-error interpretations do not transfer directly.

Choosing between scikit-learn and statsmodels

Need Better starting point Reason
Preprocessing, pipelines, cross-validation, hyperparameter tuning, and production prediction scikit-learn Consistent estimator API and model-selection tooling
Coefficient tables, standard errors, tests, confidence intervals, and covariance-aware regression statsmodels Statistical results objects and diagnostics
Both reliable inference and predictive evaluation Use both Understand and diagnose a statistical specification with statsmodels, then evaluate a leakage-safe predictive pipeline with scikit-learn

Frequently Asked Questions

Do I need to standardize variables for linear regression in Python?

Standardization is not required for ordinary least squares itself, but it is important for ridge and lasso because their penalties depend on coefficient scale. Put scaling inside the pipeline so it is learned within each training fold.

Why can test-set R² be negative?

R² compares your predictions with the baseline that always predicts the evaluation-set mean. A negative value means the model performed worse than that baseline on those unseen observations.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can regression coefficients prove that one variable causes another?

No. A coefficient describes a conditional association under the model. Causal interpretation requires an appropriate design, identification assumptions, and control of confounding beyond ordinary regression output.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.