Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Backward feature elimination starts with all candidate predictors and removes them one at a time until a stopping rule is met. The term covers several different methods: statistical elimination based on p-values, predictive selection based on cross-validation, and recursive feature elimination (RFE) based on an estimator’s feature-importance ranking. They can choose different subsets because they optimize different goals.

Use p-value elimination for carefully specified statistical models, backward sequential selection when the goal is predictive validation performance, and RFE or RFECV when an estimator’s importance ranking is appropriate. Whichever method you choose, learn the selected features from training data only and evaluate the complete selection process on data it did not use.

What backward feature elimination does

Feature selection keeps a subset of the original variables. It differs from feature extraction, which creates a new representation such as principal components; feature engineering, which creates variables from existing data; and regularization, which penalizes coefficients while fitting a model rather than necessarily removing variables outright. Backward elimination is a model-based or wrapper feature-selection strategy: it repeatedly fits or evaluates a model as features are removed.

Removing predictors can make a model cheaper to train and serve, easier to interpret, smaller to store, or less dependent on costly-to-collect data. It may improve generalization when irrelevant variables encourage overfitting, but it can also hurt performance when weak predictors work together. Selection itself can overfit the validation data, so improvement is never guaranteed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.

Which method should you use?

Method How it removes features Best starting point Main limitation
Statistical backward elimination Removes the predictor with the largest p-value above a chosen threshold, then refits the statistical model. A carefully specified explanatory model, such as OLS regression or a generalized linear model. Repeated selection complicates ordinary p-value interpretation; assumptions and collinearity matter.
Backward sequential feature selection Tests candidate removals and removes the one that produces the best cross-validated estimator score. Scikit-learn documents this greedy procedure in its SequentialFeatureSelector API. Predictive model selection when you have a target feature count and an appropriate validation metric. Can be computationally expensive and can overfit the validation process.
RFE Fits an estimator, ranks features using its coefficients or feature importances, and removes the least important. See scikit-learn’s RFE documentation. When the estimator provides a meaningful importance ranking. Ranking depends on the estimator and its representation of importance.
RFECV Uses recursive elimination and cross-validation to choose a feature count. See RFECV. When you want validation to choose the number of features. More computationally costly than a single fit; validation design must match the data.
Lasso or Elastic Net Fits a regularized model that shrinks coefficients; Lasso can set some to zero. High-dimensional or correlated predictors where regularization is a suitable modeling choice. Regularization is not the same procedure as backward elimination; scaling and tuning matter.
Forward selection Starts with no predictors and adds variables according to a chosen criterion. Situations where fitting the full candidate model is impractical. It explores a different path and need not find the same subset as backward methods.

Scikit-learn’s feature-selection guide compares methods and notes that backward sequential selection may require roughly m × k model fits for one removal step with m current features and k-fold cross-validation. RFE can obtain a ranking from a single estimator fit per elimination step, though total cost depends on the estimator, number of steps, and configuration.

How the elimination loop works

P-value version

  1. Begin with the candidate predictors you have decided are appropriate for the model.
  2. Fit the statistical model and inspect the predictors’ p-values.
  3. Find the largest p-value among features eligible for removal.
  4. If it exceeds the chosen threshold, remove that predictor and refit. Otherwise, stop.
  5. Stop as well if a predetermined minimum feature count is reached, then inspect the final model and the elimination history.

For example, suppose a model begins with five predictors and the largest eligible p-value belongs to x4. If that p-value exceeds the specified threshold, remove x4, refit using the other four predictors, and reassess. The next model’s p-values must be recomputed; they are not the original model’s values.

Predictive backward version

  1. Start with all eligible predictors.
  2. For each remaining predictor, temporarily remove it and evaluate the resulting subset using cross-validation.
  3. Keep the removal that gives the best score under the selected metric.
  4. Continue until the target feature count or other stopping rule is reached.

The objective is predictive score, not statistical significance. A feature with a large p-value may still improve cross-validated prediction, while a statistically significant coefficient need not materially improve a chosen predictive metric.

Statistical backward elimination with OLS in Python

The following function removes the eligible predictor with the largest p-value while it exceeds alpha. It allows protected columns, enforces a minimum feature count, and records a removal log. It uses statsmodels’ OLS; fitted-result p-values are exposed through RegressionResults.pvalues.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import numpy as np
import pandas as pd
import statsmodels.api as sm


def backward_elimination_pvalues(
    X,
    y,
    alpha=0.05,
    keep=None,
    min_features=1,
    verbose=True,
):
    """Remove the eligible predictor with the largest p-value."""
    if not isinstance(X, pd.DataFrame):
        X = pd.DataFrame(X)

    if X.columns.duplicated().any():
        raise ValueError("X contains duplicate column names.")

    features = list(X.columns)
    protected = set(keep or [])
    missing_protected = protected.difference(features)
    if missing_protected:
        raise ValueError(
            f"Protected columns are not present in X: {missing_protected}"
        )
    if min_features < 1:
        raise ValueError("min_features must be at least 1.")

    history = []
    while len(features) > min_features:
        X_model = sm.add_constant(X[features], has_constant="add")
        model = sm.OLS(y, X_model, missing="drop").fit()
        pvalues = model.pvalues.drop(labels="const", errors="ignore")
        removable = pvalues.drop(labels=list(protected), errors="ignore")
        if removable.empty:
            break

        worst_feature = removable.idxmax()
        worst_pvalue = removable.loc[worst_feature]
        if not np.isfinite(worst_pvalue) or worst_pvalue <= alpha:
            break

        history.append({
            "removed_feature": worst_feature,
            "p_value": worst_pvalue,
            "features_before": len(features),
            "adjusted_r_squared": model.rsquared_adj,
            "aic": model.aic,
            "bic": model.bic,
        })
        if verbose:
            print(f"Removing {worst_feature!r}; p-value={worst_pvalue:.6g}")
        features.remove(worst_feature)

    final_model = sm.OLS(
        y,
        sm.add_constant(X[features], has_constant="add"),
        missing="drop",
    ).fit()
    return features, final_model, pd.DataFrame(history)

Apply it to training data, not the full dataset:

selected_features, final_model, elimination_log = (
    backward_elimination_pvalues(
        X_train,
        y_train,
        alpha=0.05,
        min_features=3,
        keep=["treatment"],
    )
)

print(selected_features)
print(final_model.summary())
print(elimination_log)

The example threshold is a parameter, not a universal recommendation. A fixed p-value cutoff is only one stopping rule. Depending on the aim, compare nested models with likelihood-ratio tests where appropriate, consider AIC or BIC, set a feature count in advance, preserve study-mandated adjustment variables, or assess stability by repeating selection on resampled training data. Evaluate prediction performance separately if prediction is also a goal.

Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

What the p-values do and do not mean

A p-value is evidence about a model-specific hypothesis under the model’s assumptions. A p-value above a threshold does not establish that a variable has no practical value, is causally irrelevant, or will be useless in another model. It also does not rule out importance in an interaction or joint contribution with correlated variables.

Because the same data are used repeatedly to choose and refit the model, the final p-values are affected by data-dependent selection. They are not ordinary, untouched p-values from a model specified before seeing the data. Treating them as definitive evidence without accounting for post-selection inference can overstate certainty.

Predictive backward selection with scikit-learn

Regression example

For regression, use a scoring rule aligned with the practical objective. Scikit-learn loss scorers use negated values so that larger scores are better.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from sklearn.feature_selection import SequentialFeatureSelector
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import KFold
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

estimator = Pipeline([
    ("scale", StandardScaler()),
    ("model", LinearRegression()),
])
cv = KFold(n_splits=5, shuffle=True, random_state=42)

selector = SequentialFeatureSelector(
    estimator=estimator,
    n_features_to_select=10,
    direction="backward",
    scoring="neg_mean_squared_error",
    cv=cv,
    n_jobs=-1,
)
selector.fit(X_train, y_train)
selected_features = X_train.columns[selector.get_support()].tolist()
print(selected_features)

Classification example

For a binary classification problem, use a metric suited to the decision. ROC AUC or average precision may be more informative than accuracy when classes are imbalanced; use a cost-sensitive or class-specific metric when the consequences of errors differ.

from sklearn.feature_selection import SequentialFeatureSelector
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

estimator = Pipeline([
    ("scale", StandardScaler()),
    ("model", LogisticRegression(max_iter=2000)),
])
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)

selector = SequentialFeatureSelector(
    estimator=estimator,
    n_features_to_select=10,
    direction="backward",
    scoring="roc_auc",
    cv=cv,
    n_jobs=-1,
)
selector.fit(X_train, y_train)
selected_features = X_train.columns[selector.get_support()].tolist()
print(selected_features)

Other scoring options include neg_mean_absolute_error or r2 for regression; average_precision, f1, or accuracy for classification; and neg_log_loss for probability quality. Do not choose a metric simply because it is familiar: it should reflect the intended use and the cost of different errors.

Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.

RFE and RFECV in Python

RFE is related to backward elimination but is not p-value elimination or backward sequential selection. It repeatedly fits an estimator and removes features according to that estimator’s importance attributes, such as coef_ or feature_importances_.

from sklearn.feature_selection import RFE
from sklearn.linear_model import LogisticRegression

estimator = LogisticRegression(max_iter=2000, solver="liblinear")
selector = RFE(
    estimator=estimator,
    n_features_to_select=10,
    step=1,
)
selector.fit(X_train, y_train)

selected_features = X_train.columns[selector.support_].tolist()
ranking = dict(zip(X_train.columns, selector.ranking_))
print(selected_features)
print(ranking)

Use RFECV when cross-validation should choose the number of features rather than having you set it directly. Its documented default when cv=None is five-fold cross-validation; choose an explicit splitter when the data structure requires one.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from sklearn.feature_selection import RFECV
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold

estimator = LogisticRegression(max_iter=2000)
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
selector = RFECV(
    estimator=estimator,
    step=1,
    min_features_to_select=1,
    scoring="roc_auc",
    cv=cv,
    n_jobs=-1,
)
selector.fit(X_train, y_train)

selected_features = X_train.columns[selector.support_].tolist()
print(selector.n_features_)
print(selector.cv_results_["mean_test_score"])

These rankings are estimator-dependent. Linear coefficients and tree importances are not interchangeable measures; impurity-based tree importance can favor continuous or high-cardinality variables. Neither should automatically be read as causal importance.

Prevent leakage during selection and evaluation

Do not fit a selector on the whole dataset before splitting it. That lets information from the eventual test set influence which predictors are retained, making the test score optimistic.

# Incorrect: selection sees the eventual test observations.
selector.fit(X, y)
X_reduced = selector.transform(X)
X_train, X_test, y_train, y_test = train_test_split(
    X_reduced, y, test_size=0.2, random_state=42
)

At minimum, split first, fit selection only on training data, and apply the fitted selector to the holdout:

Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)
selector.fit(X_train, y_train)
X_train_selected = selector.transform(X_train)
X_test_selected = selector.transform(X_test)
final_model.fit(X_train_selected, y_train)
predictions = final_model.predict(X_test_selected)

For cross-validation or tuning, put preprocessing and selection inside the pipeline passed to the validation process, so each training fold learns its own transformations and feature subset. A selector can itself cross-validate candidate subsets; nesting it in a pipeline used for final evaluation prevents the outer held-out fold from influencing selection.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from sklearn.feature_selection import SequentialFeatureSelector
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

base_model = Pipeline([
    ("scale", StandardScaler()),
    ("classifier", LogisticRegression(max_iter=2000)),
])
selector = SequentialFeatureSelector(
    estimator=base_model,
    n_features_to_select=10,
    direction="backward",
    scoring="roc_auc",
    cv=StratifiedKFold(n_splits=5, shuffle=True, random_state=42),
    n_jobs=-1,
)
pipeline = Pipeline([
    ("selection", selector),
    ("model", base_model),
])
pipeline.fit(X_train, y_train)
test_score = pipeline.score(X_test, y_test)

If you tune model hyperparameters and feature selection against the same validation results, the combined search can overfit those results. Use nested cross-validation or keep a final untouched test set for the final assessment.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Handle assumptions and difficult data carefully

Correlated predictors and unstable subsets

With multicollinearity, two predictors can each look weak in a p-value model even when their group is useful. A selection procedure may arbitrarily keep one, and a different sample, split, threshold, estimator, or cross-validation partition may keep another. Compare performance and selection frequencies across resamples; consider keeping a correlated group, using domain knowledge, or applying regularization rather than treating one chosen member as uniquely important.

Interactions, confounders, and model specification

A predictor can matter through an interaction, nonlinear transformation, threshold, or subgroup effect while appearing weak as a simple main effect. Specify scientifically plausible terms before elimination. For explanatory or causal work, retain required confounders or design variables even if their p-values are large; the code’s keep argument can protect such variables, but their role should be documented. Automated selection alone does not establish a causal effect.

OLS assumptions and data preparation

P-value elimination is most defensible when the statistical model is well specified and its inferential assumptions fit the data. Check independence or model dependence, functional form, residual behavior appropriate to the intended inference, multicollinearity, missingness, sample size, categorical-variable coding, and target leakage. OLS p-values can be unstable or unavailable when predictors approach or exceed the sample size. Repeated selection in such settings can create severe false-discovery risk.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

Missing and categorical data

Imputation and encoding should be learned on training folds, typically within a preprocessing pipeline. One-hot encoding can turn a single categorical variable into several columns; independently removing dummy columns can make the selected representation hard to interpret. Consider selecting or retaining the original variable as a group where possible, and keep an explicit mapping between transformed features and their source columns.

from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

numeric_pipeline = Pipeline([
    ("imputer", SimpleImputer(strategy="median")),
    ("scale", StandardScaler()),
])
categorical_pipeline = Pipeline([
    ("imputer", SimpleImputer(strategy="most_frequent")),
    ("encode", OneHotEncoder(handle_unknown="ignore")),
])
preprocessor = ColumnTransformer([
    ("numeric", numeric_pipeline, numeric_columns),
    ("categorical", categorical_pipeline, categorical_columns),
])

Apply selection after preprocessing if the estimator operates on the transformed matrix. The transformed columns may be reordered or expanded, so use available output-feature-name support and maintain the mapping back to original variables.

Time-dependent and grouped observations

Shuffled K-fold splits can leak structure when observations are ordered in time or share a subject, customer, patient, or device. Use TimeSeriesSplit for ordered temporal data or a group-aware splitter such as GroupKFold when observations share a group. The selector and final evaluation should respect the same dependency structure.

Validate whether selection helped

Compare the reduced model against a full-feature baseline using the same leakage-safe evaluation design and the metric that matters for the task. Record enough to distinguish a useful reduction from a fragile one:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Validation and final holdout performance, with the split strategy stated.
  • Number of features retained and selection frequencies across resamples.
  • Training, inference, storage, or data-collection savings relevant to deployment.
  • Whether the reduced set is easier to explain or maintain.
  • Uncertainty in performance and subset choice, especially when correlated predictors are present.

If training performance rises but test performance falls, check first that selection did not see the test set; then consider validation overfitting, a small sample, or an overly broad search. Move selection inside the pipeline, use nested validation or an untouched test set, and simplify the search. If runs select different variables, report that instability rather than presenting one subset as definitive.

When backward elimination is the wrong tool

For very wide datasets, wrapper searches may be impractical and OLS p-values may not be reliable. Consider univariate filters or mutual information for inexpensive screening, Lasso or Elastic Net for regularized modeling, estimator-embedded methods such as SelectFromModel, domain-based feature groups, or dimensionality reduction when a transformed representation is acceptable. These approaches answer different questions and should still be evaluated without leakage.

For RFE errors about feature importance, check whether the estimator exposes coef_ or feature_importances_, or configure the documented importance_getter for the estimator or pipeline. With a pipeline, the exact attribute path depends on its step names and structure.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.