Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Develop a Super Learner by generating out-of-fold predictions from a prespecified set of candidate models, then choosing how to combine them under a loss function. In Python, scikit-learn’s StackingRegressor and StackingClassifier handle the out-of-fold stacking workflow, but their ordinary defaults are not the exact constrained Super Learner: its meta-weights are nonnegative, sum to one, and have no intercept. Whichever approach you use, evaluate the finished procedure on data that did not train its component models or choose its settings.

What a Super Learner does

A Super Learner combines predictions from a prespecified library of candidate algorithms. It uses V-fold cross-validation to produce predictions for rows each candidate did not train on, then selects combination weights by minimizing a chosen loss. Van der Laan, Polley, and Hubbard introduced the method in their 2007 paper, “Super Learner.”

The cross-validation step matters because the combiner should learn from predictions that resemble the errors it will encounter on new data. If it instead learns from each base model’s predictions on rows used to fit that model, those predictions can be unrealistically good, and the combiner may overfit.

How it differs from other ensembles

  • Bagging combines models trained on resampled versions of data, often to reduce variance.
  • Boosting builds a sequence of models, with later models responding to earlier errors.
  • Voting or fixed-weight averaging combines predictions using weights chosen in advance or by a fixed rule.
  • Super Learner uses cross-validated predictions to learn a loss-minimizing combination from a chosen candidate library.

It does not guarantee a better result than the strongest candidate on every dataset. The outcome depends on the data, prediction target, loss, validation design, and whether the library contains models whose errors complement one another.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a useful candidate library

The library is a substantive modeling choice, not a list that can be selected once for every problem. Include candidates that are plausible for the target and data and that bring meaningfully different inductive biases. Depending on the task and sample size, examples may include regularized linear models, tree ensembles, and support-vector methods.

Keep preprocessing inside each model pipeline

Fit learned preprocessing—such as imputation, scaling, or feature selection—inside a scikit-learn Pipeline paired with its estimator. Then the preprocessing is refit within each training fold instead of learning from validation-fold rows. Apply the same principle to any feature engineering that estimates quantities from the data.

Choose validation splits to match the data

For independent, identically distributed observations, shuffled folds may be suitable. For time-ordered data, grouped observations, or other dependent samples, use a splitter that respects that structure. Random folds can leak information across time or across related observations. Check that every training fold contains the classes and data needed to fit every candidate.

Build a standard stack with scikit-learn

Use StackingRegressor for a continuous target and StackingClassifier for classification. In the stable API documentation displayed as version 1.9.1 on September 30, 2026, both use five folds when cv=None; treat that as a documented default, not a universal recommendation. Supply a splitter appropriate to the data and the installed scikit-learn version.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Regression example

from sklearn.ensemble import StackingRegressor
from sklearn.linear_model import LinearRegression, Ridge, Lasso
from sklearn.ensemble import RandomForestRegressor

base_estimators = [
    ("ridge", Ridge(alpha=1.0)),
    ("lasso", Lasso(alpha=0.01)),
    ("forest", RandomForestRegressor(
        n_estimators=300,
        min_samples_leaf=2,
        random_state=42,
    )),
]

# Nonnegative coefficients, but no exact sum-to-one constraint.
meta = LinearRegression(fit_intercept=False, positive=True)

ensemble = StackingRegressor(
    estimators=base_estimators,
    final_estimator=meta,
    cv=5,
)
ensemble.fit(X_train, y_train)
predictions = ensemble.predict(X_test)

Here cv=5 creates the folds used to generate training features for the final estimator. The stack then fits the base estimators on all of X_train. In the version 1.9.1 API documented on September 30, 2026, the regressor’s default final estimator is RidgeCV; specifying one explicitly makes the modeling choice visible.

Classification example and prediction methods

from sklearn.ensemble import StackingClassifier, RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.svm import SVC

base_classifiers = [
    ("logistic", LogisticRegression(max_iter=2000)),
    ("forest", RandomForestClassifier(
        n_estimators=300,
        min_samples_leaf=2,
        random_state=42,
    )),
    ("svc", SVC(probability=True, random_state=42)),
]

classifier = StackingClassifier(
    estimators=base_classifiers,
    final_estimator=LogisticRegression(max_iter=2000),
    cv=5,
    stack_method="auto",
)
classifier.fit(X_train, y_train)
probabilities = classifier.predict_proba(X_test)

With stack_method="auto", StackingClassifier tries each base estimator’s predict_proba, then decision_function, then predict. These outputs have different meanings: probabilities represent estimated class likelihoods, decision scores are margins, and hard predictions discard confidence information. In binary classification, scikit-learn drops the first probability column when it would be perfectly collinear with the second. If probabilities drive decisions, assess their calibration on appropriate held-out data; a probability-producing method does not by itself establish calibration.

The classifier’s documented default final estimator in the version 1.9.1 API displayed September 30, 2026, is LogisticRegression. As with regression, set the final estimator explicitly when you need a particular model or loss behavior.

Make the blend an exact constrained Super Learner

A Super Learner-style convex blend predicts with weights that are each at least zero and whose sum is exactly one, without an intercept. This makes the result a weighted average of candidate predictions: if all candidates predict the same value for a row, the blend returns that value. The positive, no-intercept LinearRegression combiner above enforces nonnegative coefficients but does not enforce that they sum to one, so it is an approximation rather than an exact convex blend.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Exact constrained regression blend

The following example builds out-of-fold predictions, minimizes squared error subject to the convex-weight constraints, and then fits each candidate on all training rows for prediction. It uses SciPy’s constrained optimizer. Pass pipelines as candidates when preprocessing must be learned within folds.

import numpy as np
from scipy.optimize import minimize
from sklearn.base import clone
from sklearn.model_selection import KFold, cross_val_predict

def fit_convex_super_learner_regressor(estimators, X, y, cv=None):
    """Fit a squared-error convex blend and refit base estimators on all rows."""
    y_array = np.asarray(y).ravel()
    if cv is None:
        cv = KFold(n_splits=5, shuffle=True, random_state=42)

    # Each column contains predictions made for rows held out of that fold.
    oof = np.column_stack([
        cross_val_predict(clone(estimator), X, y_array, cv=cv, method="predict")
        for estimator in estimators
    ])

    n_models = oof.shape[1]
    objective = lambda weights: np.mean((y_array - oof @ weights) ** 2)
    result = minimize(
        objective,
        x0=np.full(n_models, 1.0 / n_models),
        method="SLSQP",
        bounds=[(0.0, 1.0)] * n_models,
        constraints={"type": "eq", "fun": lambda weights: weights.sum() - 1.0},
    )
    if not result.success:
        raise RuntimeError(f"Weight optimization failed: {result.message}")

    fitted = [clone(estimator).fit(X, y_array) for estimator in estimators]
    return fitted, result.x

def predict_convex_super_learner_regressor(fitted, weights, X):
    base_predictions = np.column_stack([
        estimator.predict(X) for estimator in fitted
    ])
    return base_predictions @ weights

This is an exact convex blend for the stated regression setup and squared-error objective, not a universal implementation of every Super Learner variant. The loss and output representation should match the prediction problem. For classification, define the candidate outputs and classification loss deliberately; if blending probabilities, verify calibration and retain valid probability behavior. For time series, groups, or other structured data, use an appropriate cross-validation design rather than the example’s shuffled KFold.

The scikit-learn developers’ stacking example notes that “The cleanest way to enforce the coefficient normalization with scikit-learn is by defining a custom estimator.” Its positive, no-intercept linear-regression example is useful when an approximation is sufficient, but do not describe it as sum-to-one constrained.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Evaluate the complete procedure honestly

The cross-validation inside stacking constructs training inputs for the final estimator; it is not an independent estimate of how the complete modeling process will perform on new data. After comparing candidates, choosing the library, tuning settings, and fitting the stack, use untouched test data or an appropriately nested validation design for performance estimation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Avoid prefit leakage

The stacking API’s cv="prefit" option fits the final estimator on predictions from already-fitted base estimators. If those estimators were trained on the same rows, the meta-model sees in-sample predictions, creating a very high risk of overfitting. Use it only when the base predictions are genuinely out of sample relative to the rows used to fit the final estimator.

Compare candidates and ensemble on the same basis

Evaluate every candidate and the ensemble with the same outer splits and task-appropriate metrics. Consider:

  • Held-out loss and metrics aligned with the task, rather than only the meta-model’s internal fit.
  • Variation across suitable resamples, to see whether apparent gains are stable.
  • Probability calibration when predictions are used as probabilities.
  • Interpretability and whether the blend’s weight constraints matter to the application.
  • Training and prediction cost, plus the operational complexity of keeping several models.

In scikit-learn’s worked synthetic regression example, stacking achieved a slight improvement over selecting the best model in that example, at greater computational expense. Its displayed fitted coefficients—including an approximate constrained combiner whose weights sum to about 0.9977—are specific to that generated dataset, not expected weights or a general performance guarantee.

When to choose stacking instead of one model

Approach What it does Trade-off
Select one learner Choose a candidate using valid outer validation. Usually simpler and cheaper to fit; performance depends on selecting a suitable candidate.
Ordinary scikit-learn stacking Fits a final estimator on cross-validated base predictions; optionally, passthrough=True also supplies original features. Flexible and built in, but the final estimator can have unconstrained weights and an intercept.
Constrained Super Learner-style blend Uses nonnegative weights summing to one and no intercept. Has interpretable convex contributions, but exact normalization requires a constrained/custom estimator.

A broader model library can help when its members make complementary errors, but it also increases fitting cost and the number of choices that need validation. Keep the stack only if its measured benefit on appropriate evaluation data justifies those costs for the task.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.