What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use rolling-origin (walk-forward) validation for forecasting: at each forecast origin, train only on data available up to that time, predict the operational horizon, then move the origin forward. Keep time order intact, refit every transformation inside each fold, and choose windows, gaps, metrics, and retraining rules that match deployment.

Why ordinary cross-validation fails for forecasts

Randomized K-fold or shuffled train-test splits can place observations from the future in the training set while evaluating an earlier date. With autocorrelated data, that breaks the information boundary of a real forecast and can make generalization look better than it is. A time-series evaluation must answer a past-to-future question: what would the model have predicted using only information that existed then?

Sort records by prediction timestamp before splitting. Check duplicate timestamps, missing intervals, revised data, and any entity or group structure. A feature is valid only if its value would have been available at the forecast origin, not merely if it is associated with that date after later revisions.

Use rolling-origin validation as the default

Rolling-origin validation (also called walk-forward or rolling-origin evaluation) repeats the deployment boundary across historical dates. For each origin, fit on the available history, forecast the next point or the required multi-step block, record errors, and advance the origin.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Choose the first origin after enough observations exist to fit the model and construct every feature.
  2. Fit the complete modeling pipeline on observations at or before that origin.
  3. Predict the same horizon used in production, without using any observations from the test block.
  4. Store point-level errors and horizon-specific errors.
  5. Advance the origin according to the intended scoring cadence and repeat.
  6. Aggregate results only after all selected origins have been evaluated.

This procedure can use an expanding history, in which every eligible past observation remains available, or a fixed-width window, in which only the most recent observations are retained.

Expanding versus fixed-width windows

Window Use it when Trade-off
Expanding Production continually accumulates all eligible history. More data can reduce variance, but old regimes may dilute current patterns.
Fixed-width Production intentionally forgets older data or the process can drift. Can adapt faster, but discards potentially useful history and increases estimate variability.

The window policy is part of the model specification. Do not select one in validation and deploy another.

Match the test block to the forecast horizon

One-step-ahead accuracy does not establish multi-step accuracy. Set each test block to the operational horizon: for example, forecast the next 24 hourly observations if that is what a scheduler consumes. Score each step separately when error changes with lead time, and report an overall horizon summary as well.

Retraining policy matters too. Simulate whether production refits at every step, on a fixed schedule, or fits once and produces a complete multi-step forecast. A single fixed fit followed by a multi-step forecast is a different experiment from refitting after every new observation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Designing folds with scikit-learn TimeSeriesSplit

Scikit-learn’s stable documentation (version 1.9.1) provides TimeSeriesSplit, an expanding-window splitter with controls for n_splits, max_train_size, test_size, and gap. It is a building block rather than a complete validation policy: you still need to choose the horizon, cadence, preprocessing, and retraining behavior.

from sklearn.model_selection import TimeSeriesSplit

splitter = TimeSeriesSplit(
    n_splits=5,
    test_size=24,       # 24 samples per test block
    max_train_size=None, # expanding history
    gap=0
)

for train_idx, test_idx in splitter.split(X):
    X_train, X_test = X.iloc[train_idx], X.iloc[test_idx]
    y_train, y_test = y.iloc[train_idx], y.iloc[test_idx]
    # fit the complete pipeline on X_train, y_train only

Comparable test durations require equally spaced samples. As the scikit-learn documentation states, “To ensure comparable metrics across folds, samples must be equally spaced.” If timestamps are irregular, construct folds by calendar dates or durations rather than treating equal row counts as equal time.

Inspect the generated indices on a small example before scoring. Confirm that each training index precedes its test indices, that test blocks have the intended duration, and that the initial training history is large enough for all lags, rolling features, and model parameters.

Choose a gap from the data-generating process

A gap separates the end of training from the start of testing. It is useful when labels overlap, feature windows share information with test targets, or data arrives with a known delay. There is no universal gap length.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Derive the gap from the target definition, maximum feature look-back, label overlap, and publication or availability delay.
  • Use gap=0 only when the split boundary itself prevents leakage.
  • Document the timing assumption and verify that the gap is applied at every origin.

For example, a target defined from outcomes over the next seven days may require a separation that prevents training labels from covering the test period. The correct value depends on the exact label and feature construction, not on a generic rule of thumb.

Prevent leakage inside every fold

The split is not enough if preprocessing can see future observations. Put learned transformations, imputation, scaling, feature selection, and target-derived encodings inside a pipeline that is refit on each training subset.

  • Compute normalization statistics from the fold’s training history only.
  • Build lagged and rolling features using an explicit prediction timestamp; shift aggregations so the target period is excluded.
  • Fit feature selection and dimensionality reduction separately for each fold.
  • Audit data revisions: a value later corrected or published after the origin was not available for that historical forecast.
  • Keep entity or group boundaries intact when observations from one entity could reveal another entity’s future.

These checks implement the same past-only information rule as the outer split.

Aggregate errors without hiding the question

State exactly how scores are combined. Pooling every point-level error weights folds with more observations more heavily; averaging fold-level metrics gives each origin equal weight. Horizon-specific scores show where forecasts degrade. These summaries answer different questions and should not be presented as interchangeable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use metrics that reflect the decision cost and remain comparable across folds. For MASE, scale absolute errors by a naive in-sample error, and recompute that scale from the training history at each origin so future values cannot enter the denominator. Report the scaling convention and seasonal-naive definition when applicable.

Evaluate a naive baseline at exactly the same origins and horizons. Beating an in-sample residual score, or beating a baseline evaluated on different dates, is not evidence of forecasting improvement.

Residuals are not forecast validation

Residuals from a model fitted on the full data set are in-sample diagnostics, not forecasts made without access to test observations. In the Google 2015 example in Forecasting: Principles and Practice, cross-validation produced RMSE 11.27, MAE 7.26, MAPE 1.19, and MASE 1.02, while training residuals were lower at RMSE 11.15, MAE 7.16, MAPE 1.18, and MASE 1.00. Those figures describe that example, not a general benchmark, but they illustrate why residuals can be optimistic.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Place origins deliberately

Use enough origins to represent meaningful historical regimes and to support stable summaries, while retaining enough initial history for fitting. Adjacent origins often produce overlapping test blocks; treat their errors as related measurements, not independent replications. Include periods affected by known seasonality, promotions, outages, or structural changes when those conditions matter in deployment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

After model and feature decisions are complete, retain a later chronological holdout when you need an untouched final check. Repeatedly tuning against every validation fold can make the selection score optimistic; the final holdout should be used once for confirmation.

When alternatives are appropriate

Irregularly sampled series

Row-based splitters do not guarantee equal calendar durations when observations are missing or unevenly spaced. Build date-based folds and define horizons in time units so each test period represents the same operational question.

Non-stationary processes

An empirical study of 62 real-world and three synthetic time series (2019) found that estimation behavior varied by scenario. In its real-world cases with non-stationary variation, methods preserving temporal order produced the most accurate estimates; cross-validation approaches could apply to stationary series. This is evidence from a defined study set, not a universal guarantee.

Bayesian time-series models

Leave-one-out can be optimistic for future prediction because observations after a held-out target may inform its prediction. Leave-future-out (LFO) instead holds out observations that occur after the training history. Exact LFO may require repeated expensive refits; PSIS-LFO approximations can reduce that cost and provide diagnostics indicating when refitting is necessary.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical pre-score checklist

  • Records are sorted by prediction timestamp, with duplicates and missing intervals investigated.
  • The target and every feature have an explicit availability time.
  • Window type, origin cadence, test duration, gap, and retraining schedule match production.
  • All learned preprocessing is inside the fold-local fitting pipeline.
  • Indices have been inspected on a small example.
  • Baselines use the same origins, horizons, and metric definitions.
  • The report identifies pooling, fold averaging, and horizon aggregation choices.
  • A later untouched holdout is reserved if model selection requires repeated tuning.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.