Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a model with development data, then estimate how the complete selection process will perform on data it has not seen. Cross-validation and parameter search help compare candidates, but the highest tuning score is not automatically a trustworthy estimate of future performance.

Model selection and model evaluation are different jobs

Model selection chooses a model family, features, preprocessing steps, and hyperparameter settings. Model evaluation estimates how the chosen workflow will perform on new data. The distinction matters because every choice informed by validation results can adapt to quirks in those results.

As the scikit-learn cross-validation guide puts it: “Learning the parameters of a prediction function and testing it on the same data is a methodological mistake.” A model can memorize its training examples and score well there while performing poorly on unseen examples.

Choose a metric and split strategy that match the task

Start with the decision you need to support

Decide what a useful prediction means before comparing models. For classification, accuracy may conceal poor performance on a rare class or errors with a high cost. For regression, a metric should reflect the size and consequences of prediction errors. The right scoring rule depends on the task and decision; scikit-learn documents separate metric sections for classification, regression, multilabel, and clustering.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make validation resemble deployment

A random split is not suitable for every dataset. If the goal is to predict future observations, randomly mixing earlier and later records can make validation unlike deployment. If multiple rows belong to the same person, device, or other group, placing related rows on both sides of a split can also make the task unrealistically easy. Choose a time-aware or group-aware strategy when future periods or independent groups are what the model must handle.

Stratified folds attempt to preserve class proportions across classification folds, which can help avoid a fold with no examples of a rare class. Stratification addresses that practical problem; it does not by itself make an evaluation statistically sound. See scikit-learn’s cross-validation guide for splitter choices and considerations.

How holdout and cross-validation compare

Method How it works Useful when Main caution
Holdout split Separate development data from an evaluation portion. You can reserve an untouched final check and want a simple workflow. The estimate can depend heavily on that one split; ensure it represents the intended prediction setting.
K-fold cross-validation Rotate which fold is held out, so each observation is validated in turn. You want multiple development scores and to use data efficiently for comparing candidates. It costs more computation than one split, and folds must respect time, groups, or other structure in the data.
Stratified folds Construct classification folds that attempt to preserve class proportions. Class imbalance might otherwise leave a fold without examples of a class. Stratification alone does not guarantee a sound estimate.
Nested cross-validation Use inner folds for selection and outer folds to evaluate that selection procedure. You need to estimate performance without a separate untouched test set and are tuning on the available data. It is computationally more expensive.

Cross-validation repeats train-and-validation splits; it does not make any splitter appropriate by default. For scikit-learn, the documented API behavior for an integer or None cross-validation setting is five folds for binary or multiclass classifiers and KFold otherwise, with shuffling disabled by default. Check the documentation for the library version you use, and specify a splitter deliberately when defaults do not match your data.

How to search hyperparameters

Hyperparameters are settings chosen outside the model’s fitted parameters, such as regularization strength or tree depth. Search methods evaluate candidate settings under a scoring rule; they do not eliminate the need for a credible evaluation.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Search method What it does Trade-off
Grid search Evaluates combinations from an explicit set of values. Clear and reproducible for a small search space, but cost rises with the number of combinations and folds. A coarse grid can miss useful regions.
Randomized search Samples settings from specified distributions or lists. Can explore broader spaces with a fixed budget, but results depend on the search space, budget, and randomness.
Successive halving Begins with many candidates and allocates more resources to the promising ones. Can avoid spending heavily on weak candidates, but the resource must be meaningful and early rankings must be reliable enough to guide elimination.

For a small, prespecified set of plausible choices, grid search is straightforward. For a broad space, randomized search offers a set budget. Successive halving is an option when its resource allocation makes sense for the estimator and task. Scikit-learn’s hyperparameter tuning guide describes these search approaches.

A practical workflow for choosing a model

  1. Define the prediction task and metric. Identify the target, the cost of different errors, and the scoring rule before searching.
  2. Reserve final evaluation data if feasible. Set it aside before development; do not use it to choose preprocessing, features, a model family, or parameters.
  3. Build a fitted pipeline. Put learned preprocessing, feature selection, and the estimator together so each training fold learns its transformations without information from its held-out fold.
  4. Compare sensible candidates. Use a simple baseline and a small set of plausible model families. Evaluate them on development data with a splitter suited to the data and deployment question.
  5. Set a search budget. Choose grid search, randomized search, or successive halving according to the size of the space and whether its computation assumptions fit.
  6. Review scores and their variation. Select by the metric that matches the task, and inspect fold-to-fold variation rather than relying only on the mean.
  7. Estimate performance independently. Use the untouched test set for a final evaluation, or use nested cross-validation when no separate final test set is available and selection must be evaluated on the same dataset.
  8. Refit for use. Once the evaluation is complete, fit the chosen workflow on all available development data. Keep the independent evaluation estimate distinct from the refitted model.

When nested cross-validation is useful

If you tune candidates and report the best score from those same folds, the result can be optimistic: among many candidates, one may score well partly because it fits noise in the validation results. Nested cross-validation separates the tasks. The inner loop selects settings; the outer loop evaluates the process that performs that selection.

Nested cross-validation is most useful when you need an estimate of the selection procedure and have not set aside a final test set. If a genuinely untouched final test set has been reserved and kept out of every development decision, nested cross-validation is not required for the final check. Scikit-learn illustrates the difference in its nested versus non-nested cross-validation example.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common mistakes that make results misleading

  • Training and scoring on the same observations: this rewards memorization rather than showing how the model handles new examples.
  • Calling the best tuning score a final test result: repeated comparisons let model selection adapt to noise in validation scores.
  • Preprocessing before splitting or cross-validation: a transformation or feature-selection step can learn from held-out examples. Fit such steps inside the pipeline and training folds.
  • Randomly splitting data with time or group structure: validation can become easier or otherwise unlike the future or independent cases of interest.
  • Optimizing accuracy by default: accuracy may not reflect rare classes or unequal error costs. Choose a metric tied to the actual outcome and decision.
  • Rechecking the final test set during development: once its results influence choices, it is no longer an independent final check.

Other model-selection criteria

Criteria such as AIC and BIC penalize model complexity while comparing fit in settings where their assumptions and implementation apply. They can support likelihood-based statistical model selection, but they are not interchangeable with predictive metrics measured on held-out data. Check whether a criterion is appropriate for the estimator and question rather than treating it as a universal ranking.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.