PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchUse scikit-learn’s DummyClassifier for classification or DummyRegressor for regression. Fit the dummy estimator on the training data, then score it and your candidate model with the same metric and the same cross-validation folds (or the same holdout set). The resulting score is a reference point: because dummy estimators ignore feature values, a useful model should beat an appropriate dummy rule under the evaluation that matters to your project.
What “automatic baseline” means in scikit-learn
Scikit-learn supplies ready-made estimators that implement simple prediction rules. You still choose the task, rule, scoring measure, and evaluation design; there is no single automatic baseline that is correct for every dataset.
The scikit-learn developers describe DummyClassifier as “a simple baseline to compare against other more complex classifiers.” Their DummyRegressor documentation describes a “Regressor that makes predictions using simple rules.” Neither estimator learns a relationship between feature values and the target.
Choose the estimator for your task
| Task | Estimator | What predictions use | Available rules |
|---|---|---|---|
| Classification | DummyClassifier |
Target labels and, for some strategies, their observed class distribution; input features are ignored. | stratified, most_frequent, prior, uniform, or constant |
| Regression | DummyRegressor |
Target values only; input features are ignored. | mean, median, quantile, or constant |
Create a classification baseline
Majority-class baseline
most_frequent always predicts the label that occurs most often in the training targets. This is the direct scikit-learn equivalent of a majority-class baseline.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
from sklearn.dummy import DummyClassifier
from sklearn.metrics import accuracy_score
baseline = DummyClassifier(strategy="most_frequent")
baseline.fit(X_train, y_train)
predictions = baseline.predict(X_test)
print(accuracy_score(y_test, predictions))
Use the matching training features in fit even though the dummy rule does not use feature values; this preserves the normal estimator interface and validates the input shape.
Other classifier strategies
prior: predicts the class with the largest training prior and provides class-prior probabilities. It is useful when you want a deterministic prior-based reference.stratified: generates random labels according to the training class distribution.uniform: generates random labels with a uniform class distribution.constant: always predicts a label you supply with theconstantparameter.
Set random_state for repeatable results with stratified or uniform. The other strategies are deterministic after fitting. A randomized baseline can vary from run to run without that seed, so compare it using the same reproducibility settings as your experiment.
Rank #2
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
random_baseline = DummyClassifier(
strategy="stratified",
random_state=42,
)
random_baseline.fit(X_train, y_train)
Create a regression baseline
Mean baseline
DummyRegressor(strategy="mean") predicts the mean of the training targets for every row.
from sklearn.dummy import DummyRegressor
from sklearn.metrics import mean_absolute_error
baseline = DummyRegressor(strategy="mean")
baseline.fit(X_train, y_train)
predictions = baseline.predict(X_test)
print(mean_absolute_error(y_test, predictions))
Median, quantile, and constant rules
median: predicts the training-target median, often a useful reference when large outliers would distort the mean.quantile: predicts a specified training-target quantile; provide the quantile value required by the estimator.constant: predicts a caller-supplied constant.
quantile_baseline = DummyRegressor(
strategy="quantile",
quantile=0.9,
)
quantile_baseline.fit(X_train, y_train)
Choose the rule to answer a defined baseline question. A mean baseline and a median baseline answer different questions, so do not treat whichever scores better as universally “correct.”
Free tools Windows power users keep installed
One-click scans. No signup required.
Compare the baseline and candidate fairly
Use one scoring rule
Fit and evaluate both estimators with the same task-appropriate metric. Accuracy can be misleading for imbalanced classification; consider a metric such as balanced accuracy, precision, recall, F1, log loss, or an appropriate area-under-curve measure when that reflects the real decision. For regression, select a measure such as mean absolute error, mean squared error, or a percentage-based metric only when its assumptions fit the target.
The estimator’s default score method is not automatically the metric your project needs. State the metric explicitly in code and in your results.
Rank #4
Use the same cross-validation design
Cross-validation gives a more stable comparison than scores from unrelated splits. Pass the same data partitions, splitter, and scoring choice to both estimators.
from sklearn.model_selection import cross_val_score, StratifiedKFold
from sklearn.linear_model import LogisticRegression
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scoring = "balanced_accuracy"
baseline = DummyClassifier(strategy="most_frequent")
candidate = LogisticRegression(max_iter=1000)
baseline_scores = cross_val_score(
baseline, X, y, cv=cv, scoring=scoring
)
candidate_scores = cross_val_score(
candidate, X, y, cv=cv, scoring=scoring
)
print("baseline mean:", baseline_scores.mean())
print("candidate mean:", candidate_scores.mean())
For regression, use a regression splitter and scoring value, then call cross_val_score in the same way for the dummy and candidate estimators. Keep preprocessing inside a pipeline so each fold learns transformations from its training portion only.
Read the result as a sanity check
- If the candidate clearly beats the dummy under the metric and split design that matter, the model has evidence of useful predictive signal relative to that simple rule.
- If it does not, inspect the features, target construction, data split, preprocessing, metric, and model configuration before claiming improvement.
- A dummy score is not a production-quality forecast: it says how a deliberately simple rule performs, not how much the features explain.
- For imbalanced labels, compare against a baseline that reflects the question you care about; a high raw accuracy from always predicting the majority class may conceal poor minority-class performance.
Practical checklist
- Identify whether the target is a class label or a numeric value.
- Instantiate
DummyClassifierorDummyRegressor. - Select a strategy that represents the baseline question: majority, prior, random, constant, mean, median, quantile, or constant.
- Set
random_statefor randomized classifier strategies when reproducibility matters. - Fit on the training targets with the matching training features.
- Choose and document a metric that represents the actual goal.
- Evaluate the dummy and candidate on identical folds or an identical held-out set.
- Investigate any failure to beat the baseline before tuning or deploying the candidate.
Version and interface note
The API details summarized here follow the scikit-learn documentation identified for versions 1.9.1 (dummy-estimator API) and 1.4.2 (evaluation guide). Strategy names, parameters, and scoring availability can change between releases, so check the documentation for the version installed in your environment and keep the code and claims aligned with it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

