Recommended Free Tools
Random oversampling duplicates randomly selected minority-class examples in the training data; random undersampling removes randomly selected majority-class examples. Neither method is a guaranteed improvement: compare both with an unsampled baseline, and evaluate on validation or test data that retain the class proportions expected in use.
What random oversampling and undersampling do
Random oversampling
A random oversampler draws minority-class examples with replacement. The added rows are copies of existing observations, not newly generated examples. The imbalanced-learn documentation describes using the augmented data set to train a classifier. In one documented three-class example, a 5,000-row data set with class weights of 0.01, 0.05 and 0.94 is resampled to 4,674 examples per class (imbalanced-learn guide, 2026).
Random undersampling
A random undersampler selects and removes majority-class examples, reducing that class’s representation. This can make training faster and reduce the class-count gap, but it also discards observations that may contain useful information. A project paper on undersampling defines it as reducing the number of majority-class samples; a 2022 PLOS ONE study describes random undersampling as randomly removing majority-class examples.
How they differ from SMOTE and ADASYN
SMOTE creates synthetic minority examples by interpolating between minority-class neighbors; ADASYN similarly synthesizes examples but concentrates more on harder cases. Those are different operations from copying existing observations. For mixed continuous and categorical features, imbalanced-learn provides SMOTENC; basic SMOTE is not designed for that feature mix.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
Choosing a method
There is no universally best sampler. Choose candidates based on what information you can afford to lose, how minority examples should be represented, and which errors matter in the application.
| Approach | What happens to the data | Main consideration |
|---|---|---|
| No sampling | The classifier trains on the observed class distribution. | Keep this as the baseline; resampling is not automatically necessary. |
| Random oversampling | Minority examples are duplicated with replacement; majority examples remain. | Retains majority-class information, but repeated minority observations can increase overfitting risk. |
| Random undersampling | Randomly selected majority examples are removed. | Reduces the majority class, but discards information and can increase result variability. |
| SMOTE or ADASYN | New minority examples are synthesized by interpolation; ADASYN focuses synthesis near harder examples. | These are not simple duplication. Match the method to the feature types and the problem. |
| Hybrid, such as SMOTETomek | Combines oversampling with a cleaning or undersampling operation. | It adds another choice to evaluate rather than guaranteeing a better result. |
A 2022 PLOS ONE study tested seven sampling methods with eight classifiers on 31 real-world imbalanced data sets, using repeated 5×2 cross-validation. Across its aggregate comparison, random oversampling performed best among the sampling methods for improving AUPRC and AUROC, while undersampling reduced performance in more cases on average than oversampling and hybrid methods. That aggregate result is not a recommendation to oversample every data set: the best result required no sampling for 29 of 31 data sets by AUPRC and 30 of 31 by AUROC.
The same study found statistically significant sampling differences in 211 of 1,736 AUPRC combinations (12.2%) and 173 of 1,736 AUROC combinations (10.0%). The authors concluded that sampling could be ineffective or harmful. These results support treating sampling as an experiment to validate for your model and metric, not as a default step.
How to evaluate imbalanced classification
Split the data before resampling. Sampling the full data set before splitting can put duplicated or closely related observations on both sides of the split, making evaluation unrepresentative. The sampler belongs inside the training process so it is fit only on each training fold. Keep validation and test data untouched, and report their original class prevalence.
Rank #3
- This guide is a perfect overview for the topics covered in introductory statistics courses.
- Set aside validation or test data before applying any sampler.
- Compare an unsampled model with random oversampling and random undersampling. Add SMOTE, ADASYN or a hybrid only when there is a reason to test it.
- For cross-validation, place sampling in a pipeline so each fold fits the sampler on its training portion only. The imbalanced-learn pipeline abstraction is compatible with scikit-learn.
- Evaluate on untouched folds or holdout data that preserve the expected deployment distribution.
- Report AUPRC and AUROC, along with class-specific precision, recall or a cost-based measure that matches the consequences of errors.
AUPRC and AUROC can favor different choices, as they did in the cited multi-data-set study. Select the metric based on the decision the model supports and the relative cost of false positives and false negatives; do not declare a sampler the winner based on only one convenient score. Accuracy alone can also conceal poor performance on the class that matters most.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Using RandomOverSampler in Python
Install the open-source imbalanced-learn package alongside scikit-learn, then put RandomOverSampler and the classifier in an imbalanced-learn pipeline. This example splits first, fits the pipeline on training rows only, and scores on the original validation distribution.
Rank #4
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import average_precision_score, roc_auc_score
from imblearn.over_sampling import RandomOverSampler
from imblearn.pipeline import Pipeline
X_train, X_valid, y_train, y_valid = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42
)
model = Pipeline([
("sampler", RandomOverSampler(random_state=42)),
("classifier", LogisticRegression(max_iter=1000)),
])
model.fit(X_train, y_train)
scores = model.predict_proba(X_valid)[:, 1]
print("AUPRC:", average_precision_score(y_valid, scores))
print("AUROC:", roc_auc_score(y_valid, scores))
Here, X and y are the feature matrix and binary target, and the positive class is assumed to be encoded as 1. For a multiclass task, choose metrics and prediction handling suited to that task. The default sampler balances the classes; if a different target ratio is appropriate, set it explicitly with the sampler’s sampling_strategy parameter. Record that ratio and the random seed so the experiment is reproducible.
To make a fair comparison, use the same train/validation split and classifier settings for the unsampled baseline and each sampler. Choose the final model using the validation results and the application’s decision costs, then report its performance on a separate, untouched test set if one is available.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Quick Recap
What to report
- The original class prevalence in the evaluation data.
- Whether training used no sampling, oversampling, undersampling or another method, including the target ratio when applicable.
- The validation design, including whether sampling was confined to each training fold.
- AUPRC and AUROC, plus class-specific or cost-based measures relevant to the decision.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

