Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PyCaret is an open-source, MIT-licensed Python library that compresses common machine-learning workflow steps into a consistent, low-code API. It can prepare data, compare models, tune candidates, evaluate predictions, and save pipelines while still exposing ordinary model and pipeline objects. That makes it useful for learning, rapid tabular prototypes, and repeatable baselines. It does not decide whether your target is valid, prevent conceptual leakage, choose the right business metric, or operate a production system for you.

What is PyCaret?

PyCaret is an orchestration layer over established Python machine-learning libraries. Depending on the task, it coordinates components from scikit-learn and specialist projects such as XGBoost, LightGBM, CatBoost, Optuna, and sktime. The algorithms remain the important computational building blocks; PyCaret gives them a shared experiment workflow. Its current positioning and MIT-licensed open-source status are described on the official PyCaret site.

This is “low-code” machine learning rather than one-click AI. A few calls can create preprocessing pipelines, run cross-validation, produce a comparison leaderboard, tune a candidate, generate predictions, and serialize the resulting pipeline. You still have to define a meaningful prediction problem, inspect the data, choose a valid split and metric, and decide whether the result is safe and useful.

What changed in PyCaret 4.0?

The current documentation describes five experiment areas built around object-oriented classes. PyCaret 4.0 is not backward-compatible with the 3.x functional API: older examples built around module-level setup() and compare_models() should not be mixed casually with 4.0 code. The official FAQ recommends pinning an existing 3.x project or migrating it deliberately.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The 4.0 installation documentation lists Python 3.11, 3.12, and 3.13 support; the FAQ cites scikit-learn 1.7 or newer. These are documentation claims that can change as dependencies evolve, so record the exact package versions used for a project. Check the changelog before treating any 4.0 release or API as permanently stable.

Install PyCaret in an isolated environment

A virtual environment prevents PyCaret’s scientific dependencies from colliding with unrelated notebook or scikit-learn installations.

  1. Create an environment:

    python -m venv .venv
  2. Activate it on macOS or Linux:

    source .venv/bin/activate

    On Windows PowerShell:

    .venvScriptsActivate.ps1
  3. Install the core package:

    python -m pip install --upgrade pip
    python -m pip install pycaret
  4. Install an extra only when you need it. The current installation page lists:

    python -m pip install "pycaret[dashboard]"
    python -m pip install "pycaret[explain]"
    python -m pip install "pycaret[forecast]"
    • dashboard adds dashboard/server components.
    • explain adds advanced explainability dependencies such as SHAP.
    • forecast adds more sktime adapters for time-series work.
  5. Capture the environment for reproducibility:

    python -m pip freeze > requirements.txt

For a published or production experiment, replace the unpinned install with the exact tested version, for example python -m pip install "pycaret==<tested-version>". The complete option list is in the installation documentation. CPU execution is the default. GPU support applies only to selected estimators with their dependencies installed; the same page describes roughly 4 GB RAM as comfortable for tutorial-scale work and 16 GB or more as preferable for serious workloads. Those figures are guidance, not hard requirements.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Your first PyCaret 4.0 experiment

This classification example follows the current object-oriented tutorial. It uses PyCaret’s sample juice dataset, so no external data download is needed.

from pycaret.classification import ClassificationExperiment
from pycaret.datasets import get_data

data = get_data("juice", verbose=False)

exp = ClassificationExperiment(
    target="Purchase",
    session_id=42,
    normalize=True,
).fit(data)

result = exp.compare_models(n_select=3)
print(result.leaderboard.head())

tuned = exp.tune_model(
    result.best,
    n_iter=20,
    optimize="AUC",
)

predictions = exp.predict_model(tuned.pipeline)
print(predictions.metrics)

The comparison produces cross-validation metrics and candidate models, tune_model searches for a better configuration against the selected metric, and predict_model returns predictions and evaluation information. Exact rankings and scores can vary with package, dependency, hardware, seed, and dataset changes; do not publish a hard-coded “winner” from this example. The source tutorial is at PyCaret’s tutorials page.

Rank #2
Sale
Deep Learning (Adaptive Computation and Machine Learning series)
  • Language Published: English
  • Binding: hardcover
  • It ensures you get the best usage for a longer period

Save the fitted pipeline

Save the complete fitted pipeline rather than only a bare estimator so preprocessing travels with the model:

from pycaret import save_model, load_model

save_model(tuned.pipeline, "juice_classifier")
loaded_model = load_model("juice_classifier")

The 4.0 changelog shows top-level pycaret.save_model and pycaret.load_model; some experiment APIs also expose persistence methods. Confirm the accepted object and import path against the version you installed, using the changelog and module documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PyCaret’s five current experiment modules

Module Experiment class Use it for
Classification ClassificationExperiment Binary or multiclass categorical targets
Regression RegressionExperiment Continuous numeric targets
Clustering ClusteringExperiment Grouping observations without a labeled target
Anomaly detection AnomalyExperiment Finding observations unusual under supplied features
Time series TimeSeriesExperiment Forecasting from time-indexed data

See the current module reference. Older 3.x documentation also listed NLP and association-rule functionality, but those lists should not be presented as the 4.0 task surface; historical context is available in the 3.x documentation.

What PyCaret automates

  • Missing-value handling and common categorical encoding.
  • Scaling or normalization when requested.
  • Cross-validation and baseline model fitting.
  • Consistent comparison of many estimators.
  • Hyperparameter tuning against a chosen metric.
  • Predictions, basic plots, and diagnostics.
  • Serialization of fitted pipelines and model objects.

Automation does not make defaults correct. A leaderboard is meaningful only when the target, split strategy, metric, preprocessing, and sampling assumptions match the real decision.

A reliable workflow beyond the demo

1. Define the prediction problem

State the target, prediction time, costly errors, and success metric before fitting anything. Decide whether the target is categorical, continuous, absent, or time-indexed. A feature created after the outcome is unavailable at prediction time even if PyCaret can technically ingest it.

2. Inspect the data

data.head()
data.dtypes
data.isna().sum()
data.describe(include="all")

Look for duplicates, impossible values, class imbalance, IDs treated as features, timestamp ordering, post-outcome columns, and train/test contamination.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

3. Choose an experiment class

from pycaret.classification import ClassificationExperiment
from pycaret.regression import RegressionExperiment
from pycaret.clustering import ClusteringExperiment
from pycaret.anomaly import AnomalyExperiment
from pycaret.time_series import TimeSeriesExperiment

Use these 4.0 classes instead of copying unlabelled 3.x snippets.

4. Establish a baseline

Start with a simple model and a business-relevant metric. A complex model is not useful if it cannot beat a reasonable baseline or if its score represents the wrong cost.

5. Compare deliberately

compare_models() coordinates repeated fits and cross-validation. Restrict candidates when latency, interpretability, licensing, or hardware matters, and use n_select=3 or another small set rather than assuming the first automatic winner is final.

6. Tune without consuming your test set

Choose the metric first, retain an untouched final holdout, and compare tuned results with the untuned baseline. Repeatedly tuning against the same validation results can overfit the selection process; serious benchmarks may need nested validation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

7. Inspect errors

Review false positives and negatives, regression residuals, probability calibration, segment-level performance, feature explanations, and stability over time. Explanations are diagnostic evidence, not proof of causality or fairness.

8. Treat deployment as a separate engineering project

A serialized artifact still needs input-schema checks, dependency locks, logging, drift monitoring, rollback, retraining rules, privacy review, security controls, and a documented model card or equivalent record.

Task-specific cautions

Classification

For imbalanced data, accuracy can hide disastrous minority-class recall. Compare precision, recall, F1, ROC AUC, PR AUC, and calibration as appropriate; set a decision threshold based on error costs. Grouped or temporal validation may be more valid than random folds.

Regression

MAE is easier to interpret, while RMSE penalizes large errors more heavily. MAPE behaves poorly around zero. Inspect outliers and skew, consider prediction intervals, and back-transform predictions correctly after log transformations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Clustering

There is no ordinary labeled target, so “best” is less objective. Scale features, justify the distance measure and cluster count, treat silhouette scores as limited evidence, and test whether clusters are stable and interpretable for the business.

Anomaly detection

An anomaly is unusual under the supplied features, not automatically fraudulent or harmful. Contamination assumptions, changing baseline behavior, false positives, human review, and missing ground-truth labels are central design issues.

Time series

Never treat forecasting as ordinary shuffled tabular validation. Preserve time order, define a forecast horizon, and use rolling or expanding-window backtesting. Account for seasonality, missing timestamps, exogenous variables, intervals, and residual diagnostics. The official time-series tutorial demonstrates TimeSeriesExperiment, horizon-based comparison, tuning, intervals, and residual analysis.

Who benefits from PyCaret?

Beginners

PyCaret exposes the shape of a complete workflow with less boilerplate and supplies task-focused tutorials and datasets. Learn in stages: pandas and inspection; train/test concepts; a classification or regression baseline; metrics and error analysis; tuning; then dependency and deployment hygiene.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Experienced data scientists

Experts can use it for rapid benchmarking, shared pipeline conventions, and candidate restriction before writing a more explicit scikit-learn implementation. Its sklearn-oriented pipelines make that transition more practical than exporting an opaque prediction service.

Experts may reject it when abstraction hides a critical transformation, a custom validation scheme does not fit the generic workflow, a search is too expensive, or strict production review requires every decision to be explicit.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and recovery

3.x code running in 4.0

Import errors or changed signatures usually indicate an API mismatch. Check the installed version, consult the matching documentation, migrate to experiment classes, or pin the existing project to a compatible 3.x release. Do not mix APIs in one environment; see the FAQ.

Dependency conflicts

Run python -m pip check and python -m pip freeze. If model imports still fail, rebuild the environment with pinned Python and PyCaret versions instead of layering packages into a general-purpose environment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leakage and imbalance

Common leakage sources include future-derived features, imputation fitted across test data, random splits on temporal records, duplicates across folds, and target-derived columns. For imbalance, choose an appropriate metric, inspect the confusion matrix, and consider thresholding or resampling.

Over-comparison and resource use

Trying many estimators reduces coding effort but can perform many fits and consume substantial CPU and memory. Limit candidates, budget compute, and preserve a final holdout. A leaderboard is not a production recommendation.

Serialization and dashboards

Never load untrusted pickle-like artifacts. Store models with access control and test them in a clean environment. The dashboard extra is optional; notebooks and scripts need only the core package.

PyCaret compared with alternatives

Choice Strength Trade-off
PyCaret Free local Python workflow, rapid baselines, inspectable pipelines Version churn, local dependency and operations burden
Plain scikit-learn Maximum control and a smaller abstraction surface More workflow code and manual model comparison
H2O Driverless AI Enterprise AutoML, feature engineering, interpretation, deployment and governance Commercial licensing and substantially higher potential cost; see product information and cloud documentation
Amazon SageMaker AI Managed AWS training, hosting, monitoring and governance Usage-based charges for compute, endpoints, storage and related services; see pricing
Google managed ML Hosted training and serving integrated with Google Cloud Cloud operations and billed training, endpoints and predictions; see pricing
DataRobot Commercial AutoML and enterprise MLOps support Generally quote-based pricing; its official pricing reference directs customers to a representative

PyCaret’s core package is free and open source; hosted compute, enterprise control planes, support, and managed alternatives can cost money. An AWS Marketplace listing for H2O AI Cloud displayed contract figures of $225,000 per GPU for 12 months and a $720,000-per-year starter configuration on August 16, 2026; those listing figures are not universal retail prices and AWS infrastructure may be additional.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is PyCaret right for you?

  • Beginner prototype or tabular baseline: usually yes.
  • Local, open-source experimentation: yes, if you pin the environment.
  • Highly customized scikit-learn architecture: plain scikit-learn may provide better control.
  • Deep-learning-first, computer-vision, or large-language-model work: usually not the natural fit.
  • Very large datasets: assess the cost of fitting every compared model.
  • Regulated production: use PyCaret only within explicit validation, security, monitoring, governance, and review controls.
  • Managed enterprise MLOps: evaluate a cloud or commercial platform instead.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.