Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AutoGluon can train a useful tabular classification or regression model with a short Python workflow, choosing and combining candidate models for you. The hard parts it cannot decide for you are what the target should be, whether the data represents future cases, which metric matters, and whether the result is safe to use.

What is AutoML?

Automated machine learning (AutoML) automates parts of the supervised-learning workflow. Depending on the tool and configuration, it can prepare common feature types, handle missing values, generate features, select algorithms, tune model settings, validate candidates, build ensembles, and save a predictor for later use.

A typical workflow is: define the target, prepare training data, fit candidate models, compare validation results, select a model or ensemble, evaluate it on separate data, and use it to predict new cases. AutoML reduces repeated implementation work; it does not decide whether your target is meaningful or your data is trustworthy. AutoGluon’s tabular toolkit describes automated data cleaning, feature engineering, hyperparameter optimization, and model selection for table-based prediction problems (official tabular overview).

  • AutoML library: A Python package you run locally or in a notebook, such as AutoGluon.
  • Managed AutoML: A hosted cloud service that may provision compute, notebooks, storage, and deployment endpoints.
  • No-code AutoML: A GUI-oriented interface that reduces the need to write code.
  • Manual machine learning: You explicitly construct preprocessing, model selection, and evaluation steps, often with scikit-learn.

What is AutoGluon?

AutoGluon is an open-source Python machine-learning toolkit developed by AWS AI. It is a framework, not a single algorithm: it can train multiple model types and combine them, with separate APIs for tabular prediction, multimodal data such as text and images, and time-series forecasting. Its project repository provides current licensing and project details (AutoGluon on GitHub).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Tabular prediction is the most approachable place to start. The usual shape is one row per observation, one target column to predict, and other columns containing information available when the prediction is made. A target with categories such as yes/no is a classification problem; a numeric target such as price or demand is a regression problem. AutoGluon can read common tabular sources including CSV files and pandas DataFrames (tabular tutorials).

What you need before training

  • Basic Python familiarity and a CSV or DataFrame.
  • A clearly identified target column and a decision about how model quality should be measured.
  • Training data that resembles the cases the model will encounter later.
  • Enough RAM and disk space for the models you choose; ensembles can use considerably more resources than one simple model.

Before fitting anything, define the moment at which a prediction would be made. Remove columns that reveal what happened afterward, and avoid features unavailable at that moment. For data with repeated people, customers, devices, or time periods, a random row split can put closely related or future records on both sides of evaluation. Use group-aware or chronological splitting when that better reflects real use. AutoML cannot correct a target that is defined incorrectly, biased sampling, or mislabeled data.

Install AutoGluon

The current stable installation guide lists support for Python 3.10–3.13 and Linux, macOS, and Windows; confirm requirements for the release you install because compatibility changes (official installation guide). A virtual environment keeps the package and its dependencies separate from other Python projects.

  1. Create and activate an environment:

    python -m venv .venv
    # macOS/Linux
    source .venv/bin/activate
    
    # Windows PowerShell
    .venvScriptsActivate.ps1
  2. Update packaging tools:

    python -m pip install --upgrade pip setuptools wheel
  3. Install the tabular dependencies, or the broader package if you need other AutoGluon task areas:

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
    python -m pip install "autogluon.tabular[all]"
    # Or install the broader package:
    python -m pip install autogluon

    The basic autogluon.tabular package alone is a skeleton installation with fewer optional dependencies; the installation guide explains the available scopes (installation options).

  4. Check the installed version in the same environment you will use to run your script or notebook:

    python -c "import autogluon; print(autogluon.__version__)"

If installation fails, start with a fresh environment, upgrade the packaging tools, and verify that your notebook kernel uses that environment’s interpreter. Consult the release-specific installation guide rather than copying commands from old tutorials. Early examples using from autogluon import TabularPrediction are obsolete; the current tabular API uses TabularPredictor and TabularDataset. The stable documentation line surfaced for this article is 1.5.0, while a development API page identifies 1.5.1, so record the version actually installed rather than assuming every interface detail is identical (AutoGluon documentation).

Prepare and inspect a classification dataset

For this example, assume you have local files named train.csv and test.csv, with the label column named class. The test file contains labels for evaluation. In a real prediction pipeline, incoming records usually do not contain the label.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from autogluon.tabular import TabularDataset, TabularPredictor

train_data = TabularDataset("train.csv")
test_data = TabularDataset("test.csv")

print(train_data.head())
print(train_data.shape)
print(train_data.dtypes)
print(train_data["class"].value_counts())

Check whether values and types match expectations. Dates sometimes arrive as strings, numeric values can be accidentally read as text, and category spelling can be inconsistent. Inspect empty strings as well as missing values, high-cardinality identifiers, columns with almost no variation, and duplicate records. Ask whether an ID represents a genuine reusable signal or merely identifies the particular rows in this dataset.

Every column other than the named label is treated as a potential feature unless you remove it. In particular, check for post-outcome status fields, target-derived aggregates, and information calculated using records from the validation or future period. For a time-dependent problem, build each feature only from information available at its prediction timestamp.

Train your first AutoGluon model

The core operation is a call to TabularPredictor(...).fit(...). This explicit version sets a metric, output location, time budget, and introductory preset:

predictor = TabularPredictor(
    label="class",
    eval_metric="accuracy",
    path="AutogluonModels/ag_classification"
).fit(
    train_data=train_data,
    time_limit=120,
    presets="medium_quality"
)
  • label="class" names the value to predict.
  • eval_metric="accuracy" tells AutoGluon which measure to optimize and report. Choose a metric that reflects the actual decision, not just a familiar default.
  • path specifies where the predictor and its model files are saved.
  • time_limit=120 gives training an approximate 120-second budget; actual completion and the models fitted depend on hardware, data, version, and configuration.
  • presets="medium_quality" is a practical first run, not a claim that this is the best setting for a final benchmark.

The fit() API also supports validation data, resource limits, and other controls. The official guidance recommends starting with presets rather than immediately adjusting many low-level hyperparameters (fit API reference).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Understand training output and compare models

During fitting, logs can show the detected problem type, feature information, candidate models, validation scores, fit and prediction times, and a final weighted ensemble. The saved model directory contains the predictor artifacts. Which models finish within a time limit, which one leads, and how much space the artifacts use depend on the dataset, version, hardware, preset, and available time.

To compare trained models on the labeled test data, use the leaderboard:

leaderboard = predictor.leaderboard(test_data, silent=True)
print(leaderboard)

Common leaderboard fields include model name, score, evaluation metric, prediction time, fit time, stack level, and fit order. The exact columns can vary by version. The leaderboard helps compare candidates, but a test set loses its role as a final untouched check if you repeatedly use it to choose features, presets, thresholds, or models. Make those decisions using training/validation data and reserve a final test set for a limited final evaluation. AutoGluon’s predictor API and deployment guide describe model comparison and artifact workflows (predictor API; deployment guide).

Choose a metric and evaluate honestly

Accuracy is the fraction of examples classified correctly. It can look strong when most examples belong to one class, even if the model misses the minority class. For a binary problem where the positive class is rare, consider recall when missed positives are costly, precision when false alarms are costly, F1 when balancing the two is useful, or ROC AUC when comparing ranking across thresholds. Balanced accuracy gives class-wise performance more equal weight; log loss evaluates probability quality. For regression, MAE expresses average absolute error in target units, while RMSE penalizes larger errors more strongly.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Configure the selected metric when fitting; this example changes the optimization objective:

predictor = TabularPredictor(
    label="class",
    eval_metric="roc_auc"
).fit(
    train_data=train_data,
    time_limit=120,
    presets="medium_quality"
)

Use an appropriate representative validation method and keep a separate, genuinely unseen test set if possible. A validation score is evidence from the validation procedure, not a production guarantee. Training performance, validation estimates, final test performance, and live production monitoring answer different questions. A tuning dataset influences model selection and ensemble weights; AutoGluon warns that it should not then be treated as fully unseen evaluation data (fit API guidance).

score = predictor.evaluate(test_data)
print(score)

For imbalanced data, also inspect class-specific results and a confusion matrix, and choose any decision threshold according to the relative cost of false positives and false negatives. A score alone does not establish that a model is suitable for a medical, financial, safety, or other consequential decision.

Generate class and probability predictions

When the test data includes its label only for evaluation, remove that column before requesting predictions:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
X_test = test_data.drop(columns=["class"])
predictions = predictor.predict(X_test)
probabilities = predictor.predict_proba(X_test)

print(predictions.head())
print(probabilities.head())

predict() returns a predicted class for classification (or a numeric estimate for regression). predict_proba() returns estimated class probabilities. Probabilities are not guaranteed to be calibrated for every use; test calibration when decisions depend on their numerical meaning. New inference data should have compatible feature columns and sensible values. Reuse the saved predictor rather than separately recreating preprocessing.

Inspect feature importance carefully

Permutation importance estimates how much a model’s score changes when a feature’s values are shuffled. Use held-out data for a more informative check than the data used to fit the model:

importance = predictor.feature_importance(data=test_data)
print(importance)

This is predictive importance, not causality. Correlated features can share or obscure importance; results can change with the dataset and model. A predictive feature may also be a proxy for a sensitive attribute, so review the result in context before removing a column or acting on it. AutoGluon documents its permutation-importance method and held-out-data guidance (feature-importance API).

Choose a preset that fits your resources

Presets change the balance among model quality, training effort, inference speed, and artifact size. The names and aliases are version-sensitive; check the documentation for the release you are using (tabular essentials and presets).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Preset Typical role Trade-off to consider
medium_quality Initial prototype Faster and lighter, with lower expected quality than more intensive settings.
good_quality Improved quality with relatively efficient inference Requires more training and storage than a quick prototype.
high_quality Higher-quality experiments with faster inference than the most accuracy-focused setting Uses more compute and disk.
best_quality or current alias best Accuracy-focused experiments Can be much slower, larger, and slower at inference.
extreme Newer tabular foundation-model workflow in current documentation Additional dependencies and a GPU may be required for best results; it is not a universal laptop setting.

For a compact deployment-oriented experiment, the documentation shows combining quality and deployment optimization presets; verify supported combinations for your version:

predictor = TabularPredictor(label="class").fit(
    train_data,
    presets=["good_quality", "optimize_for_deployment"]
)

For a machine with limited resources, use a shorter time budget, CPU-only training where supported, and a lighter preset:

predictor = TabularPredictor(label="class").fit(
    train_data,
    presets="medium_quality",
    time_limit=60,
    num_gpus=0
)

Resource controls such as num_cpus, num_gpus, and memory_limit are available in fit(), but actual resource usage depends on model choices and environment. Stacking and bagging can increase training and inference cost substantially; a high-scoring ensemble is not automatically the best deployment choice (fit API resource controls).

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Save, reload, and deploy the predictor

The predictor is saved at its configured path. Reload it later to generate predictions with the fitted preprocessing and models intact:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
predictor.save()

loaded_predictor = TabularPredictor.load(
    "AutogluonModels/ag_classification"
)
new_predictions = loaded_predictor.predict(X_test)

Treat the model directory as a versioned artifact. Record the AutoGluon and Python versions, operating system, dataset snapshot or hash, feature list, training configuration, metric, and training time. A local saved predictor supports inference but is not, by itself, a monitored production service. Production use additionally requires testing the input contract, packaging, access controls, monitoring for data or performance drift, and a plan for updating or rolling back the model. The deployment guide covers reducing artifacts and preparing predictors for deployment (deployment documentation).

If you need hosted notebooks, managed compute, AWS identity integration, or hosted endpoints, AWS documents AutoGluon-Tabular for SageMaker AI (SageMaker AutoGluon-Tabular). That is a different operational choice from running the library locally and can involve charges for compute, storage, notebooks, endpoints, and related services. A local open-source package avoids a software subscription but does not make the computer or cloud infrastructure free.

Common problems and fixes

  • Installation errors: Use a new virtual environment, update packaging tools, confirm the active interpreter, install the tabular extra if that is all you need, and follow release-specific instructions.
  • Out-of-memory or disk exhaustion: Start with a lighter preset and shorter time limit, install only needed components, and avoid resource-intensive stacking unless the gain justifies its cost.
  • Poor score: Recheck label quality, feature types, target definition, class balance, split design, and metric choice before increasing training time.
  • Suspiciously excellent score: Look for target leakage, duplicate entities across splits, future information, or aggregates computed across the full dataset.
  • Unexpected columns or categories at prediction: Check that incoming data follows the training feature schema and reuse the saved predictor. Do not silently pass a label column as an input feature.
  • Dates or categories look wrong: Inspect inferred data types, inconsistent spelling, empty strings, and values loaded as text; automatic handling does not replace domain checks.
  • Time-dependent target: Do not randomly mix past and future rows. Use chronological validation or the dedicated time-series tools when forecasting future values (AutoGluon time-series tutorials).
  • Test performance keeps changing after experimentation: Stop using the test set to choose modeling decisions and reserve a fresh final holdout if possible.

AutoGluon or manual scikit-learn?

Consideration AutoGluon Manual scikit-learn workflow
Getting a first baseline Quick: automates much of model search and preprocessing. Requires more pipeline implementation.
Control and transparency Many choices are automated; ensembles can be complex. Explicit pipelines are easier to inspect and control.
Ensembling Built into the framework. Must be assembled manually or with additional tools.
Artifact and inference footprint Can be large, especially for ensembles. Often smaller for a deliberately simple model.
Good fit Rapid experimentation and a strong tabular baseline. Teaching fundamentals, tight deployment constraints, or a closely controlled production pipeline.

A useful hybrid is to establish an AutoGluon baseline, then compare it with a simple explicit model such as logistic regression or a decision tree. The comparison can reveal whether added ensemble complexity earns its cost in quality, latency, and maintainability.

When AutoML is not the right answer

  • Causal questions: Prediction and feature importance do not establish what would happen if someone intervened on a feature.
  • Strict interpretability or tiny deployment targets: A large ensemble may be harder to explain and unsuitable for an embedded device; a simpler explicit model may be preferable.
  • Time-dependent or grouped data: Validation must respect chronology or entity boundaries, or the score can overstate future performance.
  • Unsupported constraints or custom objectives: A specialized modeling workflow may require controls that the standard API does not provide directly.
  • Weak or biased data: Automation cannot make unreliable labels, unrepresentative sampling, or unavailable prediction-time inputs trustworthy.
  • Governance-heavy deployment: A Python model artifact alone does not provide organizational approval, monitoring, audit, or managed endpoint workflows.

For an AWS-managed environment, SageMaker provides a path to hosted infrastructure and integration, while a local AutoGluon installation is simpler for learning and experimentation. Choose based on operational needs, not on an assumption that managed hosting is inherently cheaper or more accurate (SageMaker documentation).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.