Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

XGBoost is a gradient-boosting framework for Python with scikit-learn-compatible estimators, a lower-level native training API, and distributed interfaces. For most scikit-learn workflows, start with XGBClassifier for classification or XGBRegressor for regression, then use a validation set and early stopping to choose how many boosting rounds to keep.

Install XGBoost in Python

For the standard stable Python package, install with pip:

python -m pip install xgboost

The default package includes support for GPU algorithms. If you only need CPU training and prefer a smaller package, the official installation guide also provides the CPU-only variant:

python -m pip install xgboost-cpu

For Conda, the documented package is py-xgboost from conda-forge. See the official installation guide for platform and GPU details. Package versions and Python compatibility change; PyPI listed XGBoost 3.4.1, released August 15, 2026, with Python 3.12+ metadata when checked. Confirm the current package metadata and installation instructions for your environment before pinning a version.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the right Python interface

XGBoost offers three useful interface families. They share the same underlying boosting project, but differ in how much control and integration they offer.

Interface Use it when Entry point
Scikit-learn estimators You want familiar fit, predict, and pipeline workflows. XGBClassifier or XGBRegressor
Native API You want lower-level control over training data, objectives, evaluation, and boosting rounds. xgboost.DMatrix and xgboost.train
Distributed interfaces Your training needs distributed data processing or multi-worker execution. Dask or Spark integrations

The official Python introduction covers native and estimator workflows; the Python API reference documents estimator parameters and supported custom objectives and metrics.

Build a reproducible scikit-learn workflow

Keep a validation set separate from the training data. It lets you monitor the metric on data the model is not fitting directly and gives early stopping a basis for selecting the boosting iteration. The example below uses a binary classification target and assumes X and y are already prepared.

from sklearn.model_selection import train_test_split
from xgboost import XGBClassifier

X_train, X_valid, y_train, y_valid = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)

model = XGBClassifier(
    objective="binary:logistic",
    eval_metric="logloss",
    n_estimators=2000,
    learning_rate=0.05,
    tree_method="hist",
    n_jobs=-1,
    early_stopping_rounds=50,
)

model.fit(
    X_train,
    y_train,
    eval_set=[(X_valid, y_valid)],
    verbose=False,
)

print(model.best_iteration)
print(model.predict_proba(X_valid)[:5])

For a continuous target, substitute XGBRegressor, choose a regression objective and metric appropriate to the task, and omit classification-only settings such as stratification when splitting. The estimator interface accepts eval_set during fitting; early stopping monitors its evaluation metric and records best_iteration. Choose the metric to match the decision you need to make, rather than relying on a default that may not fit the problem.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inspect the recorded validation history and best iteration before using the model. The estimator’s prediction behavior and iteration-range options are documented in the Python introduction. Keep evaluation data distinct from final test data if you need an unbiased final performance estimate.

Use the native API when you need lower-level control

The native interface represents data with DMatrix and trains with xgboost.train. Its evaluation sets are called evals, and the training rounds are specified directly. For example, a binary classification setup can be written as:

import xgboost as xgb

train_matrix = xgb.DMatrix(X_train, label=y_train)
valid_matrix = xgb.DMatrix(X_valid, label=y_valid)

params = {
    "objective": "binary:logistic",
    "eval_metric": "logloss",
    "tree_method": "hist",
}

booster = xgb.train(
    params,
    train_matrix,
    num_boost_round=2000,
    evals=[(valid_matrix, "validation")],
    early_stopping_rounds=50,
    verbose_eval=False,
)

print(booster.best_iteration)
booster.save_model("model.json")

The native API is useful when you want explicit control over the training matrix and booster parameters or need to work directly with a Booster. XGBoost also documents saving models in JSON, plotting feature importance, and plotting trees; see its Python introduction for the current API details.

Tune the parameters that shape the model

There is no universally best XGBoost parameter set: dataset size, missingness and sparsity, class balance, chosen metric, and compute budget all affect the trade-offs. Tune a small number of meaningful axes against validation performance rather than treating a copied parameter grid as a guarantee.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Tree construction: tree_method selects the construction approach. Histogram training with hist is a common starting point; compare alternatives only when they suit the data and compute environment.
  • Tree complexity: max_depth limits depth, while min_child_weight constrains how readily the model creates child nodes. gamma adds a minimum loss-reduction requirement for a split. Together these influence overfitting and tree size.
  • Learning rate and rounds: learning_rate controls the contribution of each boosting step; n_estimators sets the maximum number of rounds in the estimator API. A lower learning rate often calls for more rounds, which is why validation and early stopping are useful together.
  • Sampling: subsample samples rows, and colsample_bytree samples features for each tree. These affect fit and variation; assess them using the target metric.
  • Regularization: XGBoost exposes regularization controls to discourage overly complex fits. Consider them alongside tree constraints rather than adjusting only depth.
  • Compute: n_jobs controls estimator thread use. More threads can increase CPU resource use, so set it in light of the rest of your workload.

These and other estimator options are described in the Python API reference. Change parameters systematically and compare validation results; a higher training score alone is not evidence of a better model.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Run training on a GPU

For a supported NVIDIA/CUDA environment, GPU training is enabled explicitly with device="cuda". A typical histogram-based estimator configuration is:

from xgboost import XGBRegressor

model = XGBRegressor(
    tree_method="hist",
    device="cuda",
)

The same settings can be supplied to the native API through its parameter dictionary. The GPU documentation describes supported workflows; the installation guide explains wheel and platform considerations, including multi-GPU constraints. GPU execution requires a compatible environment, and it is not a guarantee of faster training for every dataset or workload. Distributed GPU workflows are also documented for Dask and Spark integrations.

Save a model for reuse

Once training and validation are complete, save the fitted estimator or native booster in a model format XGBoost supports. For a native booster, the example above writes JSON with save_model("model.json"). The official documentation also describes prediction with an iteration range, which matters when training used early stopping. Consult the Python introduction for model loading and prediction details, and retain the preprocessing and feature order your application needs to provide at inference time.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When to compare XGBoost with scikit-learn boosting

XGBoost and scikit-learn’s gradient-boosting estimators are not interchangeable winners across all datasets. Compare the specific needs that matter to your application: tree construction method, GPU or distributed-training requirements, missing and categorical data handling, evaluation and early-stopping workflow, model serialization, and operational complexity. Scikit-learn documents HistGradientBoostingClassifier as a faster option for intermediate and large datasets and explains the learning-rate/estimator-count trade-off in its ensemble documentation. Benchmark both on the same split, metric, and hardware if performance is the deciding factor.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.