Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A decision tree is a supervised-learning model that repeatedly divides a dataset into smaller regions, ending in leaves that predict a class or a numeric value. CART (Classification and Regression Trees) chooses each division by testing feature-and-threshold candidates, scoring the resulting child nodes with a weighted impurity or loss, and selecting the lowest-scoring split. Standard CART produces a binary tree and supports both classification and regression.

What is a decision tree?

A decision tree represents a sequence of if–then tests. Each internal node asks a question about one feature, each branch records the answer, and each leaf produces the prediction. For classification, a leaf usually predicts the most common class (or class probabilities). For regression, it predicts a numeric value, commonly the average target among the training samples that reach that leaf.

The model is non-parametric: it does not assume that the data follow a particular linear, normal, or other fixed distribution. Training recursively partitions feature space so that observations grouped in the same leaf have similar targets.

How CART chooses a split

1. Generate candidate feature tests

For a numeric feature, CART considers thresholds that divide observations into two groups, such as income ≤ 52,000 and income > 52,000. It evaluates candidates for each feature at the current node. The exact handling of categorical or missing values depends on the library and version; verify those behaviors in the documentation for the estimator you are using.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Score the two child nodes

For a candidate split, the weighted impurity (or loss) is:

(nL/n) × I(L) + (nR/n) × I(R)

Here, n is the number of samples in the parent, nL and nR are the child sizes, and I is the criterion chosen for the estimator. Weighting prevents a tiny child from dominating the score.

3. Select the best improvement

CART selects the candidate with the smallest weighted score, equivalently the largest reduction from the parent node’s impurity or loss. It then repeats the search independently in both children. Splitting stops when a configured limit is reached, a node is pure enough, or no permitted split provides sufficient improvement.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

4. Predict at a leaf

At prediction time, a row follows the stored tests from the root to one leaf. Classification returns a class or probabilities; regression returns the leaf’s fitted numeric value according to the regression criterion.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Gini impurity, entropy, and log loss

For classification, the criterion changes how “mixed” a node is. No criterion is universally best; validation data or cross-validation should decide.

Criterion What it measures Practical interpretation
Gini impurity 1 − ∑ pk2, where pk is the proportion of class k Expected misclassification tendency when a label is drawn according to the node’s class distribution. Often a fast default.
Shannon entropy / information gain −∑ pk log(pk) Measures uncertainty reduction. It can rank splits slightly differently from Gini.
Log loss Penalizes poorly calibrated class probabilities, with especially large penalties for confident wrong predictions Useful when probability quality matters and the implementation supports it.

Scikit-learn’s classification-tree documentation lists Gini impurity, Shannon entropy/information gain, and log loss as supported criteria. Supported names and options can change by library version, so check the reference for the version installed in your project.

For regression, tree estimators use regression losses such as squared error; the exact set of available losses is version-dependent. Compare criteria with the metric that matters for the application rather than assuming a classification criterion applies to a numeric target.

CART versus C4.5

CART and C4.5 are greedy, recursive tree-building families, but they make different design choices. The scikit-learn guide summarizes the distinction this way: “CART (Classification and Regression Trees) is very similar to C4.5, but it differs in that it supports numerical target variables (regression) and does not compute rule sets.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Comparison axis CART C4.5
Target types Classification and regression Primarily classification in the original algorithm
Branching Binary splits: every internal node has at most two children Classically allows multiway tests, especially for categorical attributes
Split evaluation Weighted impurity or regression loss; common classification choices include Gini and entropy Information-gain ratio and related entropy-based tests
Output form A tree used directly for predictions Can be converted into a rule set; standard CART does not compute rule sets
Categorical features Handling is implementation- and version-specific; some libraries require encoding first Designed with categorical attributes in mind
Pruning and regularization Depth, minimum sample counts, leaf limits, impurity thresholds, and minimal cost-complexity pruning are common controls Uses its own pruning strategy and implementation details
Reproducibility Library behavior matters; feature permutation and tied split improvements can make results vary unless a seed is fixed Depends on the particular implementation and tie-breaking rules

The categorical-feature row should not be read as a promise about every modern CART library. For example, the scikit-learn 1.2 documentation stated that its tree implementation did not support categorical variables directly. Check the exact version and preprocessing requirements before choosing an estimator.

Why fully grown trees overfit

A tree can keep splitting until leaves contain very few observations. Such a tree may fit noise, outliers, or accidental interactions in the training set, producing low training error but poor performance on unseen data. A small tree can underfit by missing real structure, so complexity must be selected rather than maximized or minimized by rule of thumb.

Controls that limit tree complexity

Control Effect Typical use
max_depth Caps the number of levels from root to leaf Directly limits interaction depth and makes the model easier to inspect
min_samples_split Requires a node to contain at least a specified number of samples before it can split Blocks late splits supported by very little data
min_samples_leaf Requires every resulting leaf to contain a minimum number of samples Stabilizes predictions and prevents tiny terminal groups
max_leaf_nodes Limits the total number of terminal leaves Controls overall model size when depth alone is not sufficient
min_impurity_decrease Allows a split only when its weighted improvement reaches a threshold Rejects changes that have negligible objective value
Minimal cost-complexity pruning Removes subtrees whose added fit is not worth their structural complexity Grow a candidate tree, then select a pruning strength using validation

These settings interact. Increasing a minimum leaf size can make a depth limit irrelevant, while a very small impurity threshold can still permit a large tree. Select them with a validation set or cross-validation, using the metric appropriate to the task; do not treat one parameter value as universally optimal.

A defensible tuning procedure

  1. Split data into training and evaluation portions, or define cross-validation folds before fitting.
  2. Fit a deliberately constrained baseline and record the chosen task metric on held-out data.
  3. Search a grid or randomized set of depth, sample-count, leaf-count, and impurity controls without using the final test set for selection.
  4. If using minimal cost-complexity pruning, obtain candidate pruning strengths from the training procedure and evaluate them through cross-validation.
  5. Refit the selected configuration on the allowed training data, then report the untouched test result once.
  6. Inspect leaf sizes, class balance, and error by important subgroups; a single aggregate score can hide unstable or unfair leaves.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Reproducibility and implementation details

In the current scikit-learn classifier reference, features are randomly permuted at each split. When several candidates produce the same improvement, the implementation may choose among them randomly. Set random_state when deterministic fitting, comparable experiments, or audit trails are required. A fixed seed does not remove data-sampling variation from cross-validation; it only controls the estimator’s documented randomness.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Record the library name and version, criterion, preprocessing, missing-value policy, random seed, and every complexity setting with the trained model. Version changes can alter supported criteria and the treatment of categorical or missing values.

When CART is a good choice

  • Use it for a transparent baseline when stakeholders need to follow a sequence of feature tests.
  • Use regression trees when the target is numeric and the relationship is nonlinear or contains interactions that a linear model would require you to specify.
  • Use a shallow or pruned tree when interpretability and stable rules matter more than squeezing out the last fraction of validation performance.
  • Consider another model or an ensemble when a single tree is too unstable, too large, or consistently less accurate under cross-validation.

For historical context, the foundational reference is Classification and Regression Trees by Breiman, Friedman, Olshen, and Stone, published as a book in 1984. The edition and bibliographic details should be checked when citing it formally.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.