Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A decision tree is a supervised machine-learning model that predicts a class or numeric value by asking a sequence of feature-based questions. Its branching, if-then structure is easy to inspect, but an unrestricted tree can memorize training data and perform poorly on new cases.

What is a decision tree?

A decision tree learns rules from labeled examples. At each internal node, it tests a feature; each branch corresponds to an outcome of that test; and each leaf gives the prediction. For classification, a leaf predicts a class. For regression, it predicts a numeric value.

For example, a classifier for loan applications might first ask whether income is above a threshold, then ask about debt for applications on one side of that split. The result is a path of decisions rather than one formula applied uniformly to every case. Trees are non-parametric: they do not assume a fixed functional form for the relationship between inputs and the target.

How does a decision tree choose a split?

Training proceeds recursively. At a node, the algorithm considers candidate questions—often a feature and a threshold—and scores how well each would separate the observations into child groups. It selects the best-scoring split available at that node, then repeats the process in each child until a stopping rule applies.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In scikit-learn’s formulation, the observations at a node are represented by Qm, and a candidate split is θ = (j, tm): feature j and threshold tm. For a binary split, the left child receives observations where xj ≤ tm; the right child receives the rest. The algorithm chooses the split that minimizes the weighted impurity of the children. Scikit-learn’s decision-tree documentation describes this procedure and notes that “scikit-learn uses an optimized version of the CART algorithm.”

This is a greedy method: each node makes the best choice it can at that point. It does not search all possible complete trees to guarantee the globally best overall tree. As a result, an early split can shape the choices available deeper in the tree.

Example: reducing classification impurity

Suppose a node contains 10 examples: 5 are positive and 5 are negative. A candidate question divides them into two child nodes, each containing 5 examples. If one child has 4 positive and 1 negative example while the other has 1 positive and 4 negative examples, both children are more class-concentrated than the original node. A classification criterion measures that improvement and compares it with other candidate splits. The tree then applies the process again within each child.

Gini impurity, entropy, and regression loss

For classification, common split criteria include Gini impurity and entropy, often expressed through information gain. Gini impurity is low when a node is dominated by one class and higher when classes are mixed. Entropy similarly measures uncertainty in the class distribution; information gain describes the reduction in entropy resulting from a split.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Regression trees use a regression criterion, such as squared error, rather than class impurity. The appropriate criterion depends on the task and implementation; there is no single universally correct split measure. Scikit-learn documents the available mathematical formulation and criteria.

Classification trees and regression trees

Tree type Target What a leaf predicts Typical split scoring
Classification tree A discrete class or category A predicted class, typically based on the examples reaching the leaf Gini impurity or entropy/information gain
Regression tree A numeric value A numeric prediction A regression loss such as squared error

The branching mechanism is similar in both cases, but the target and the criterion used to judge a split differ. Evaluate each with a metric suited to its task—for instance, a classification metric for class labels or a regression metric for numeric predictions.

Decision-tree families: ID3, C4.5, C5.0, and CART

These names refer to related but distinct algorithm families, not interchangeable settings. Their differences include supported targets, split types, and how they handle feature values.

  • ID3 uses information gain and is associated with categorical features and multiway splits.
  • C4.5 extends the family to continuous-feature thresholds and can convert trees into rules.
  • C5.0 is a later proprietary Quinlan family.
  • CART produces binary splits and supports classification and regression. Scikit-learn uses an optimized CART implementation.

When comparing implementations, check the task type, split criterion, binary versus multiway branching, handling of categorical or missing values, interpretability, stability, computational cost, and overfitting controls. A family name alone does not establish which approach will perform best on a particular dataset.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why can a decision tree overfit?

A tree can keep splitting until its leaves contain very few training examples. That may capture genuine patterns, but it can also capture noise or quirks specific to the training set. A fully grown tree can therefore look highly accurate on training data while generalizing poorly. Individual trees are also unstable: small changes in the data can produce different splits and a substantially different tree.

Control growth and choose complexity using data not used to fit the tree. Useful controls include:

  • max_depth: limits the number of levels.
  • min_samples_split: requires a minimum number of observations before a node can be split.
  • min_samples_leaf: requires a minimum number of observations in a leaf.
  • ccp_alpha: controls minimal cost-complexity post-pruning, which removes branches that do not justify their added complexity.

Pruning removes branches to simplify the model and can reduce overfitting; the right setting depends on the data and evaluation metric. Scikit-learn documents minimal cost-complexity pruning, and IBM’s decision-tree overview explains pruning as removing branches of low importance to reduce complexity and overfitting.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to build and evaluate a tree in scikit-learn

  1. Start with a shallow tree. Set a modest max_depth so the initial model is inspectable rather than allowing unrestricted growth.
  2. Fit on training data and visualize it. Inspect the questions, thresholds, and leaf predictions to see how the model divides the examples.
  3. Tune complexity on validation data or with cross-validation. Compare depth, min_samples_split, min_samples_leaf, and, when appropriate, ccp_alpha.
  4. Evaluate on held-out data. Report a metric appropriate to classification or regression; do not treat training accuracy alone as evidence that the model will generalize.

Impurity-based feature importance needs care: it can favor features with many possible split points, and an overfit tree can make its apparent importances unreliable. Check explanations on held-out data and consider permutation importance where appropriate. Scikit-learn’s permutation-importance guide describes that alternative.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When is a decision tree a good choice?

A single tree is useful when a readable sequence of rules matters, when relationships may be nonlinear, or when interactions between features should be discovered automatically. Trees generally need little feature scaling and can handle both classification and regression.

The trade-off is variance and limited predictive stability: a small data change can alter the learned structure, and greedy splitting does not guarantee a globally optimal tree. Assess performance with a held-out test set or cross-validation, and tune complexity on validation data rather than choosing a tree solely because its rules look plausible.

Decision tree or random forest?

Choose a single decision tree when compact, inspectable rules are a priority and its measured performance is adequate. Consider a random forest when robustness and predictive performance matter more than having one short explanation: it combines multiple trees, which generally reduces the instability of relying on one tree, but produces a less compact explanation.

Neither choice is automatically superior for every dataset. Compare them using the same train/validation procedure and task-appropriate metric, while accounting for whether a concise rule path or a more robust ensemble better suits the use case.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.