Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To calculate feature importance in Python, first confirm that your fitted model predicts well, then choose a method that matches your question. For scikit-learn tree models, feature_importances_ gives impurity-based importance from the fitted trees. For an evaluation-set view tied to a metric, use sklearn.inspection.permutation_importance to measure how much the model’s score falls when each feature is shuffled.

What feature importance tells you—and what it does not

A feature-importance score describes a fitted model’s reliance on input columns under a particular method, dataset, and (for permutation importance) scoring metric. It is not a causal effect, a universal measure of a feature’s value, or a guarantee that the same ranking will hold for another model or dataset.

Validate predictive performance before interpreting importance. As the scikit-learn documentation puts it, “Indeed, there would be little interest in inspecting the important features of a non-predictive model.” Its Titanic random-forest example reports training accuracy of 1.000 and test accuracy of 0.814; those are outputs from that illustration, not typical or expected results.

Choose between tree impurity importance and permutation importance

Method Estimator coverage and data basis What it measures Key limitations and cost
feature_importances_ (mean decrease in impurity, or MDI) Available on supported tree estimators; derived from splits in the fitted trees and their training data. How the fitted trees used features to reduce impurity during training. Inexpensive to read, but can favor high-cardinality features and features involved in overfitting.
permutation_importance Model-agnostic API; computed on the dataset you supply, such as held-out evaluation data. Change in a selected score after shuffling one feature column at a time. Requires repeated rescoring; results depend on the metric and dataset, and correlated features can mask one another.

Use MDI for a quick summary of how a tree ensemble made its training splits. Prefer held-out permutation importance when the question is how much disrupting a feature changes generalization-oriented performance. The methods answer different questions, so their rankings need not match.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to get feature importance from a Random Forest

After fitting a scikit-learn random forest, pair its feature_importances_ values with the input feature names, sort them, and plot or inspect the result. Label this as MDI: it describes feature use in the fitted trees, not a held-out score drop.

import pandas as pd
import matplotlib.pyplot as plt

# Assumes model is a fitted RandomForestClassifier or RandomForestRegressor
# and X_train is a DataFrame whose columns were used to fit model.
importance = pd.Series(
    model.feature_importances_,
    index=X_train.columns,
    name="mean decrease in impurity",
).sort_values(ascending=False)

importance.plot(kind="bar")
plt.ylabel("Mean decrease in impurity")
plt.tight_layout()
plt.show()

If the estimator was fit from an array rather than a DataFrame, use the exact feature-name sequence supplied at fit time. Confirm that its order and length match the estimator’s inputs before attaching names to scores.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

How to calculate permutation importance on held-out data

Keep an evaluation set that was not used to fit the model. The API evaluates the fitted estimator on the original data, shuffles one feature column at a time, and reports the score decrease across repeated shuffles. Choose a metric that reflects the task; setting random_state makes the shuffling reproducible.

  1. Fit the model using training data, keeping the evaluation data separate from fitting and model selection where possible.
  2. Call permutation_importance with the fitted estimator, evaluation features and labels, a task-appropriate scoring value, and explicit repeat and randomness settings.
  3. Sort importances_mean alongside the original feature names. Inspect importances_std and, when useful, the individual repeat values in importances to see variability.
import pandas as pd
from sklearn.inspection import permutation_importance

result = permutation_importance(
    model,
    X_test,
    y_test,
    scoring="accuracy",  # choose a metric appropriate to the task
    n_repeats=30,
    random_state=42,
    n_jobs=-1,
)

permutation_scores = pd.DataFrame({
    "feature": X_test.columns,
    "mean_score_decrease": result.importances_mean,
    "std": result.importances_std,
}).sort_values("mean_score_decrease", ascending=False)

print(permutation_scores)

The example uses accuracy for illustration; select another scorer when accuracy is not the objective you need to explain. With scoring=None, the API uses the estimator’s default score, and its default is five repeats. Explicitly setting the scorer and repeat count makes the analysis easier to interpret.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When preprocessing is part of the model, keep it in the fitted pipeline where appropriate so shuffled inputs pass through the same transformations as the model’s ordinary inputs. Check the API documentation for the estimator and data interface you are using.

Interpret scores with their main failure modes in mind

High-cardinality features and overfitting can inflate MDI

MDI can assign high importance to numerical or other high-cardinality features, including noise, because it is based on training-set splits. In the scikit-learn Titanic example, a random numerical feature receives misleadingly high MDI importance, while held-out permutation importance places it near zero. A held-out permutation check is more informative when the concern is whether the feature supports performance beyond training data.

Correlated predictors can make individual permutation scores look small

If two columns carry similar information, shuffling one may leave the model able to rely on the other. The remaining feature can preserve much of the score, making each individual permutation result look modest even when the model predicts well. The scikit-learn correlated-features example demonstrates this issue with the Breast Cancer Wisconsin diagnostic dataset. Consider a defensible grouping or feature-selection strategy when the interpretation question concerns shared information; do not treat every low individual score as proof of irrelevance.

The selected metric changes the answer

Permutation importance is specific to the scorer: a feature may have a larger effect on accuracy than on a different objective. State the metric used. The API can calculate results for multiple scorers, allowing comparison under more than one objective.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Repeats and sample size affect time and variability

More repeats require more model evaluations. The API exposes n_repeats, n_jobs, and max_samples; using fewer samples can reduce runtime but may make estimates less accurate. Review the repeat-level values and variability rather than presenting a single mean as exact.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why feature-importance rankings differ

  • Different question: MDI summarizes training split use; permutation importance estimates score change on the supplied evaluation data.
  • Different metric: A ranking based on accuracy may differ from one based on another scorer.
  • Different data: Importance depends on the dataset used, including its sample composition and feature relationships.
  • Shared information: Correlated predictors can split or mask individual permutation scores.
  • Model and fit variation: A ranking belongs to the fitted model and can change with another model or fit.

For reproducibility, record the estimator, dataset split, feature names and order, metric, repeat count, randomness setting, and relevant scikit-learn version. The examples here use the stable scikit-learn documentation available on October 4, 2026; check the documentation matching your installed version because APIs and documentation can change.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.