Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Naive Bayes is a supervised classification algorithm that uses Bayes’ theorem to estimate which class is most likely for each example. Its “naive” assumption treats features as conditionally independent once the class is known. That simplification is often useful, but it is a modeling assumption—not a claim that real-world features are truly unrelated.

This six-step tutorial explains the idea, helps you choose a scikit-learn variant, and builds a complete training-and-evaluation workflow in Python.

Step 1: Understand the classification problem

Classification starts with labeled examples. Each row contains input features X, and a target label y identifies the class. For example, a model might classify a flower species from measurements, or decide whether an email is spam from word features.

  • Features (X): the measurements or encoded attributes used to make a prediction.
  • Labels (y): the known classes learned during training.
  • Training data: examples used to estimate class probabilities and feature likelihoods.
  • Test data: held-out examples used only after training to estimate generalization.

Naive Bayes is a classifier, so it predicts a discrete class rather than a continuous number.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Step 2: See how Bayes’ theorem becomes a classifier

For a feature vector x and class c, Bayes’ theorem is:

P(c | x) = P(x | c) P(c) / P(x)

P(c | x) is the posterior probability of the class after observing the features. P(c) is the prior probability of that class, and P(x | c) is the likelihood of seeing the features in that class.

Naive Bayes simplifies the likelihood by assuming conditional independence:

P(x | c) = P(x1 | c) × P(x2 | c) × … × P(xn | c)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The estimator calculates a score for every class and predicts the class with the greatest score. Because probabilities can become extremely small when many features are multiplied, implementations generally perform the calculation in log space.

Step 3: Match the estimator to your data

Scikit-learn provides several Naive Bayes estimators. Select one from the representation and distribution of your features, then validate that choice on held-out data.

Estimator Best fit Input example Important detail
GaussianNB Continuous features whose class-conditional distributions can be approximated as Gaussian Measurements such as length, weight, or sensor values Works with numeric arrays; the Gaussian assumption is an approximation.
MultinomialNB Multinomial or count-style features Word-count vectors for text TF-IDF features can also work in practice; validate the representation.
BernoulliNB Binary-valued features Indicators showing whether a word occurs Models both feature presence and non-occurrence.
CategoricalNB Categorical feature distributions Encoded categories such as color or browser type Each feature must use non-negative integer category indices.
ComplementNB A MultinomialNB adaptation useful to investigate with imbalanced data Imbalanced count-style text features It is not automatically best; compare it on your task.

For the beginner example below, the Iris dataset supplies continuous measurements, so GaussianNB is a natural starting point.

Step 4: Prepare data and create a leakage-safe split

Install the libraries if necessary:

python -m pip install scikit-learn

Then load labeled data and reserve test examples before fitting the model:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

iris = load_iris()
X = iris.data
y = iris.target

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.2,
    random_state=42,
    stratify=y,
)

test_size=0.2 reserves 20 percent for evaluation. random_state=42 makes this particular split repeatable, while stratify=y preserves class proportions. These settings are choices for this example, not universal requirements.

If preprocessing learns values from the data—such as scaling, imputation, feature selection, or vocabulary construction—fit it only on the training portion. A scikit-learn Pipeline is the safest way to keep those operations inside the training workflow.

Step 5: Fit Gaussian Naive Bayes and make predictions

Instantiate the estimator, train it with the training data, and predict labels for examples the model has not seen:

from sklearn.naive_bayes import GaussianNB

model = GaussianNB()
model.fit(X_train, y_train)

y_pred = model.predict(X_test)

print("Predicted labels:", y_pred)
print("Actual labels:   ", y_test)

You can also inspect class probabilities. The columns returned by predict_proba follow model.classes_:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
probabilities = model.predict_proba(X_test)

print("Classes:", model.classes_)
print("First example probabilities:", probabilities[0])

For a new Iris measurement, pass a two-dimensional array because scikit-learn expects one row per example:

new_flower = [[5.1, 3.5, 1.4, 0.2]]
predicted_class = model.predict(new_flower)[0]
print("Predicted class index:", predicted_class)
print("Predicted class name:", iris.target_names[predicted_class])

Step 6: Evaluate the held-out predictions

Accuracy is easy to read when classes are reasonably balanced, but it should not be the only metric for an imbalanced problem. Print a classification report and a confusion matrix:

from sklearn.metrics import accuracy_score, classification_report, confusion_matrix

print("Accuracy:", accuracy_score(y_test, y_pred))
print("nClassification report:")
print(
    classification_report(
        y_test,
        y_pred,
        target_names=iris.target_names,
    )
)
print("Confusion matrix:")
print(confusion_matrix(y_test, y_pred))

The report includes precision, recall, F1 score, and support for each class. A confusion matrix shows which classes were confused. Do not treat the score from this one split as a universal performance guarantee; change the split or use cross-validation when you need a more reliable estimate.

Choosing between plausible Naive Bayes variants

When more than one estimator fits your data, compare them using the same held-out split and metric. For text classification, a count representation commonly pairs with MultinomialNB, while word-occurrence indicators pair naturally with BernoulliNB.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import MultinomialNB
from sklearn.pipeline import make_pipeline

texts = [
    "cheap offer available now",
    "team meeting moved to tomorrow",
    "claim your prize today",
    "please review the project report",
]
labels = ["spam", "work", "spam", "work"]

X_train, X_test, y_train, y_test = train_test_split(
    texts,
    labels,
    test_size=0.5,
    random_state=42,
    stratify=labels,
)

text_model = make_pipeline(
    CountVectorizer(),
    MultinomialNB(),
)
text_model.fit(X_train, y_train)
print(text_model.predict(X_test))

The example demonstrates the workflow, not a meaningful benchmark: four documents are far too few to establish production performance. With a real corpus, keep vectorization in the pipeline so vocabulary statistics are learned from training data only.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Limitations and practical checks

  • Feature dependence: correlated features can violate the naive independence assumption. The model may still work, but compare alternatives on your data.
  • Probability interpretation: predicted probabilities can be poorly calibrated even when class ranking is useful. Calibrate only when your application needs trustworthy probability values.
  • Preprocessing leakage: never let test-set statistics influence transformations fitted during training.
  • Class imbalance: inspect per-class precision and recall rather than relying only on accuracy; ComplementNB is one candidate to evaluate for imbalanced count data.
  • Variant mismatch: do not feed arbitrary encodings into an estimator without checking its distributional assumptions.

Incremental fitting for larger datasets

MultinomialNB, BernoulliNB, and GaussianNB expose partial_fit for incremental learning. On the first call, provide the complete list of possible class labels so later batches use a consistent class space:

import numpy as np
from sklearn.naive_bayes import GaussianNB

classes = np.unique(y)
stream_model = GaussianNB()

stream_model.partial_fit(X_train[:50], y_train[:50], classes=classes)
stream_model.partial_fit(X_train[50:], y_train[50:])

Use this pattern only when batches genuinely arrive over time or the full dataset does not fit comfortably in memory. Keep a separate evaluation set that is not used for incremental updates.

What to learn next

After this tutorial, practice changing one element at a time: try a different split, inspect the confusion matrix, compare GaussianNB with a reasonable baseline, or build a text pipeline with MultinomialNB and BernoulliNB. For a broader companion, O’Reilly lists Introduction to Machine Learning with Python by Andreas C. Müller and Sarah Guido as a beginner-to-intermediate, 400-page practical guide focused on Python and scikit-learn. It was first published in October 2016, so verify examples against the scikit-learn version you use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.