Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can build a basic spam-versus-ham classifier in Python with labeled message text, TfidfVectorizer, and scikit-learn’s MultinomialNB. The example below uses the UCI SMS Spam Collection to demonstrate the workflow; it is a reproducible starting point, not proof of how a model will perform on a current email inbox.

What you need to build a spam filter

A text classifier needs four pieces: labeled examples, a way to turn text into numeric features, a classification model, and an evaluation method. In this example, the labels are ham (not spam) and spam; TF-IDF supplies the features, and Naive Bayes makes the predictions.

The example uses the UCI SMS Spam Collection, which UCI describes as a public set of labeled SMS messages collected for mobile-phone spam research. UCI reports 5,574 instances and a donation date of June 21, 2012. Each line contains the correct class followed by the raw message. The collection is useful for learning binary text classification, but SMS is not equivalent to a modern email stream with headers, HTML, attachments, and different language or attack patterns. See the SMS Spam Collection dataset at UCI. The dataset page lists Almeida, Hidalgo, and Yamakami’s 2011 paper, Contributions to the study of SMS spam filtering: new collection and results, as an introductory reference. View the paper.

Load and inspect the labeled messages

Download the dataset file and place it at SMSSpamCollection relative to the Python script or notebook. The tab-separated format has a label, then a message; splitting only at the first tab preserves any tabs that may occur in message text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from pathlib import Path
import pandas as pd

rows = []
for line in Path("SMSSpamCollection").read_text(encoding="utf-8").splitlines():
    label, message = line.split("t", 1)
    rows.append((label, message))

df = pd.DataFrame(rows, columns=["label", "message"])
print(df["label"].value_counts())

Check that the labels and message column look as expected before training. A split that preserves the proportions of both labels helps ensure that spam and ham examples appear in both training and test data.

Build the TF-IDF and Naive Bayes pipeline

TfidfVectorizer converts raw documents into a TF-IDF feature matrix. With its standard settings, it lowercases text, tokenizes words, applies smoothed inverse document frequency, and normalizes each row with the L2 norm. Conceptually, TF-IDF combines how often a term occurs in a message with how uncommon it is across the training corpus, so a term found in almost every message generally carries less distinguishing weight than one concentrated in a smaller subset. Scikit-learn documents the smoothed IDF calculation as log((1 + n) / (1 + df)) + 1; resulting weights depend on the training corpus and vectorizer configuration. Read scikit-learn’s text feature extraction documentation.

A Pipeline keeps vectorization and classification together. When the pipeline is fit on training messages, the vectorizer learns its vocabulary and IDF values from those messages rather than from the entire dataset. This prevents test-set information from influencing training. Scikit-learn’s text tutorial demonstrates this general pattern of chaining vectorization and Naive Bayes. See the scikit-learn text-classification tutorial.

from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.naive_bayes import MultinomialNB
from sklearn.metrics import classification_report, confusion_matrix

X_train, X_test, y_train, y_test = train_test_split(
    df["message"],
    df["label"],
    test_size=0.20,
    random_state=42,
    stratify=df["label"],
)

model = Pipeline([
    ("tfidf", TfidfVectorizer(
        lowercase=True,
        ngram_range=(1, 2),
        min_df=1,
    )),
    ("classifier", MultinomialNB()),
])

model.fit(X_train, y_train)
predicted = model.predict(X_test)

print(classification_report(y_test, predicted, digits=3))
print(confusion_matrix(y_test, predicted, labels=["ham", "spam"]))

The 20% test split, seed of 42, and stratification rule make this run’s split reproducible for the same data and software behavior. They do not make its score universal. MultinomialNB is a transparent baseline for sparse text features; treat it as a point of comparison, not a guaranteed production choice.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Evaluate errors, not just a headline score

The classification report gives per-class precision, recall, and F1, as well as overall summaries. The confusion matrix uses the explicit order ham, then spam: rows are the true labels and columns are predicted labels. Read the off-diagonal counts as the two kinds of mistake:

  • A ham message predicted as spam is a false positive: a wanted message may be hidden or diverted.
  • A spam message predicted as ham is a false negative: unwanted mail remains visible.

Which error matters more depends on the mailbox and how filtered messages are handled. Decide that before adjusting the model or a decision threshold. Keep the test set untouched until the end; if you tune parameters, do so with cross-validation on training data and reserve the test set for the final evaluation. No performance score is assumed here: use the metrics from your own run, and record the corpus version, label mapping, split rule, and random seed so the result has context.

Try the trained model on new messages

After fitting, pass a list of raw strings to the pipeline. It applies the already-fitted vectorizer and classifier in sequence and returns one of the labels it learned.

examples = [
    "Congratulations, you have won a prize. Call now!",
    "Can we meet for lunch tomorrow?",
]
print(model.predict(examples))

These sample predictions illustrate the interface, not a guarantee that every similar message will be classified correctly. Inspect errors from validation data that resembles the messages the filter is meant to handle.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose useful follow-up experiments

Once the baseline runs, compare alternatives on the same validation procedure rather than assuming one feature or model is better. Scikit-learn’s vectorizer supports word and character analyzers, as well as controls such as ngram_range, min_df, max_df, and max_features. Check the TfidfVectorizer API.

  • Compare word unigrams with word unigrams and bigrams.
  • Compare word features with character or character-boundary n-grams. Character features may be worth testing for obfuscated text, but measure their effect on held-out messages.
  • Compare MultinomialNB with a linear classifier using training-only tuning and the same evaluation data.
  • Compare spam and ham precision and recall, not only one aggregate metric.
  • For a deployment candidate, measure training time, model size, inference latency, and behavior on obfuscation, HTML, and changing message distributions.

What this example does not cover in production

This model classifies the text strings supplied to it. It does not parse MIME structure, inspect attachments safely, authenticate senders, manage allowlists, or incorporate user feedback. A deployed email filter needs representative, consented labeled email data and operational safeguards appropriate to the organization.

  • Use subject and body fields from representative email rather than assuming SMS performance transfers to email.
  • Set privacy controls for message data and restrict access to training and prediction logs.
  • Log model and data versions, monitor abuse and changes in message patterns, and review false positives before making filtering more aggressive.
  • Check performance over time and retrain when the message distribution changes.

The 2012 SMS corpus is an educational dataset, not evidence of current email-filter accuracy. An organization can replace the SMS file with its own appropriately handled labeled email fields while keeping the leakage-safe pipeline and disciplined evaluation approach.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.