Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A naive search scorer can rank a document higher simply because it repeats a query term: with one query token, a document containing “python” three times can outscore one containing it once. BM25F offers a more structured alternative by normalizing term frequency separately in fields such as title and body, weighting those fields, and limiting the gain from repetition. It can mitigate this failure mode, but it cannot guarantee a better ranking without suitable fields, parameters, and relevance judgments.

Why repetition can win a naive search

Consider a simple scorer that adds a document’s count for each query term:

score(doc, query) = sum(count(term, doc) for term in query)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For the example below, the query is treated as one token, python, rather than three duplicate query tokens. The scorer therefore rewards repeated occurrences in the document, not repetition in the query:

Hypothetical document Occurrences in document Raw-count score
“Python” appears three times in the body 3 3
“Python” appears once in the title 1 1

Under this deliberately simple rule, the first document wins. That does not establish that any particular search engine or live search for “python python python” produces that result: the title alone specifies no corpus, candidate documents, tokenizer, or ranking system. Real implementations may also deduplicate query terms or count duplicates differently.

How BM25F changes the scoring

BM25-family methods temper the value of additional occurrences through term-frequency saturation and account for document length. BM25F extends that approach across multiple document fields, or streams, such as title and body. It normalizes each field against its own length and the collection’s average length, multiplies the normalized frequency by a field weight, combines the fields, and then applies saturation. A relevance-weighted IDF term can also affect the contribution of each query term.

For field s, the length-normalization factor in the reviewed formulation is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

B_s = (1 - b_s) + b_s * (field_length / average_field_length)

The parameter b_s controls how strongly that field’s length affects normalization. A field’s term frequency is adjusted using this factor, scaled by its weight, and combined with the adjusted frequencies from other fields before saturation. This is field-aware term-frequency scoring—not a claim that the system understands a word’s meaning or context.

In practical terms, a title match can be assigned more weight than a body match, while a long field’s raw count is not automatically treated as equivalent to the same count in a short field. Because the combined frequency saturates, each additional occurrence has diminishing impact instead of adding the same amount forever.

Naive counts and BM25F compared

Scoring aspect Naive raw-count scorer BM25F
Term frequency Adds occurrences according to its counting rule; repetition can keep increasing the score. Applies saturation, so extra occurrences have diminishing influence.
Document length May give longer documents more chances to accumulate matches. Normalizes term frequency separately by field length and average field length.
Document structure Often treats content as undifferentiated text. Can combine weighted fields such as title and body.
IDF May omit term rarity altogether. Includes an inverse-document-frequency component in the reviewed formulation.
Tuning May have few relevance-specific settings. Requires choices for field weights and per-field normalization, which need evaluation on the target collection.

Implement a compact BM25F scorer in pure Python

This example uses only Python’s standard library: dictionaries, regular expressions, and sorting. It assumes documents already have consistently parsed title and body strings. The lowercase word tokenizer is intentionally simple; production systems should use tokenization and normalization suited to their language and corpus.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The implementation calculates field lengths and average lengths from the same collection, then combines weighted, length-adjusted field frequencies for each query term before applying saturation. Its IDF calculation uses document frequency across the collection.

import math
import re


def tokenize(text):
    return re.findall(r"w+", text.lower())


def bm25f(documents, query, k=1.5, b=None, weights=None):
    fields = ("title", "body")
    b = b or {"title": 0.75, "body": 0.75}
    weights = weights or {"title": 3.0, "body": 1.0}

    tokens = {
        field: [tokenize(doc.get(field, "")) for doc in documents]
        for field in fields
    }
    lengths = {
        field: [len(doc_tokens) for doc_tokens in tokens[field]]
        for field in fields
    }
    averages = {
        field: (sum(lengths[field]) / len(documents) if documents else 0.0)
        for field in fields
    }

    # Count each distinct query token once; duplicate query words do not
    # multiply their contribution in this example.
    query_terms = set(tokenize(query))
    scores = [0.0] * len(documents)
    n_docs = len(documents)

    for term in query_terms:
        df = sum(
            any(term in tokenize(doc.get(field, "")) for field in fields)
            for doc in documents
        )
        if not df:
            continue
        idf = math.log(1.0 + (n_docs - df + 0.5) / (df + 0.5))

        for i in range(n_docs):
            combined_tf = 0.0
            for field in fields:
                field_tokens = tokens[field][i]
                tf = field_tokens.count(term)
                avg_len = averages[field]
                norm = (1.0 - b[field])
                if avg_len:
                    norm += b[field] * lengths[field][i] / avg_len
                if norm:
                    combined_tf += weights[field] * tf / norm

            if combined_tf:
                scores[i] += idf * (combined_tf * (k + 1.0)) / (combined_tf + k)

    return sorted(enumerate(scores), key=lambda pair: pair[1], reverse=True)


documents = [
    {"title": "Snakes and habitats", "body": "python python python"},
    {"title": "Python programming language", "body": "A guide to code"},
]

for index, score in bm25f(documents, "python"):
    print(index, score)

The shown parameters, k=1.5, b=0.75 for each field, and title/body weights of 3.0 and 1.0, are example values documented by the BM25-Search project, not universal recommendations. The code is a compact teaching implementation rather than a tested production package; it makes explicit choices about tokenization, duplicate query terms, field names, and collection-level IDF.

Its printed scores depend on the exact code, sample documents, and parameters, so they should be treated as outputs of this illustrative example—not as measured evidence that BM25F improves search quality. For a small corpus, compare those outputs with raw-count scores on the same documents and query, inspect the per-field contributions, and judge rankings against relevance labels appropriate to the task.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose fields and parameters for the collection

A title boost is a modeling assumption: it says that a match in a title should count more for this search task. It is not universally correct. Choose fields that are consistently available and meaningful, then evaluate weights and normalization settings against relevance judgments. If documents have no useful structure, BM25F’s field-weighting advantage may be limited.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The reviewed BM25F account describes collection-wide IDF, but warns that it can create degenerate cases when one field is unusually verbose and contains most terms for most documents. That makes field design and corpus statistics important, rather than details to set once and ignore.

For the probabilistic framework and its equations, see Robertson and Zaragoza’s 2009 review, “The Probabilistic Relevance Framework: BM25 and Beyond”. For a Python-specific field example and its documented parameter values, see the BM25-Search project documentation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.