Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To build a semantic-search baseline with Sentence Transformers, encode each passage into a vector, encode an incoming query, and rank the passages by similarity to that query. This tutorial uses an asymmetric retrieval task—a short question searching longer passages—and shows the basic workflow, its limits, and how to add an optional reranking stage.

What this baseline does

A bi-encoder maps each text to a fixed-size vector. Once the corpus is encoded, the query vector can be compared with those stored passage vectors to find semantically related candidates. Sentence Transformers describes this as an efficient first stage for semantic retrieval; it is not a guarantee that the top-ranked passage is correct.

The example below is asymmetric retrieval: a short user question is matched against longer answer passages. That differs from symmetric search, such as finding questions similar in length and form to a query. A model that suits one task is not automatically the best choice for the other. See the Sentence Transformers semantic-search guide.

Build a small, inspectable corpus

Keep a stable ID beside each passage so ranked results can be mapped back to their original text. Chunking also affects what the system can retrieve: a very broad passage may contain several topics, while a tiny fragment may omit context needed to judge relevance. The sample uses short passages for clarity; choose chunk boundaries that preserve useful context in your own material.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
corpus = [
    {"id": "p1", "text": "Semantic search retrieves text based on meaning, not only exact word matches."},
    {"id": "p2", "text": "A bi-encoder converts a query and each document into vectors that can be compared."},
    {"id": "p3", "text": "A CrossEncoder scores a query and candidate passage together for reranking."},
]
query = "What is semantic search?"

Encode passages and the query

Install the package in your Python environment with pip install -U sentence-transformers. Then load a model suited to the retrieval task and encode passages with encode_document() and the incoming question with encode_query().

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("sentence-transformers/multi-qa-mpnet-base-cos-v1")
documents = [item["text"] for item in corpus]
document_embeddings = model.encode_document(documents)
query_embedding = model.encode_query(query)

multi-qa-mpnet-base-cos-v1 is one catalog example specifically trained for semantic search, not a universal best model. Sentence Transformers recommends task-appropriate models for retrieval. The query/document methods can apply prompts or task routing when a model defines them; for models without specialized prompts or task settings, they may behave the same as encode(). Consult the project’s usage guide and semantic-search guide when selecting a model and encoding method.

Rank passages by similarity

Use the model’s similarity function, then sort the resulting scores from highest to lowest. The indices refer to the original corpus order, so they can be used to recover IDs and text. Similarity scores are ranking signals, not calibrated probabilities that a result is relevant.

scores = model.similarity(query_embedding, document_embeddings)[0]
ranked_indices = scores.argsort(descending=True)

for index in ranked_indices:
    item = corpus[int(index)]
    print(item["id"], float(scores[index]), item["text"])

For larger collections, the Sentence Transformers retrieval utilities provide a documented alternative to manually comparing all embeddings. The project’s documentation describes its manual embedding-and-similarity route as suitable for small corpora up to about one million entries. Treat that as approximate guidance, not a hardware-independent capacity limit or latency promise. For a production collection, assess indexing, memory, update patterns, and response-time needs against your own data and workload.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Add a CrossEncoder reranking stage when useful

A bi-encoder scores query and passage separately, making it practical to encode passages ahead of time. A CrossEncoder instead scores a query and candidate passage together. A common retrieve-and-rerank workflow first obtains candidates with lexical search or dense bi-encoder retrieval, then applies CrossEncoder scores to reorder that smaller set. This adds pairwise inference, so it costs more computation than returning the first-stage ranking alone; whether the tradeoff is worthwhile depends on the task and must be evaluated. The retrieve-and-rerank guide describes both lexical and dense candidate retrieval.

from sentence_transformers import CrossEncoder

# `candidate_texts` should contain passages returned by your first-stage retriever.
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
pairs = [(query, text) for text in candidate_texts]
rerank_scores = reranker.predict(pairs)
reranked = sorted(
    zip(rerank_scores, candidate_texts),
    key=lambda item: item[0],
    reverse=True,
)

This snippet illustrates the second-stage pattern; choose a CrossEncoder appropriate to your data and verify the model’s expected use. The scores from a reranker are also not automatically probabilities.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Evaluate before choosing a model or configuration

Run representative queries from the actual use case and judge whether the retrieved passages answer them. Compare candidate models and configurations on those same queries, and record relevance judgments so changes in ranking can be assessed. The documentation examples and model catalog identify possible approaches, but do not establish which model wins on your corpus.

  • Choose between symmetric and asymmetric retrieval based on the relationship between query and candidate text.
  • Compare lexical retrieval, dense retrieval, or a combination as candidate-generation methods.
  • Check whether the model uses query/document prompts or task routing and call its retrieval methods accordingly.
  • Account for collection size, memory, latency, and how the corpus changes.
  • Test whether CrossEncoder reranking justifies the extra scoring work for your relevance needs.

These choices are covered across the Sentence Transformers semantic-search, retrieve-and-rerank, and retrieval API documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.