Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To generate a sentence embedding with Transformers, tokenize the text, run a compatible model to get contextual token vectors, then apply the pooling and output-processing steps specified for that checkpoint. In the sentence-transformers/all-mpnet-base-v2 model card, the example uses attention-mask-aware mean pooling followed by L2 normalization. The model’s token outputs are not, by themselves, a sentence embedding.

What a text embedding represents

A Transformer produces a contextual representation for each token in an input. Its hidden states have batch, sequence-length, and hidden-size axes; the sequence axis means an input can produce many token vectors, not one sentence vector. A pooling operation combines those token representations into one fixed-size vector for each text.

That distinction matters when using the general feature-extraction interface: it exposes model features, but does not automatically guarantee a task-appropriate sentence embedding. The pooling and any normalization depend on the checkpoint and intended use. See the Transformers feature-extraction pipeline documentation and the model’s own instructions.

Generate embeddings with all-mpnet-base-v2

The following implementation follows the model card’s example. It loads the checkpoint’s tokenizer and base model, tokenizes a batch with padding and truncation, computes token representations, pools only unmasked positions, and normalizes each resulting vector.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import torch
import torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer

model_name = "sentence-transformers/all-mpnet-base-v2"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name)

sentences = [
    "Transformers produce contextual token representations.",
    "Pooling combines token representations into a sentence vector.",
]

encoded_input = tokenizer(
    sentences,
    padding=True,
    truncation=True,
    return_tensors="pt",
)

with torch.no_grad():
    model_output = model(**encoded_input)

# Exclude padding positions from the mean.
token_embeddings = model_output[0]
input_mask_expanded = (
    encoded_input["attention_mask"]
    .unsqueeze(-1)
    .expand(token_embeddings.size())
    .float()
)
sum_embeddings = torch.sum(token_embeddings * input_mask_expanded, dim=1)
sum_mask = torch.clamp(input_mask_expanded.sum(dim=1), min=1e-9)
mean_embeddings = sum_embeddings / sum_mask

# The model-card example normalizes each sentence vector.
sentence_embeddings = F.normalize(mean_embeddings, p=2, dim=1)

print(sentence_embeddings.shape)

For a batch of two sentences, the output has one vector per sentence. The model-card example’s pooling calculation expands the attention mask to match the token-vector dimensions, weights token representations by that mask, sums along the sequence axis, and divides by the number of unmasked positions. The small clamp minimum prevents division by zero. Its final normalization operates along the embedding dimension.

Why the attention mask matters

Batch tokenization commonly pads shorter inputs so every item has the same sequence length. Those padding positions are not part of the sentence. If a mean includes them, the result is influenced by padding rather than only by the input text. In the example above, the attention mask gives real tokens weight 1 and padded positions weight 0 before averaging.

Do not assume every model uses this recipe

The all-mpnet-base-v2 card specifies mean pooling and normalization for its example; that is not a universal Transformers rule. Another checkpoint may require a different pooling strategy, input formatting, or output treatment. The model card describes the all-mpnet-base-v2 sequence this way: “First, you pass your input through the transformer model, then you have to apply the right pooling-operation on-top of the contextualized word embeddings.”

Before using a checkpoint, check its Hub model card and metadata. Hugging Face describes model cards as a place for examples, architecture information, and metadata such as license. Follow the checkpoint’s directions rather than applying the code above unchanged to an unrelated model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose an embedding approach for the intended task

Sentence embeddings are used for tasks such as semantic search, clustering, and retrieval. The appropriate checkpoint and output handling depend on what you need the vectors to do. Use these checks when evaluating an approach:

  • Task alignment: Confirm that the checkpoint is intended for sentence similarity, retrieval, or your relevant downstream objective.
  • Pooling contract: Look for the pooling method required by the model card; do not assume mean pooling or a first-token representation is interchangeable.
  • Input handling: Follow the tokenizer, truncation behavior, padding, attention-mask use, and any model-specific formatting instructions.
  • Output handling: Check the vector dimension and whether the checkpoint’s directions call for normalized embeddings before your similarity calculation.
  • License and provenance: Review the Hub metadata and model card before adopting a checkpoint.

These checks identify what to evaluate; they do not establish a best model for a particular language, domain, latency target, or retrieval benchmark. Choose among candidates with evaluation data representative of your own use case rather than assuming one checkpoint wins universally.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What to do with the resulting vectors

Once you have one vector per text, you can use those representations in a semantic search, clustering, or retrieval workflow. Keep the model and preprocessing consistent when creating vectors for both stored documents and later queries. If you change checkpoints or the embedding recipe, treat the resulting vectors as a different representation and evaluate the effect on your application.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.