Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To do named entity recognition (NER) with BERT, fine-tune a token-classification model on text whose words have entity labels. The essential steps are to choose a suitable labeled dataset, align its word-level labels to BERT’s subword tokens, train the model, evaluate entity-level precision, recall, and F1, then load the checkpoint for inference. Hugging Face documents a BERT example using CoNLL-2003; its current token-classification tutorial demonstrates the same workflow mechanics with DistilBERT.

What BERT does in an NER model

NER identifies spans of text and assigns them categories such as person, organization, or location. In a BERT implementation, this is a token-classification task: the model predicts a label for each input token, and those labels are interpreted together to identify entity spans. The label inventory and conventions come from your dataset; they are not built into BERT.

Hugging Face’s token-classification guide walks through the current API pattern using WNUT 17 and DistilBERT. Its Transformers repository example documents fine-tuning BERT on CoNLL-2003 and supports custom train and validation files. The code pattern below follows the guide’s tokenization and alignment approach while using a BERT checkpoint; confirm that your selected checkpoint and tokenizer are compatible with token classification.

Choose a dataset and label scheme

Training data should resemble the text and entity types the model will encounter after deployment. For example, the Hugging Face tutorial uses WNUT 17, which includes emerging entities. The repository example uses CoNLL-2003. Neither dataset is automatically the right choice for another domain, language, or annotation policy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Choice Documented example What to check
WNUT 17 The current token-classification guide loads flaitenberger/wnut_17; its sample data has tokens and integer ner_tags. Check whether its entity types and text match your intended use.
CoNLL-2003 The repository example pairs google-bert/bert-base-uncased with tomaarsen/conll2003. Check the dataset’s annotation conventions, licensing or terms, and fit for your task.
Custom data The repository example includes a path for custom training and validation files. Confirm that preprocessing yields token sequences and labels in the format expected by the training script.

Before training, inspect the dataset’s label names and how entities are represented. BIO-style schemes commonly distinguish the beginning and inside of an entity—such as B-PER and I-PER—from non-entity tokens marked O. Use the exact inventory your data provides rather than assuming every corpus has identical classes.

Install the software dependencies

The Hugging Face guide lists these packages for its walkthrough:

  • transformers for the model, tokenizer, and training utilities
  • datasets for loading and handling data
  • evaluate for computing metrics
  • seqeval for sequence-labeling evaluation

Install them in the Python environment you plan to use, following the package installation instructions for your platform. The documented workflow does not establish a specific hardware requirement.

Tokenize the text and align labels to BERT subwords

A dataset may label whole words, while BERT’s tokenizer can split a word into multiple subword tokens. It also inserts special tokens. The training labels therefore need to line up with the tokenized input positions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Tokenize each example with the tokenizer associated with the BERT checkpoint, preserving the original word boundaries. For word-tokenized data, use the tokenizer’s pre-tokenized input mode.
  2. Use the tokenizer’s word_ids() mapping to find which source word produced each tokenized position.
  3. Assign -100 to special tokens so the loss ignores them.
  4. For each source word, keep its original NER label on the first subtoken and assign -100 to any later subtokens from that word.

The ignored value -100 is the convention used in the guide’s illustrated alignment method. It prevents special tokens and repeated subword pieces from contributing separate labels to the loss. Other label-propagation schemes are possible, but training and evaluation must use the same convention.

Configure the BERT token-classification model

Build mappings between each label string and its integer ID, and pass both mappings and the label count when loading the model. The model’s output head must have one class for every label in your dataset.

from transformers import AutoModelForTokenClassification, AutoTokenizer

checkpoint = "google-bert/bert-base-uncased"
labels = ["O", "B-PER", "I-PER", "B-ORG", "I-ORG"]  # Replace with your dataset's complete label list.
id2label = {i: label for i, label in enumerate(labels)}
label2id = {label: i for i, label in id2label.items()}

tokenizer = AutoTokenizer.from_pretrained(checkpoint)
model = AutoModelForTokenClassification.from_pretrained(
    checkpoint,
    num_labels=len(labels),
    id2label=id2label,
    label2id=label2id,
)

The label list shown is illustrative, not a complete inventory for WNUT 17 or CoNLL-2003. Replace it with the exact labels from your chosen dataset. A mismatch between the label IDs in preprocessing and the model mappings can make both training and predictions incorrect.

Fine-tune the model

Prepare the training and validation splits, tokenize and align their labels, then pass the processed data to the Transformers training workflow. The guide’s displayed settings are an example configuration, not a universal recommendation: learning rate 2e-5, per-device training and evaluation batch sizes of 16, 2 epochs, and weight decay 0.01. Your suitable settings depend on your data and training environment; those example values do not guarantee a particular result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The repository example uses a BERT checkpoint with CoNLL-2003 through run_ner.py. Its script relies on fast-tokenizer features, so verify tokenizer compatibility if you adapt that path. For a custom dataset, ensure the train and validation files, label names, and preprocessing all agree before starting the run.

Evaluate entity recognition, not just token accuracy

Use the held-out validation or test split to measure how well the model identifies entities. The Hugging Face guide uses Evaluate’s seqeval metric and reports precision, recall, F1, and accuracy after excluding positions labeled -100.

  • Precision indicates how many predicted entities are correct.
  • Recall indicates how many reference entities the model finds.
  • F1 combines precision and recall, making it a useful summary when assessing entity extraction.
  • Accuracy measures correct labels at token positions, but by itself does not show whether complete entity spans were identified correctly.

Report the dataset, split, label scheme, and metric with any result. Scores from different corpora or annotation schemes are not directly comparable, and the cited implementation examples do not establish a general expected BERT NER score.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Run inference with the fine-tuned checkpoint

For straightforward predictions, load the saved model through the Transformers NER pipeline and pass it text:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from transformers import pipeline

ner = pipeline("ner", model="path/to/saved-checkpoint", tokenizer="path/to/saved-checkpoint")
results = ner("Ada Lovelace worked with Charles Babbage in London.")
print(results)

Replace the paths with the directory where you saved the fine-tuned model and tokenizer. The guide’s pipeline output includes token text, a predicted label, a confidence score, and character start and end positions. Depending on the aggregation setting, output may represent individual tokens or merged spans.

The Hugging Face Inference Providers token-classification guide describes these aggregation choices:

Strategy Effect
none Leaves predictions ungrouped at token level.
simple Groups consecutive tokens with the same label.
first Preserves word integrity by using the first token’s label.
average Uses averaged scores across a word.
max Uses the highest score across a word.

Choose output granularity based on what the consuming application needs. A token-level result can contain subword fragments; do not present those fragments as separate complete entities without interpreting or grouping them.

Common implementation checks

  • Predictions use unexpected labels: verify that label2id, id2label, and the dataset’s integer tag mapping use the same ordering.
  • Subword tokens receive inconsistent treatment: confirm that word IDs are used during alignment and that the chosen scheme is applied consistently in training and evaluation.
  • Metrics look misleadingly strong: inspect entity-level precision, recall, and F1 in addition to token accuracy, especially when non-entity O labels dominate.
  • Entities appear split in inference results: check the selected aggregation strategy and whether the pipeline is returning token-level or grouped predictions.
  • Results do not transfer to your application: assess whether training and evaluation data match the target domain, language, text style, and entity definitions.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.