Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To train BERT to predict masked words in Keras, choose one of two paths: build a small BERT-like encoder to learn the mechanics, or use KerasHub’s BertMaskedLM task with a BERT preset for a more streamlined workflow. The first path is educational and compact; it does not reproduce full-scale BERT pretraining. The second provides a documented MLM task and preprocessing route, but does not by itself add every objective used in original BERT pretraining.

What masked language modeling trains

Masked language modeling (MLM) is a self-supervised objective: select token positions in a text sequence, hide or otherwise corrupt the corresponding inputs, and train the model to predict the original token IDs at those positions. It is a fill-in-the-blank task over token IDs, not a prediction of arbitrary words independent of the tokenizer.

The representation must stay consistent end to end. Tokenizer vocabulary and special-token conventions, sequence length, padding mask, selected mask positions, and target labels must all agree. If the model predicts a token ID from a different vocabulary mapping than the one used to encode the input, the learning target is invalid.

Choose a Keras implementation path

Path Best for What it provides
Build a compact encoder from scratch Learning how token embeddings, attention, and masked-token prediction fit together A small BERT-like model and a custom MLM training example
Use KerasHub BertMaskedLM Using a BERT preset and its supported preprocessing workflow A task API with preset loading and either raw-text preprocessing or explicit preprocessed features

Use the from-scratch route when transparency is the priority. Use the preset route when you want to start from an existing BERT configuration and weights instead of implementing the encoder mechanics yourself.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall

Path 1: build a compact BERT-like MLM model

The Keras example constructs a small encoder using TextVectorization and Keras attention layers, trains with an MLM objective on IMDB reviews, and later demonstrates downstream sentiment fine-tuning. The example page was created on 2020-09-18 and last modified on 2024-03-15; its setup notes include tf-nightly. Check the current example and your installed package compatibility rather than treating that setup line as a current version matrix. See Keras’s end-to-end masked language modeling example.

Understand what the sample configuration means

The tutorial’s illustrative configuration uses maximum sequence length 256, batch size 32, learning rate 0.001, vocabulary size 30,000, embedding dimension 128, eight attention heads, feed-forward dimension 128, and one encoder layer. These are values chosen for that compact tutorial, not BERT-base specifications or generally recommended production settings.

In a from-scratch implementation, the essential flow is to vectorize text, construct the encoder inputs, corrupt selected token positions, and compute the prediction loss for the original IDs at those positions. The example is useful for seeing how the pieces connect, but its compact dimensions and dataset do not establish the results, runtime, or scale of full BERT pretraining.

Path 2: use KerasHub’s BERT masked-LM task

The KerasHub API defines keras_hub.models.BertMaskedLM(backbone, preprocessor=None, **kwargs). Its backbone is a BertBackbone; a BertMaskedLMPreprocessor may be supplied for input preparation. The documented preset example loads bert_base_en_uncased, with preprocessing enabled by default when constructing from a preset. With preprocessing enabled, raw strings can be passed to fitting and evaluation, where they are tokenized and dynamically masked.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import keras_hub

masked_lm = keras_hub.models.BertMaskedLM.from_preset(
    "bert_base_en_uncased",
)
masked_lm.fit(x=text_features, batch_size=batch_size)

Here, text_features represents your text inputs and batch_size must be set for your run. Consult the installed KerasHub API and preset documentation for version-specific usage. The official model API is documented at KerasHub’s BertMaskedLM page.

Supply explicit preprocessed features when you need control

The API also supports explicit preprocessed inputs. Its example uses a feature mapping containing token_ids, padding_mask, mask_positions, and segment_ids; labels are the original token IDs at the selected masked positions. These components must describe the same tokenized sequence.

Do not assume the mask token ID is always zero. The API example uses zero for its illustrative input, but the correct ID depends on the tokenizer or preprocessor vocabulary and conventions. Use the matching preprocessing configuration to obtain or encode the right token and labels.

Masking rate and prediction positions are recipe choices

Masking percentages are source-specific settings, not a universal Keras constant. Google Research’s original BERT repository states, “We mask out 15% of the words in the input, run the entire sequence through a deep bidirectional Transformer encoder, and then predict only the masked words.” The repository’s description gives that original recipe; see the Google Research BERT repository.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A separate KerasHub pretraining guide uses MASK_RATE = 0.25, sequence length 128, and PREDICTIONS_PER_SEQ = 32 in its sample configuration. Those are guide values, not universal defaults. The same guide demonstrates WordPiece tokenization and MaskedLMMaskGenerator in a custom input pipeline, with masking mapped over tf.data so selected positions can be generated as batches are iterated. See Keras’s Transformer pretraining with KerasHub guide.

For data generation, Google’s repository recommends setting maximum predictions per sequence to around maximum sequence length multiplied by the masked-LM probability, and passing the value consistently to data generation and training. If you change sequence length or masking rate in a custom pipeline, ensure that prediction-position capacity and labels still match the generated examples.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the KerasHub pretraining pipeline does

  1. Tokenize: apply the tokenizer and vocabulary that match the BERT model, including its special-token conventions.
  2. Prepare sequences: create token IDs, padding masks, and segment IDs in the format expected by the model.
  3. Select masked positions: use MaskedLMMaskGenerator to produce corrupted inputs, positions to predict, and the original-token labels.
  4. Encode and predict: the model encodes token IDs; its MaskedLMHead gathers the encodings at selected positions and projects them to vocabulary predictions.
  5. Train: the guide’s example compiles with sparse categorical cross-entropy, AdamW, and weighted sparse categorical accuracy.

This separation helps diagnose data bugs: an incorrect vocabulary mapping or mismatched labels cannot be repaired by changing the optimizer. The general KerasHub masked-Language-Model task behavior is described in the MaskedLM API documentation.

MLM is not all of original BERT pretraining

Original BERT pretraining material describes both masked language modeling and next sentence prediction. KerasHub’s BertMaskedLM is documented as an MLM task. It should not be treated as automatically recreating every objective or data-preparation step in the original BERT workflow. If your goal is specifically to reproduce that broader workflow, MLM alone is not the complete objective.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Plan for compute without assuming a runtime

Transformer pretraining can be computationally intensive. Actual training cost depends on the model, dataset, sequence length, and hardware, so these examples do not support a generic runtime or minimum hardware promise. Start with a configuration your environment can process, then measure your own pipeline before scaling it up.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.