Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
You can build a GPT-2-small-scale, decoder-only Transformer in PyTorch with a few hundred lines of model code, and the result is a real language model: it takes token IDs, predicts the next token, and can be trained with ordinary cross-entropy loss. What takes longer is the data pipeline and training run. The “124M” label is also less exact than it looks. It depends on how you count parameters, and the original GPT-2 paper and the nanoGPT repository use different numbers for the same architecture.
This guide walks through the configuration, the forward pass, the next-token objective, and the difference between a small learning run and a full reproduction. It is written for readers who want to understand each component well enough to write or modify the code themselves.
What “124M” actually counts
The GPT-2 paper’s architecture table lists its smallest model at 117M parameters, with 12 layers and 768-dimensional states. The nanoGPT repository uses the same architecture and labels it GPT-2 (124M). The two labels describe the same shape of network but disagree on the number, and the sources do not reconcile the 117M figure with the nanoGPT count. Treat 117M as the paper’s own label and 124M as the label used by the implementation most readers will copy.
You can check the 124M figure yourself. Counting the trainable tensors in the configuration below, with the output projection tied to the token embedding and biases included, gives about 124.47 million. This arithmetic is derived from the configuration values, not from a run of code.
#1 Best Overall
| Component | Shape or setting | Parameters |
|---|---|---|
| Token embedding (wte) | 50,257 × 768 | 38,597,376 |
| Position embedding (wpe) | 1,024 × 768 | 786,432 |
| One block: attention QKV projection | 768 → 2,304, with bias | 1,774,080 |
| One block: attention output projection | 768 → 768, with bias | 590,592 |
| One block: two LayerNorms | 2 × (768 weight + 768 bias) | 3,072 |
| One block: MLP up-projection | 768 → 3,072, with bias | 2,362,368 |
| One block: MLP down-projection | 3,072 → 768, with bias | 2,360,064 |
| Twelve blocks | 12 × 7,090,176 | 85,082,112 |
| Final LayerNorm | 768 weight + 768 bias | 1,536 |
| Language-model head | Weight tied to token embedding | 0 additional |
| Total | 124,467,456 |
Two counting choices move the total most. If you untie the language-model head from the token embedding, you add another 38,597,376 parameters and the model becomes about 163.1 million. If you omit biases or use a different position-embedding size, the number changes again. When you publish a count, state the convention: tied or untied head, whether biases and LayerNorm parameters are included, and the context length used for position embeddings.
The reference configuration
The GPT-2-small reference values, as set in the nanoGPT configuration, are listed below. Each attention head works on 768 ÷ 12 = 64 channels, so the embedding width must divide evenly across heads.
| Setting | Value | Meaning |
|---|---|---|
n_layer |
12 | Number of stacked decoder blocks |
n_head |
12 | Attention heads per block |
n_embd |
768 | Hidden width of every residual stream |
| Head dimension | 64 | 768 ÷ 12 |
| Feed-forward inner width | 3,072 | Four times the hidden width, per the minGPT reference description of GPT-2 |
vocab_size |
50,257 | GPT-2 byte-pair-encoding vocabulary |
block_size |
1,024 | Maximum context length in tokens |
How a batch flows through the model
Use batch-first notation throughout. Let B be the batch size and T the sequence length, with T no greater than 1,024.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall1. Token and position embeddings
The input is an integer tensor of shape (B, T). The token embedding looks up one 768-dimensional vector per token ID, and a learned position embedding adds one 768-dimensional vector per position. Their sum has shape (B, T, 768). Position information is learned here rather than computed from a fixed sinusoidal formula, which is the GPT-2 style.
2. Masked (causal) self-attention
Each block first normalizes the hidden state, then projects it to queries, keys, and values. The projected tensors are split into 12 heads of 64 channels, so each has shape (B, 12, T, 64). Attention scores are computed for every pair of positions, but a causal mask sets scores for later positions to negative infinity before the softmax. Position t can therefore only mix information from positions 0 through t.
Rank #2
The mask does not hide future tokens from the training labels. It controls what each position may use when it forms its prediction. The labels are always the shifted sequence, described below.
In PyTorch 2.x you can compute the masked attention core with a single call:
Recommended Free Tools
qkv = self.c_attn(x) # (B, T, 3*768)
q, k, v = qkv.split(768, dim=2)
q = q.view(B, T, 12, 64).transpose(1, 2) # (B, 12, T, 64)
k = k.view(B, T, 12, 64).transpose(1, 2)
v = v.view(B, T, 12, 64).transpose(1, 2)
y = F.scaled_dot_product_attention(q, k, v, is_causal=True)
y = y.transpose(1, 2).contiguous().view(B, T, 768) # merge heads
y = self.c_proj(y) # output projection, (B, T, 768)
The is_causal=True argument applies the lower-triangular mask. It assumes queries and keys have the same length, which holds during training on full blocks.
3. Feed-forward network
The MLP expands each position from 768 to 3,072 dimensions, applies a GELU activation, and projects back to 768. It acts on each position independently; attention is the only step where positions exchange information.
4. Residual connections and layer normalization
GPT-2 uses pre-normalization. Each sub-block normalizes its input before computing attention or the MLP, and the result is added back to the residual stream:
Rank #3
x = x + self.attn(self.ln_1(x))
x = x + self.mlp(self.ln_2(x))
The GPT-2 paper describes moving layer normalization to the input of each sub-block and adding a final layer normalization after the last block. The residual additions keep the (B, T, 768) shape unchanged through the stack, which is what lets the blocks be stacked 12 deep without the input signal being overwritten at each step.
5. Language-model head and logits
After the final LayerNorm, a linear layer maps each 768-dimensional state to 50,257 scores, one per vocabulary entry. The output has shape (B, T, 50257). These are logits, not probabilities. In the tied configuration the head reuses the token-embedding matrix, which is the reason for the 124M count above.
The training objective: predict the next token
The model is trained to predict the next token from the preceding context. Take a token sequence x of length T+1 from your corpus. The inputs are x[:, :-1], the first T tokens, and the targets are x[:, 1:], the same sequence shifted left by one. The logits at input position t are compared with target token t+1.
The loss is cross-entropy over the vocabulary, averaged over all positions:
logits = model(inputs) # (B, T, 50257)
loss = F.cross_entropy(logits.view(-1, 50257), targets.reshape(-1))
Where the shift happens varies between implementations. Some data loaders return a window of T+1 tokens and split it into inputs and targets; others shift inside the training loop. Either works, but the alignment must be checked once. A common mistake is comparing the logits at position t with the token at position t instead of t+1, which produces a model that appears to learn quickly but cannot generate coherent text.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsRank #4
Sanity checks before a long run
- Initial loss. With near-uniform logits at initialization, cross-entropy should start close to ln(50,257) ≈ 10.82. A much higher starting loss usually points to a scaling problem in the initialization or the head.
- Single-batch overfit. Train on one small batch repeatedly. The loss should fall toward zero. If it does not, the bug is in the model, the loss, or the target alignment, not in the data volume.
- Causality test. Change a token at position t and confirm the logits at positions before t do not change. Any change before t indicates the mask is missing or misapplied.
- Shape checks. Confirm (B, T, 768) after the embeddings and blocks, and (B, T, 50257) at the head.
Preparing text data
Tokenize text with the GPT-2 byte-pair-encoding tokenizer so that token IDs match the 50,257-entry vocabulary. Cut the token stream into fixed windows no longer than 1,024 tokens. Decide how document boundaries are marked, whether short final windows are padded or dropped, and how the train and validation splits are made. Keep those choices explicit in the code, because they change the loss numbers you will compare later.
The nanoGPT README describes preprocessing OpenWebText into GPT-2 token IDs stored as raw uint16 bytes. That works here because the vocabulary is smaller than 65,536. The build-nanoGPT tutorial notes a PyTorch conversion problem with uint16 in its version and a workaround that converts through NumPy int32. Treat that as a note about that repository’s code, and check the current versions of PyTorch and NumPy before you copy it.
Three goals, three different projects
Readers usually fall into one of three groups, and each needs a different plan.
| Goal | What to run | Hardware and time | Claim you can make |
|---|---|---|---|
| Learn the architecture | A small model or reduced configuration on a tiny corpus, with the single-batch and causality checks above | Not stated by the sources; size batches and sequence length to your memory and measure | “Implements a GPT-2-style decoder-only Transformer” |
| Adapt an existing model | Your own data and training loop, starting from the 124M configuration | Depends on data size and sequence length; the sources give no minimum | “Trains a 124M-configuration model on my dataset,” with the parameter-counting convention stated |
| Reproduce the reference run | nanoGPT’s documented OpenWebText recipe | 8 × A100 40GB for about four days, per the nanoGPT README | “Follows the cited nanoGPT reproduction setup.” It does not recreate GPT-2 exactly. |
A tutorial build can be finished on modest hardware, and it is the right way to verify that your forward and backward passes are correct. A reproduction is a separate, compute-intensive project. Keep the two apart in your planning so that a slow run on one GPU is not mistaken for a failed implementation.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →What the reproduction numbers mean
The nanoGPT README reports a loss of about 2.85 for its OpenWebText reproduction on the 8 × A100 40GB setup. It compares this with a validation loss of about 3.11 that it attributes to GPT-2 evaluated on OpenWebText. The README states that the original GPT-2 was trained on WebText, while the reproduction uses OpenWebText, a best-effort reproduction of that dataset. It notes a domain gap that affects direct loss comparisons, so the difference between 2.85 and 3.11 should not be read as a measured gain over GPT-2.
These are the repository’s reported numbers for its own setup, not current benchmarks. A different dataset, tokenization pipeline, batch size, or number of training steps will produce a different loss.
Sampling from the model
Generation runs one token at a time. Feed the current context, take the logits at the last position only, convert them to probabilities with a softmax (optionally after temperature scaling or top-k filtering), sample one token, append it, and repeat. Stop when you reach your end token or the 1,024-token limit. Sample from a model that has trained for a short time and the output will be repetitive or incoherent; that is expected for a small run and does not by itself indicate a bug.
Choosing reference code and checking its status
Two repositories are the usual reading list for this architecture, and both need a status check before you rely on their commands.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →- nanoGPT. Its README carries a November 2025 update describing the project as old and deprecated and pointing readers to nanochat. It remains a compact, readable implementation of the GPT-2 configuration discussed here, and it documents the reproduction recipe above. Confirm that its commands run with your PyTorch version before you depend on them.
- minGPT. Its README carries a January 2023 note calling the project semi-archived. It is useful for seeing how the model, dataset, and trainer are separated into different modules, which many learners find clearer than a single training script.
Neither repository should be presented as the current standard for training code. Read them for structure and then check the active documentation for compatible dependencies.
Quick Recap
body_placeholder_removed
Next steps for a first build
- Write the model with the configuration table above, and print the parameter count using the convention you chose.
- Run the causality test and the shape checks on random token IDs.
- Tokenize a small text file, form shifted input and target windows, and confirm the first loss is near 10.8.
- Overfit a single batch until the loss approaches zero.
- Train on your real corpus with a fixed validation split, then sample to check the output.
- Only then consider a multi-GPU reproduction, and budget for the hardware and time described above.
“
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

