You can build a working handwritten-digit classifier with Python and NumPy by implementing four operations yourself: weighted sums, an activation function, a loss calculation, and backpropagation updates. “From scratch” here means you write the model and training logic instead of calling a ready-made neural-network estimator; NumPy still provides arrays and matrix multiplication.
The example below uses a feedforward network with one hidden layer. It is intentionally small so every value and derivative can be inspected.
What you will build
The network accepts one 28×28 grayscale image, flattens it to 784 input values, and produces 10 scores corresponding to digits 0 through 9. The MNIST framing used by the NumPy tutorial contains 60,000 training images and 10,000 test images. Those figures describe the dataset split, not a performance guarantee for this implementation.
- Input: 784 normalized pixel values.
- Hidden layer: 64 units with ReLU activation.
- Output: 10 scores.
- Parameters: two weight matrices, with biases omitted for clarity.
- Training: squared error and single-example gradient descent.
Biases, minibatches, improved initialization, and a classification-specific softmax cross-entropy loss are useful extensions after the basic mechanics are clear.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →#1 Best Overall
Prerequisites and data representation
Use Python and NumPy. A NumPy arrays and linear-algebra refresher is useful if shapes such as (784, 64) and matrix multiplication are unfamiliar. Matplotlib is useful for displaying images but is not required for the network’s computations.
Before training, arrange the data as:
X_train: an array shaped(60000, 784), with pixel values scaled to a small range such as 0 to 1.y_train: one-hot targets shaped(60000, 10).X_testandy_test: the corresponding held-out examples.
If labels are integers from 0 to 9, convert each label k to a row whose kth element is 1 and all other elements are 0. Keep the test set separate while fitting weights; training accuracy does not measure performance on unseen data.
Represent the network with NumPy
For one input row x, the first matrix product must be valid: (784,) @ (784, 64) produces 64 hidden pre-activation values. The second product is (64,) @ (64, 10), producing 10 output scores.
import numpy as np
rng = np.random.default_rng(7)
input_size = 784
hidden_size = 64
output_size = 10
# Small random weights make the initial activations manageable.
W1 = rng.normal(0.0, 0.05, size=(input_size, hidden_size))
W2 = rng.normal(0.0, 0.05, size=(hidden_size, output_size))
The fixed seed makes this particular initialization reproducible. Omitting biases is a tutorial simplification: a practical network normally gives each unit a trainable offset as well as weights.
Free tools Windows power users keep installed
One-click scans. No signup required.
Forward propagation
A forward pass carries the input through the network. The hidden layer first computes a weighted sum, then ReLU replaces negative values with zero. The nonlinearity is what lets the network represent relationships that a single linear transformation cannot.
def relu(z):
return np.maximum(0, z)
def relu_derivative(z):
return (z > 0).astype(float)
def forward(x, W1, W2):
z1 = x @ W1
a1 = relu(z1)
z2 = a1 @ W2
return z1, a1, z2
Here z1 is the hidden layer’s value before activation, a1 is its output, and z2 is the final vector of ten scores. Save both pre-activation and activated values because backpropagation needs them.
Measure the error
To mirror the simple teaching example, use half the total squared error:
def loss(prediction, target):
return 0.5 * np.sum((prediction - target) ** 2)
This is a clear pedagogical loss, not the only or usual choice for multiclass classification. Softmax with cross-entropy is generally a better next step because it models class probabilities and gives a more suitable gradient for mutually exclusive labels.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Rank #3
Backpropagation with the chain rule
Backpropagation computes how much each parameter contributed to the loss. Start at the output and apply the chain rule one operation at a time.
- The derivative of the squared error with respect to the output scores is
z2 - target. - Each element of
W2is affected by a hidden activation and an output error, so its gradient is an outer product. - Propagate the output error through
W2to reach the hidden layer. - Multiply by the ReLU derivative. A unit whose pre-activation was non-positive receives zero gradient in this pass.
- Use another outer product to obtain the gradient for
W1.
def gradients(x, target, W1, W2):
z1, a1, prediction = forward(x, W1, W2)
# dL/dz2 for 0.5 * sum((z2 - target)^2)
error_out = prediction - target
# z2 = a1 @ W2
grad_W2 = np.outer(a1, error_out)
# Propagate through W2, then through ReLU
error_hidden = (error_out @ W2.T) * relu_derivative(z1)
# z1 = x @ W1
grad_W1 = np.outer(x, error_hidden)
return prediction, grad_W1, grad_W2
The forward values and their derivatives are different things: a1 is a value used to compute the output, while relu_derivative(z1) is a local slope used to route the error backward.
Update weights with gradient descent
Gradient descent moves each parameter in the opposite direction of its gradient. With learning rate η, the update is conceptually parameter = parameter - η × gradient.
def train_one_epoch(X, Y, W1, W2, learning_rate=0.01):
total_loss = 0.0
order = rng.permutation(len(X))
for i in order:
x = X[i]
target = Y[i]
prediction, grad_W1, grad_W2 = gradients(x, target, W1, W2)
total_loss += loss(prediction, target)
W1 -= learning_rate * grad_W1
W2 -= learning_rate * grad_W2
return total_loss / len(X), W1, W2
This is stochastic gradient descent: one example produces one update. It is easy to understand but slower and noisier than minibatch training. Run several epochs and record the loss:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #4
for epoch in range(10):
average_loss, W1, W2 = train_one_epoch(
X_train, y_train, W1, W2, learning_rate=0.01
)
print(f"epoch {epoch + 1}: loss={average_loss:.4f}")
A falling loss indicates that the current examples are producing smaller errors, but it does not by itself establish good generalization.
Evaluate on held-out digits
For this score-based output, choose the index of the largest score.
def predict(X, W1, W2):
predictions = []
for x in X:
_, _, scores = forward(x, W1, W2)
predictions.append(np.argmax(scores))
return np.array(predictions)
predicted = predict(X_test, W1, W2)
actual = np.argmax(y_test, axis=1)
test_accuracy = np.mean(predicted == actual)
print(f"test accuracy: {test_accuracy:.4f}")
Evaluate only after training with X_test and its labels. Do not report an expected accuracy for this listing without running the exact preprocessing, initialization, learning rate, epoch count, and dataset version you use.
Debugging checklist
- Shape errors: print
X_train.shape,y_train.shape,W1.shape, andW2.shape. The expected dimensions are 784 inputs, 64 hidden units, and 10 outputs. - Loss is not changing: verify that pixels are numeric and scaled, the learning rate is nonzero, and updates use the gradients returned for the same example.
- Loss becomes enormous: lower the learning rate or reduce the initial weight scale.
- All predictions are one class: inspect score ranges and confirm labels are correctly one-hot encoded.
- Zero hidden gradients: ReLU passes no gradient where its input is non-positive. This can leave units inactive; alternative activations or better initialization can help.
- Training looks good but test performance is poor: the model may be overfitting, or preprocessing may differ between the two splits.
What this minimal implementation leaves out
The example writes the essential computations explicitly, but it is not a production training system. A fuller implementation would add bias vectors, minibatches, a more robust initialization scheme, vectorized batch gradients, softmax cross-entropy, validation tracking, and an optimizer with features such as momentum. Larger networks also need attention to numerical stability, memory use, and reproducibility.
Best Value
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
Backpropagation is the standard gradient-based mechanism for multilayer networks, yet gradients can vanish or explode as networks deepen. ReLU avoids some saturation problems but has failure cases when units remain on its inactive side.
NumPy versus framework tutorials
Learning paths differ in what they automate. The NumPy example above manually builds a one-hidden-layer classifier. PyTorch’s minimal tensor exercise is logistic regression with no hidden layer, while its neural-network tutorial uses the framework’s neural-network abstractions and training workflow. They are useful for different purposes and are not controlled speed or accuracy comparisons.
| Approach | What you write | Architecture or focus |
|---|---|---|
| NumPy implementation | Forward pass, loss, derivatives, and parameter updates | One-hidden-layer digit classifier |
| PyTorch manual tensor example | Tensor calculations while the framework handles tensor infrastructure | Logistic regression without a hidden layer |
| PyTorch neural-network tutorial | Model and training workflow using framework abstractions and optimizers | Broader neural-network workflow |
An optional deeper read is Neural Networks from Scratch in Python by Harrison Kinsley and Daniel Kukieła. It covers derivatives, gradients, gradient descent, and backpropagation; verify the edition and availability before purchasing.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minute

