ReLU, short for rectified linear unit, is an activation function that returns the larger of zero and its input: f(x) = max(0, x). It turns negative values into zero and lets nonnegative values pass through. Neural networks use it to add nonlinearity between computations, while its simple shape gives active units a gradient of 1.
How the ReLU activation function works
For a scalar input x, ReLU is defined as:
f(x) = max(0, x)
| Input | Output | Behavior |
|---|---|---|
x < 0 |
0 |
Negative input is set to zero. |
x > 0 |
x |
Positive input passes through unchanged. |
x = 0 |
0 |
The function meets at a kink. |
In a neural network, ReLU commonly follows an affine transformation such as Wx + b. The transformation combines inputs and weights; ReLU then changes the result according to the rule above. As Google for Developers explains, “The rectified linear unit activation function (or ReLU, for short) transforms output using the following algorithm:” (Google Machine Learning Crash Course: Activation functions).
Without an activation function, stacking linear transformations still produces a linear transformation. Nonlinear activations such as ReLU let a network represent relationships that a purely linear stack cannot. The Deep Learning textbook’s neural-network chapter discusses ReLU and related activation functions.
What is the derivative of ReLU?
On either side of zero, the derivative is straightforward:
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
- For
x < 0, the slope is 0. - For
x > 0, the slope is 1.
At exactly zero, ReLU is not classically differentiable because the left and right slopes differ. A machine-learning framework may choose a convention for the backward pass at that point; that is an implementation choice, not a unique ordinary derivative.
Why ReLU is widely used—and what it does not fix
ReLU is computationally simple: it amounts to comparing a value with zero and selecting the larger one. Its derivative is also constant at 1 wherever the unit is active, meaning its input is positive. Compared with sigmoid or tanh, this can make ReLU less susceptible to vanishing gradients in that active region.
Rank #2
That advantage is conditional. ReLU does not prevent every gradient problem: inactive units have zero gradient, and networks can still experience issues such as exploding gradients. Its simple computation and useful active-side gradient help explain its appeal, but they do not guarantee that every model will train easily.
What is the dying ReLU problem?
A ReLU unit is inactive whenever its input is negative: it outputs zero and has a zero derivative there. If its weighted input remains below zero for the examples it sees, the unit can keep outputting zero and stop passing gradient through itself. This is called a dead or dying ReLU.
Google’s neural-network training guide describes this failure mode and notes that lowering the learning rate may help. It is a possible remedy, not a guarantee; LeakyReLU is another option because it preserves a nonzero slope on the negative side.
How LeakyReLU and PReLU differ from ReLU
ReLU’s negative-side slope is zero. LeakyReLU and PReLU instead allow a negative input to produce a small, nonzero output. This gives gradients a path through units that would otherwise be inactive. PReLU makes the negative-side slope learnable, while LeakyReLU uses a fixed slope.
Rank #4
| Activation | Negative input | Negative-side slope | What changes |
|---|---|---|---|
| ReLU | Output is zero | 0 | Simple, but an inactive unit passes no gradient. |
| LeakyReLU | Output follows a small negative-side slope | Fixed | Retains a gradient path for negative inputs. |
| PReLU | Output follows a negative-side slope | Learned | The model learns the slope rather than keeping it fixed. |
Neither variant is automatically better. The useful choice depends on whether inactive units are a concern, how the model performs on the particular task, and the implementation’s added cost. Compare them in the setting where the network will be used rather than assuming one activation always wins.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What the historical ImageNet comparison does—and does not—show
In their 2015 paper, Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun reported 4.94% top-5 test error on ImageNet 2012 for their PReLU networks. The paper compared this with a cited 6.66% result for GoogLeNet, the ILSVRC 2014 winner, and described the difference as a 26% relative improvement. It also cited 5.1% as human-level performance in that benchmark context (He et al., 2015).
Best Value
These are historical figures from a specific ImageNet comparison, not current general-purpose accuracy expectations. They are results for PReLU networks and do not isolate standard ReLU as the cause of a performance difference.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

