Calculus can initially feel like a collection of complicated formulas. But at the heart of many machine learning algorithms is one simple idea:
A derivative tells us how a function is changing at a particular point.
Once we understand derivatives, slopes, gradients, and optimization become much easier to connect.
1. What Is a Derivative?
Suppose we have a function:
A derivative tells us how much the output changes when the input changes by a very small amount.
In other words:
The derivative measures the instantaneous rate of change of a function.
We write the derivative as:
or:
A simple example
Consider:
At :
Now slightly increase to :
The change in is:
The change in is:
Therefore, the approximate rate of change is:
As the change in becomes smaller and smaller, this value approaches .
Therefore:
This means that at , the function is changing at a rate of .
2. Where Does the Derivative Formula Come From?
To understand the derivative, first think about the slope between two points.
Suppose we have two points on the curve:
and:
The slope between these two points is:
which simplifies to:
This is called the difference quotient.
But this only gives us the slope between two points.
What if we want the slope at exactly ?
We move point closer and closer to by making approach zero:
The slope between the two points then approaches the slope of the curve at exactly .
Therefore, the derivative is defined as:
This is the mathematical foundation of the derivative.
3. The Derivative Is the Slope of a Curve
For a straight line, the slope is constant.
For example:
has a slope of everywhere.
But a curve does not have the same slope everywhere.
Consider:
Its derivative is:
Therefore:
The slope changes depending on where we are on the curve.
Geometrically, the derivative gives us the slope of the tangent line at a particular point.
So we can think of a derivative in three equivalent ways:
- Rate of change
- Slope of the tangent line
- Instantaneous change
4. What Does the Sign of a Derivative Tell Us?
The derivative also tells us which direction a function is moving.
Positive derivative
If:
the function is increasing.
Negative derivative
If:
the function is decreasing.
Zero derivative
If:
the function is momentarily flat.
This last case becomes particularly important when we study optimization.
5. From Derivatives to Optimization
Optimization is about finding the best possible value.
In machine learning, we commonly want to minimize a loss function.
Suppose our loss is:
where represents a model parameter.
Our goal is:
Imagine the loss function as a landscape:
Loss
^
| /\
| / \
| / \
|_____/ \____
↑
minimum
We want to find the bottom of the valley.
This is where gradient descent comes in.
6. What Is Gradient Descent?
Gradient descent is an optimization algorithm that repeatedly moves toward a direction that reduces the function’s value.
For a single parameter , the update rule is:
where:
- is the parameter
- is the derivative of the loss
- is the learning rate
The derivative tells us the direction in which the loss is increasing.
Therefore, we move in the opposite direction to decrease the loss.
That is why we subtract the derivative.
7. Why Do We Subtract the Derivative?
Suppose:
The loss is increasing as increases.
So we want to move toward smaller values of .
The update:
moves to the left.
On the other hand, if:
the loss is decreasing as increases.
Subtracting a negative value causes to increase:
So we move to the right.
In both cases, we are moving downhill.
8. What Is the Learning Rate?
The learning rate, represented by , controls how large each step is.
If is too small:
The algorithm takes very small steps and learning can be extremely slow.
If is too large:
The algorithm may jump over the minimum and fail to converge.
A good learning rate allows us to gradually move toward the minimum.
9. Local Minimum
A local minimum is a point that is lower than the points immediately surrounding it.
Think of it as the bottom of a small valley:
\ /
\ /
\___/
↑
local minimum
For example:
has a minimum at:
because:
while:
and:
The derivative is:
Therefore:
At the bottom of the valley, the slope is flat.
10. Local Maximum
A local maximum is the opposite of a local minimum.
It is a point that is higher than the points immediately surrounding it.
Think of it as the top of a hill:
/\
/ \
/ \
/ \
↑
local maximum
For example:
has a maximum at:
Its derivative is:
Therefore:
Again, the slope is zero at the top.
11. Does Derivative = 0 Always Mean a Minimum?
No.
This is an important point.
A derivative of zero only tells us that the function is flat at that point.
Consider:
Its derivative is:
At :
But is neither a local minimum nor a local maximum.
The function simply becomes flat momentarily and continues increasing.
Therefore:
A zero derivative is a candidate for a minimum or maximum, but it does not guarantee one.
12. From Derivative to Gradient
So far, we have considered functions with one variable.
Machine-learning models usually have many parameters:
Our loss function might therefore look like:
We calculate the derivative with respect to each parameter.
These partial derivatives are collected into a vector called the gradient:
The gradient points in the direction of steepest increase of the function.
Therefore, the negative gradient points toward the direction of steepest decrease.
Gradient descent uses this idea:
13. The Big Picture
All these concepts are connected:
Tells us how a function is changing.
Tells us the direction and rate of change at a point.
Extends this idea to functions with multiple variables.
Points toward decreasing values.
Repeatedly moves in the direction that decreases the loss.
A point where the function is lower than its nearby points.
Final Mental Model
Imagine a ball sitting somewhere on a mountain landscape.
The gradient tells you which direction is uphill.
If your goal is to minimize the height, you go in the opposite direction.
You take a step:
Then you calculate the gradient again and take another step.
You repeat this process until you approach a minimum.
That simple idea — measure the slope, move downhill, and repeat — is one of the fundamental mathematical ideas behind how machine-learning models learn.
Share this article