When people say:
“I used machine learning to predict the price using linear regression.”
it can sound more complicated than it actually is.
There may be a lot of terminology around machine learning, models, training, parameters, and predictions. But if we strip away the terminology, simple linear regression is fundamentally a mathematical problem.
At its core, we are trying to find just two numbers:
- The intercept —
- The slope —
Once we have those two numbers, making a prediction is simply a matter of putting a value into an equation.
1. The basic idea
Suppose we want to predict someone’s salary based on their years of experience.
We have some historical data:
| Years of Experience | Salary |
|---|---|
| 1 | 30,000 |
| 2 | 35,000 |
| 3 | 42,000 |
| 4 | 48,000 |
| 5 | 55,000 |
We can plot these observations on a graph.
Scatter plot of salary versus years of experience with the fitted regression line

The goal of linear regression is to find a straight line that represents the relationship between experience and salary.
The equation of that line is:
Where:
- = predicted value
- = input value
- = intercept
- = slope
That’s essentially our entire prediction model.
2. What does the slope mean?
The slope tells us how much the predicted output changes when the input increases by one unit.
For example, suppose we calculate:
This means that for every additional year of experience, the predicted salary increases by approximately $5,000.
So if our model is:
then:
- 1 year → $30,000
- 2 years → $35,000
- 3 years → $40,000
- 4 years → $45,000
The slope controls how steep the line is.
3. What does the intercept mean?
The intercept is the predicted value of when .
Suppose:
Then our model says:
When experience is zero:
Therefore:
So the intercept is:
The point where the regression line crosses the -axis.
Depending on the problem, the intercept may or may not have a meaningful real-world interpretation. Mathematically, however, it is one of the two parameters needed to define the line.
4. So where do the slope and intercept come from?
This is where the interesting mathematics happens.
We don’t simply guess:
and
Instead, we calculate values that produce the best-fitting line for the observed data.
One common approach is called Ordinary Least Squares (OLS).
The idea behind OLS is simple:
Find the line that minimizes the total squared difference between the actual values and the predicted values.
For each observation, we have an error:
Since some errors are positive and others are negative, simply adding them together could cause them to cancel each other out.
Instead, we square each error.
The objective becomes:
Since:
we can write:
OLS finds the values of and that minimize this quantity.
5. Deriving and using partial derivatives
We have already defined the Sum of Squared Errors:
Our goal is to find the values of and that make this SSE as small as possible.
Because we have two unknowns, and , we take the partial derivative with respect to each one.
Step 1: Find
Start with:
We differentiate with respect to :
The summation can remain outside:
Now we use the chain rule. If:
then:
Since:
we get:
At the minimum SSE, the derivative must be zero:
Divide by :
Expand the summation:
Because and are constants:
Rearrange:
Divide by :
But:
and:
Therefore:
So this is where the intercept formula comes from.
Step 2: Find
Now we differentiate the same SSE with respect to :
Again, using the chain rule, the derivative of the inside is:
Therefore:
At the minimum:
Divide by :
Expand:
Therefore:
Now we already know that:
Substitute this into the equation:
Since:
we get:
Expand:
Move to the left:
Therefore:
This can be rewritten using deviations from the mean:
This is the familiar OLS slope formula.
Setting the derivatives to zero only finds a stationary point, but here it is the minimum: SSE is a convex function of and (a sum of squared linear terms), so its only stationary point is the global minimum.
6. The whole idea in one flow
So the entire process is:
Specifically:
Then:
Differentiate with respect to :
which gives:
Differentiate with respect to :
which gives:
Therefore, linear regression is not simply guessing a line.
It uses calculus to find the values of and that produce the minimum possible squared error for the training data.
5. Calculating the slope
For simple linear regression, the optimal slope is:
At first, this formula can look intimidating. Let’s break it down.
This is the slope of our regression line.
This represents an individual observed input value.
For example, if our input is years of experience:
could represent an employee with 3 years of experience.
This represents the corresponding observed output.
For example:
could be that employee’s actual salary.
and
— This is the mean of all the input values:
— This is the mean of all the output values:
Numerator
The numerator is:
It measures how and vary together around their respective means.
- If larger values of generally occur with larger values of , this quantity tends to be positive.
- If larger values of generally occur with smaller values of , it tends to be negative.
Denominator
The denominator is:
This measures how much the input varies around its mean.
So, conceptually, the slope can be thought of as:
Calculating the intercept
Once we have calculated the slope, the intercept is much simpler. The formula is:
Again, let’s understand the terms:
- = intercept
- = mean of the target values
- = calculated slope
- = mean of the input values
So the process is essentially:
Once we have and , our model is ready.
7. Now prediction becomes extremely simple
Suppose our calculations give us:
and:
Our model becomes:
Now suppose someone has 6 years of experience.
We simply substitute:
Therefore:
That’s the prediction.
There is no mysterious “AI thinking” happening here.
The model has learned two numbers:
and
Then it applies the equation.
8. So what does model.fit(X, y) actually do?
This is where the machine-learning terminology can make things sound more complicated than they are.
In Python, we might write:
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X, y)
When we call:
model.fit(X, y)
we are essentially asking the algorithm:
“Look at these examples and find the parameters that give the best-fitting linear relationship.”
For simple linear regression, those parameters are:
and
After fitting the model, we can inspect them:
print(model.intercept_)
print(model.coef_)
Conceptually:
model.fit(X, y)
↓
Find β₀ and β₁
↓
ŷ = β₀ + β₁x
↓
Use the equation for prediction
Scikit-learn uses numerical linear-algebra methods internally rather than necessarily evaluating the formulas above literally step by step. But mathematically, for ordinary simple linear regression with an intercept, it is solving the same least-squares problem.
9. Why do we call this “Machine Learning”?
This is perhaps the most interesting part.
The mathematics itself is not new.
Linear regression has a long history in statistics and mathematics.
What makes it part of a machine-learning workflow is the idea that:
- We provide data.
- An algorithm estimates parameters from that data.
- The resulting model is used to make predictions on new data.
For linear regression, the learned parameters happen to be:
For more complicated machine-learning models, there may be thousands, millions, or even billions of parameters.
But the basic idea remains:
Use data to estimate parameters, then use those parameters to produce predictions.
10. The important mental model
When you hear:
“I trained a linear regression model.”
don’t imagine a mysterious black box.
For simple linear regression, imagine this:
Training Data
│
▼
Find the best β₀, β₁
│
▼
ŷ = β₀ + β₁x
│
▼
New input x
│
▼
Prediction ŷ
The “learning” part is essentially finding the parameters.
The “prediction” part is simply using those parameters in the equation.
11. The bigger picture
Linear regression is a beautiful example because it lets us see what is happening underneath the machine-learning terminology.
At the surface, we might write:
model.fit(X, y)
and:
model.predict(X_new)
But underneath those convenient APIs is mathematics.
The fundamental process is:
For simple linear regression:
and then:
That’s really what is happening behind the scenes.
Final takeaway
The next time someone says:
“I used machine learning to predict something using linear regression.”
remember what is actually happening.
There is a dataset.
We find the line that best fits the data using a mathematical optimization method such as Ordinary Least Squares.
That gives us two important parameters:
and:
Then we put them into:
And that’s the prediction.
The “machine learning” label doesn’t make the mathematics disappear. In this case, the machine is essentially helping us estimate the parameters from data. The prediction itself is just the equation.
Share this article