What Really Happens Behind the Scenes in Linear Regression?

What Really Happens Behind the Scenes in Linear Regression?

Share
Machine LearningRegression
10 min read

When people say:

“I used machine learning to predict the price using linear regression.”

it can sound more complicated than it actually is.

There may be a lot of terminology around machine learning, models, training, parameters, and predictions. But if we strip away the terminology, simple linear regression is fundamentally a mathematical problem.

At its core, we are trying to find just two numbers:

  • The intercept — β0\beta_0
  • The slope — β1\beta_1

Once we have those two numbers, making a prediction is simply a matter of putting a value into an equation.


1. The basic idea

Suppose we want to predict someone’s salary based on their years of experience.

We have some historical data:

Years of ExperienceSalary
130,000
235,000
342,000
448,000
555,000

We can plot these observations on a graph.

Scatter plot of salary versus years of experience with the fitted regression line

image

The goal of linear regression is to find a straight line that represents the relationship between experience and salary.

The equation of that line is:

y^=β0+β1x\hat{y} = \beta_0 + \beta_1 x

Where:

  • y^\hat{y} = predicted value
  • xx = input value
  • β0\beta_0 = intercept
  • β1\beta_1 = slope

That’s essentially our entire prediction model.


2. What does the slope mean?

The slope tells us how much the predicted output changes when the input increases by one unit.

For example, suppose we calculate:

β1=5,000\beta_1 = 5{,}000

This means that for every additional year of experience, the predicted salary increases by approximately $5,000.

So if our model is:

y^=25,000+5,000x\hat{y} = 25{,}000 + 5{,}000x

then:

  • 1 year → $30,000
  • 2 years → $35,000
  • 3 years → $40,000
  • 4 years → $45,000

The slope controls how steep the line is.


3. What does the intercept mean?

The intercept is the predicted value of yy when x=0x=0.

Suppose:

β0=25,000\beta_0 = 25{,}000

Then our model says:

y^=25,000+5,000x\hat{y} = 25{,}000 + 5{,}000x

When experience is zero:

y^=25,000+5,000(0)\hat{y} = 25{,}000 + 5{,}000(0)

Therefore:

y^=25,000\hat{y} = 25{,}000

So the intercept is:

The point where the regression line crosses the yy-axis.

Depending on the problem, the intercept may or may not have a meaningful real-world interpretation. Mathematically, however, it is one of the two parameters needed to define the line.


4. So where do the slope and intercept come from?

This is where the interesting mathematics happens.

We don’t simply guess:

β0=25,000\beta_0 = 25{,}000

and

β1=5,000\beta_1 = 5{,}000

Instead, we calculate values that produce the best-fitting line for the observed data.

One common approach is called Ordinary Least Squares (OLS).

The idea behind OLS is simple:

Find the line that minimizes the total squared difference between the actual values and the predicted values.

For each observation, we have an error:

ei=yi−yi^e_i = y_i - \hat{y_i}

Since some errors are positive and others are negative, simply adding them together could cause them to cancel each other out.

Instead, we square each error.

The objective becomes:

SSE=∑i=1n(yi−y^i)2\text{SSE} = \sum_{i=1}^{n} (y_i - \hat{y}_i)^2

Since:

yi^=β0+β1xi\hat{y_i} = \beta_0+\beta_1x_i

we can write:

SSE=∑i=1n(yi−β0−β1xi)2\text{SSE} = \sum_{i=1}^{n} \left( y_i - \beta_0 - \beta_1 x_i \right)^2

OLS finds the values of β0\beta_0 and β1\beta_1 that minimize this quantity.


5. Deriving β0\beta_0 and β1\beta_1 using partial derivatives

We have already defined the Sum of Squared Errors:

SSE=∑i=1n(yi−β0−β1xi)2\text{SSE} = \sum_{i=1}^{n} \left(y_i - \beta_0 - \beta_1 x_i\right)^2

Our goal is to find the values of β0\beta_0 and β1\beta_1 that make this SSE as small as possible.

Because we have two unknowns, β0\beta_0 and β1\beta_1, we take the partial derivative with respect to each one.


Step 1: Find β0\beta_0

Start with:

SSE=∑i=1n(yi−β0−β1xi)2\text{SSE} = \sum_{i=1}^{n} \left(y_i - \beta_0 - \beta_1 x_i\right)^2

We differentiate with respect to β0\beta_0:

∂SSE∂β0=∂∂β0∑i=1n(yi−β0−β1xi)2\frac{\partial \text{SSE}}{\partial \beta_0} = \frac{\partial}{\partial \beta_0} \sum_{i=1}^{n} \left(y_i - \beta_0 - \beta_1 x_i\right)^2

The summation can remain outside:

=∑i=1n∂∂β0(yi−β0−β1xi)2= \sum_{i=1}^{n} \frac{\partial}{\partial \beta_0} \left(y_i - \beta_0 - \beta_1 x_i\right)^2

Now we use the chain rule. If:

u=yi−β0−β1xiu = y_i - \beta_0 - \beta_1 x_i

then:

∂∂β0(u2)=2u ∂u∂β0\frac{\partial}{\partial \beta_0}\left(u^2\right) = 2u \, \frac{\partial u}{\partial \beta_0}

Since:

∂∂β0(yi−β0−β1xi)=−1\frac{\partial}{\partial \beta_0} \left(y_i - \beta_0 - \beta_1 x_i\right) = -1

we get:

∂SSE∂β0=−2∑i=1n(yi−β0−β1xi)\frac{\partial \text{SSE}}{\partial \beta_0} = -2 \sum_{i=1}^{n} \left(y_i - \beta_0 - \beta_1 x_i\right)

At the minimum SSE, the derivative must be zero: 2∑i=1n(yi−β0−β1xi)=02 \sum_{i=1}^{n} \left(y_i - \beta_0 - \beta_1 x_i\right) = 0

Divide by −2-2:

∑i=1n(yi−β0−β1xi)=0\sum_{i=1}^{n} \left(y_i - \beta_0 - \beta_1 x_i\right) = 0

Expand the summation:

∑yi−∑β0−∑β1xi=0\sum y_i - \sum \beta_0 - \sum \beta_1 x_i = 0

Because β0\beta_0 and β1\beta_1 are constants:

∑yi−nβ0−β1∑xi=0\sum y_i - n\beta_0 - \beta_1 \sum x_i = 0

Rearrange:

nβ0=∑yi−β1∑xin\beta_0 = \sum y_i - \beta_1 \sum x_i

Divide by nn:

β0=∑yin−β1∑xin\beta_0 = \frac{\sum y_i}{n} - \beta_1 \frac{\sum x_i}{n}

But:

yˉ=∑yin\bar{y} = \frac{\sum y_i}{n}

and:

xˉ=∑xin\bar{x} = \frac{\sum x_i}{n}

Therefore:

β0=yˉ−β1xˉ\boxed{\beta_0 = \bar{y} - \beta_1 \bar{x}}

So this is where the intercept formula comes from.


Step 2: Find β1\beta_1

Now we differentiate the same SSE with respect to β1\beta_1:

∂SSE∂β1=∂∂β1∑i=1n(yi−β0−β1xi)2\frac{\partial \text{SSE}}{\partial \beta_1} = \frac{\partial}{\partial \beta_1} \sum_{i=1}^{n} \left(y_i - \beta_0 - \beta_1 x_i\right)^2

Again, using the chain rule, the derivative of the inside is:

∂∂β1(yi−β0−β1xi)=−xi\frac{\partial}{\partial \beta_1} \left(y_i - \beta_0 - \beta_1 x_i\right) = -x_i

Therefore:

∂SSE∂β1=−2∑i=1nxi(yi−β0−β1xi)\frac{\partial \text{SSE}}{\partial \beta_1} = -2 \sum_{i=1}^{n} x_i \left(y_i - \beta_0 - \beta_1 x_i\right)

At the minimum: 2∑i=1nxi(yi−β0−β1xi)=02 \sum_{i=1}^{n} x_i \left(y_i - \beta_0 - \beta_1 x_i\right) = 0

Divide by −2-2:

∑i=1nxi(yi−β0−β1xi)=0\sum_{i=1}^{n} x_i \left(y_i - \beta_0 - \beta_1 x_i\right) = 0

Expand:

∑xiyi−β0∑xi−β1∑xi2=0\sum x_i y_i - \beta_0 \sum x_i - \beta_1 \sum x_i^2 = 0

Therefore:

∑xiyi=β0∑xi+β1∑xi2\sum x_i y_i = \beta_0 \sum x_i + \beta_1 \sum x_i^2

Now we already know that:

β0=yˉ−β1xˉ\beta_0 = \bar{y} - \beta_1 \bar{x}

Substitute this into the equation:

∑xiyi=(yˉ−β1xˉ)∑xi+β1∑xi2\sum x_i y_i = \left(\bar{y} - \beta_1 \bar{x}\right) \sum x_i + \beta_1 \sum x_i^2

Since:

∑xi=nxˉ\sum x_i = n\bar{x}

we get:

∑xiyi=(yˉ−β1xˉ)(nxˉ)+β1∑xi2\sum x_i y_i = \left(\bar{y} - \beta_1 \bar{x}\right)\left(n\bar{x}\right) + \beta_1 \sum x_i^2

Expand:

∑xiyi=nxˉyˉ−nβ1xˉ2+β1∑xi2\sum x_i y_i = n\bar{x}\bar{y} - n\beta_1 \bar{x}^2 + \beta_1 \sum x_i^2

Move nxˉyˉn\bar{x}\bar{y} to the left:

∑xiyi−nxˉyˉ=β1(∑xi2−nxˉ2)\sum x_i y_i - n\bar{x}\bar{y} = \beta_1 \left(\sum x_i^2 - n\bar{x}^2\right)

Therefore:

β1=∑xiyi−nxˉyˉ∑xi2−nxˉ2\beta_1 = \frac{\sum x_i y_i - n\bar{x}\bar{y}}{\sum x_i^2 - n\bar{x}^2}

This can be rewritten using deviations from the mean:

β1=∑i=1n(xi−xˉ)(yi−yˉ)∑i=1n(xi−xˉ)2\boxed{\beta_1 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})(y_i - \bar{y})}{\sum_{i=1}^{n} (x_i - \bar{x})^2}}

This is the familiar OLS slope formula.

Setting the derivatives to zero only finds a stationary point, but here it is the minimum: SSE is a convex function of β0\beta_0 and β1\beta_1 (a sum of squared linear terms), so its only stationary point is the global minimum.


6. The whole idea in one flow

So the entire process is:

Prediction equation→Residual→SSE→Partial derivatives→Set derivatives to zero→β0,β1\text{Prediction equation} \rightarrow \text{Residual} \rightarrow \text{SSE} \rightarrow \text{Partial derivatives} \rightarrow \text{Set derivatives to zero} \rightarrow \beta_0, \beta_1

Specifically:

y^=β0+β1x\hat{y} = \beta_0 + \beta_1 x

Then:

SSE=∑i=1n(yi−β0−β1xi)2\text{SSE} = \sum_{i=1}^{n} \left(y_i - \beta_0 - \beta_1 x_i\right)^2

Differentiate with respect to β0\beta_0:

∂SSE∂β0=0\frac{\partial \text{SSE}}{\partial \beta_0} = 0

which gives:

β0=yˉ−β1xˉ\boxed{\beta_0 = \bar{y} - \beta_1 \bar{x}}

Differentiate with respect to β1\beta_1:

∂SSE∂β1=0\frac{\partial \text{SSE}}{\partial \beta_1} = 0

which gives:

β1=∑(xi−xˉ)(yi−yˉ)∑(xi−xˉ)2\boxed{\beta_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2}}

Therefore, linear regression is not simply guessing a line.

It uses calculus to find the values of β0\beta_0 and β1\beta_1 that produce the minimum possible squared error for the training data.


5. Calculating the slope

For simple linear regression, the optimal slope is:

β1=∑i=1n(xi−xˉ)(yi−yˉ)∑i=1n(xi−xˉ)2\boxed{\beta_1 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})(y_i - \bar{y})}{\sum_{i=1}^{n} (x_i - \bar{x})^2}}

At first, this formula can look intimidating. Let’s break it down.

β1\beta_1

This is the slope of our regression line.

xix_i

This represents an individual observed input value.

For example, if our input is years of experience:

xi=3x_i = 3

could represent an employee with 3 years of experience.

yiy_i

This represents the corresponding observed output.

For example:

yi=42,000y_i = 42{,}000

could be that employee’s actual salary.

xˉ\bar{x} and yˉ\bar{y}

xˉ\bar{x} — This is the mean of all the input values:

1n∑i=1nxi\frac{1}{n} \sum_{i=1}^{n} x_i

yˉ\bar{y} — This is the mean of all the output values:

1n∑i=1nyi\frac{1}{n} \sum_{i=1}^{n} y_i

Numerator

The numerator is:

∑i=1n(xi−xˉ)(yi−yˉ)\sum_{i=1}^{n} (x_i - \bar{x})(y_i - \bar{y})

It measures how xx and yy vary together around their respective means.

  • If larger values of xx generally occur with larger values of yy, this quantity tends to be positive.
  • If larger values of xx generally occur with smaller values of yy, it tends to be negative.

Denominator

The denominator is:

∑i=1n(xi−xˉ)2\sum_{i=1}^{n} (x_i - \bar{x})^2

This measures how much the input xx varies around its mean.

So, conceptually, the slope can be thought of as:

β1=joint variation of X and Yvariation of X\beta_1 = \frac{\text{joint variation of } X \text{ and } Y}{\text{variation of } X}

Calculating the intercept

Once we have calculated the slope, the intercept is much simpler. The formula is:

β0=yˉ−β1xˉ\boxed{\beta_0 = \bar{y} - \beta_1 \bar{x}}

Again, let’s understand the terms:

  • β0\beta_0 = intercept
  • yˉ\bar{y} = mean of the target values
  • β1\beta_1 = calculated slope
  • xˉ\bar{x} = mean of the input values

So the process is essentially:

Data→β1→β0→Regression Equation\boxed{\text{Data} \rightarrow \beta_1 \rightarrow \beta_0 \rightarrow \text{Regression Equation}}

Once we have β0\beta_0 and β1\beta_1, our model is ready.


7. Now prediction becomes extremely simple

Suppose our calculations give us:

β0=25,000\beta_0 = 25{,}000

and:

β1=5,000\beta_1 = 5{,}000

Our model becomes:

y^=25,000+5,000x\hat{y} = 25{,}000 + 5{,}000x

Now suppose someone has 6 years of experience.

We simply substitute:

x=6x = 6

Therefore:

25,000+(5,000)(6)25{,}000 + (5{,}000)(6) 55,00055{,}000

That’s the prediction.

There is no mysterious “AI thinking” happening here.

The model has learned two numbers:

β0=25,000\boxed{\beta_0 = 25{,}000}

and

β1=5,000\boxed{\beta_1 = 5{,}000}

Then it applies the equation.


8. So what does model.fit(X, y) actually do?

This is where the machine-learning terminology can make things sound more complicated than they are.

In Python, we might write:

from sklearn.linear_model import LinearRegression

model = LinearRegression()

model.fit(X, y)

When we call:

model.fit(X, y)

we are essentially asking the algorithm:

“Look at these examples and find the parameters that give the best-fitting linear relationship.”

For simple linear regression, those parameters are:

β0\beta_0

and

β1\beta_1

After fitting the model, we can inspect them:

print(model.intercept_)
print(model.coef_)

Conceptually:

model.fit(X, y)

        ↓

Find β₀ and β₁

        ↓

ŷ = β₀ + β₁x

        ↓

Use the equation for prediction

Scikit-learn uses numerical linear-algebra methods internally rather than necessarily evaluating the formulas above literally step by step. But mathematically, for ordinary simple linear regression with an intercept, it is solving the same least-squares problem.


9. Why do we call this “Machine Learning”?

This is perhaps the most interesting part.

The mathematics itself is not new.

Linear regression has a long history in statistics and mathematics.

What makes it part of a machine-learning workflow is the idea that:

  1. We provide data.
  2. An algorithm estimates parameters from that data.
  3. The resulting model is used to make predictions on new data.

For linear regression, the learned parameters happen to be:

β0,β1\beta_0,\quad\beta_1

For more complicated machine-learning models, there may be thousands, millions, or even billions of parameters.

But the basic idea remains:

Use data to estimate parameters, then use those parameters to produce predictions.


10. The important mental model

When you hear:

“I trained a linear regression model.”

don’t imagine a mysterious black box.

For simple linear regression, imagine this:

                 Training Data
                      │
                      ▼
             Find the best β₀, β₁
                      │
                      ▼
              ŷ = β₀ + β₁x
                      │
                      ▼
                 New input x
                      │
                      ▼
                  Prediction ŷ

The “learning” part is essentially finding the parameters.

The “prediction” part is simply using those parameters in the equation.


11. The bigger picture

Linear regression is a beautiful example because it lets us see what is happening underneath the machine-learning terminology.

At the surface, we might write:

model.fit(X, y)

and:

model.predict(X_new)

But underneath those convenient APIs is mathematics.

The fundamental process is:

 Data → Optimization → Parameters → Prediction \boxed{\ \text{Data}\ \rightarrow\ \text{Optimization}\ \rightarrow\ \text{Parameters}\ \rightarrow\ \text{Prediction}\ }

For simple linear regression:

 Optimization → β0,β1 \boxed{\ \text{Optimization}\ \rightarrow\ \beta_0,\beta_1\ }

and then:

 y^=β0+β1x \boxed{\ \hat{y}=\beta_0+\beta_1x\ }

That’s really what is happening behind the scenes.


Final takeaway

The next time someone says:

“I used machine learning to predict something using linear regression.”

remember what is actually happening.

There is a dataset.

We find the line that best fits the data using a mathematical optimization method such as Ordinary Least Squares.

That gives us two important parameters:

β1=∑i=1n(xi−xˉ)(yi−yˉ)∑i=1n(xi−xˉ)2\boxed{\beta_1 = \frac{\sum_{i=1}^{n}(x_i-\bar{x})(y_i-\bar{y})}{\sum_{i=1}^{n}(x_i-\bar{x})^2}}

and:

 β0=yˉ−β1xˉ \boxed{\ \beta_0=\bar{y}-\beta_1\bar{x}\ }

Then we put them into:

 y^=β0+β1x \boxed{\ \hat{y}=\beta_0+\beta_1x\ }

And that’s the prediction.

The “machine learning” label doesn’t make the mathematics disappear. In this case, the machine is essentially helping us estimate the parameters from data. The prediction itself is just the equation.

Share this article

What Really Happens Behind the Scenes in Linear Regression? | Anish Joshi