When people say:
“I used machine learning to predict the price using several features.”
it can sound like something far more complicated than the simple line we started with.
But if we strip away the terminology, multiple linear regression is the same mathematical problem as simple linear regression, with more inputs.
At its core, we are still just trying to find a handful of numbers:
- The intercept —
- One coefficient per input —
Once we have those numbers, making a prediction is simply a matter of putting values into an equation.
This article builds on simple linear regression. If you have not read that one yet, start there. Here we reuse the same ideas: the error, the SSE, partial derivatives, and setting them to zero.
1. The basic idea
Suppose we want to predict someone’s salary. Years of experience alone is a good start, but it is not the whole story. Let’s add a second input: the number of professional certifications they hold.
We have some historical data (salary in dollars):
| Years of Experience | Certifications | Salary |
|---|---|---|
| 1 | 0 | 24,500 |
| 2 | 1 | 30,300 |
| 3 | 0 | 32,300 |
| 4 | 2 | 41,600 |
| 5 | 1 | 43,600 |
| 6 | 3 | 52,800 |
| 7 | 2 | 53,500 |
| 8 | 3 | 61,400 |
With one input, we fitted a line. With two inputs, we fit a plane. With three or more, we fit a hyperplane, which is hard to draw but works exactly the same way.
The equation of our model is:
Where:
- = predicted salary
- = years of experience
- = number of certifications
- = intercept
- = coefficient of years of experience
- = coefficient of certifications
For any number of inputs , the general form is:
That’s essentially our entire prediction model.
2. What does each coefficient mean?
Each coefficient tells us how much the predicted output changes when that one input increases by one unit, while all the other inputs stay the same.
That last part is the important difference from simple regression.
For example, suppose we calculate:
This means that for every additional year of experience, with the number of certifications held fixed, the predicted salary increases by approximately 4,161.
Similarly:
means that each additional certification adds approximately 2,596 to the predicted salary, with years of experience held fixed.
So if our model is:
then:
- 3 years, 0 certifications → 32,364
- 4 years, 0 certifications → 36,525
- 4 years, 1 certification → 39,121
Going from the first row to the second changes only experience, and the prediction moves by . Going from the second row to the third changes only certifications, and the prediction moves by .
Each coefficient controls how steep the plane is in the direction of its own input.
3. What does the intercept mean?
The intercept is the predicted value of when all inputs are zero.
Suppose:
Then our model says:
When experience and certifications are both zero:
Therefore:
So the intercept is:
The point where the regression plane crosses the -axis.
Depending on the problem, the intercept may or may not have a meaningful real-world interpretation. Mathematically, however, it is one of the parameters needed to define the plane.
4. So where do the coefficients come from?
This is where the interesting mathematics happens.
We don’t simply guess:
Instead, we calculate values that produce the best-fitting plane for the observed data.
We use the same approach as before: Ordinary Least Squares (OLS).
The idea behind OLS is unchanged:
Find the plane that minimizes the total squared difference between the actual values and the predicted values.
For each observation, we have an error:
Since some errors are positive and others are negative, simply adding them together could cause them to cancel each other out.
Instead, we square each error.
The objective becomes:
Since:
we can write:
Here means the value of input 1 for observation , and is the value of input 2 for the same observation.
OLS finds the values of , and that minimize this quantity.
5. Deriving the coefficients using partial derivatives
We have already defined the Sum of Squared Errors:
Our goal is to find the values of , and that make this SSE as small as possible.
Because we now have three unknowns, we take the partial derivative with respect to each one. That gives us three equations, one for each unknown.
Step 1: Find
We differentiate with respect to . The summation can stay outside, and we use the chain rule. If:
then:
Since:
we get:
At the minimum, the derivative must be zero:
2
Divide by and expand the summation:
Divide by :
Therefore:
Compare this with simple regression, where we had . The idea is identical: the fitted plane always passes through the point of means.
Step 2: Find
Now we differentiate the same SSE with respect to . Using the chain rule again, the derivative of the inside is:
Therefore:
Set it to zero, divide by , and expand:
Step 3: Find
The same steps for , where the derivative of the inside is , give:
Step 4: The normal equations
Now we collect all three equations, moving the terms with to the left:
These are called the normal equations. There are three equations and three unknowns, so we can solve for , and .
Notice what changed compared with simple regression. With one input, we could solve by hand with a couple of substitutions and get a tidy formula. With more inputs, the equations are coupled: every coefficient appears in every equation, through terms like . That cross term measures how the inputs move together.
Setting the derivatives to zero only finds a stationary point, but here it is the minimum: SSE is a convex function of the coefficients (a sum of squared linear terms), so when the inputs are not perfectly redundant, its only stationary point is the global minimum.
6. The matrix form
Writing out every sum gets messy as the number of inputs grows. Linear algebra gives us a compact way to say exactly the same thing.
Stack the data into a matrix with one row per observation. The first column is all ones, which handles the intercept:
Then all predictions at once are:
and the SSE becomes:
Expanding it:
Taking the derivative with respect to (this is the same three partial derivatives, stacked into one vector):
Set it to zero:
This is the normal equations in one line. If can be inverted, we multiply both sides by its inverse:
This is the OLS solution for any number of inputs.
The convexity argument also has a clean matrix version. The second derivative of the SSE is , which is positive definite when the columns of are linearly independent. That guarantees a single, unique minimum.
If one input is a perfect combination of the others, for example , then cannot be inverted and there is no unique answer. This problem is called perfect multicollinearity.
7. The whole idea in one flow
So the entire process is:
Specifically:
Then:
Differentiate with respect to each coefficient and set to zero:
which gives the normal equations, and in matrix form:
Therefore, multiple linear regression is not simply guessing a plane.
It uses calculus to find the values of that produce the minimum possible squared error for the training data.
8. Working through the example
Let’s apply this to our salary data. First, we compute the sums we need. With :
| Quantity | Value |
|---|---|
| 36 | |
| 12 | |
| 204 | |
| 28 | |
| 71 | |
| 340,000 | |
| 1,748,900 | |
| 606,700 |
Plugging these into the normal equations gives:
In matrix form, this is with:
Solving this system of three equations gives:
We can check the intercept with the formula from Step 1. The means are , and , so:
It matches.
Now let’s look at how well the fitted plane matches the data. The chart below plots each employee’s actual salary against the salary the model predicts. If the model were perfect, every point would sit on the dashed line.
Actual salary versus predicted salary for the multiple regression model, with points close to the perfect-prediction line

9. Now prediction becomes extremely simple
Our calculations give us:
Now suppose someone has 6 years of experience and 2 certifications.
We simply substitute:
Therefore:
That’s the prediction.
There is no mysterious “AI thinking” happening here.
The model has learned three numbers:
Then it applies the equation.
10. What changes when we add more inputs?
The mathematics is the same, but a few ideas deserve attention.
A coefficient depends on which other inputs are in the model.
If we fit salary using years of experience alone, the slope comes out at about 5,212 per year. With certifications added, the experience coefficient drops to about 4,161.
Why? In our data, people with more experience also tend to hold more certifications. In the simple model, the experience slope absorbs part of the certification effect. Once certifications are included as their own input, the model can separate the two effects.
This is why we say each coefficient is the effect of one input holding the others constant.
Inputs that move together cause trouble.
When two inputs are strongly correlated, is close to being non-invertible. The model may still predict well, but the individual coefficients become unstable: small changes in the data can swing them widely. This is called multicollinearity.
Inputs on different scales are hard to compare.
A coefficient of 4,161 for years and 2,596 for certifications does not tell us which input matters more, because they are measured in different units. Standardizing the inputs is a common way to make coefficients comparable.
11. So what does model.fit(X, y) actually do?
This is where the machine-learning terminology can make things sound more complicated than they are.
In Python, we might write:
import pandas as pd
from sklearn.linear_model import LinearRegression
df = pd.DataFrame({
"experience": [1, 2, 3, 4, 5, 6, 7, 8],
"certifications": [0, 1, 0, 2, 1, 3, 2, 3],
"salary": [24500, 30300, 32300, 41600, 43600, 52800, 53500, 61400],
})
X = df[["experience", "certifications"]]
y = df["salary"]
model = LinearRegression()
model.fit(X, y)
When we call:
model.fit(X, y)
we are essentially asking the algorithm:
“Look at these examples and find the parameters that give the best-fitting linear relationship.”
For multiple linear regression, those parameters are:
After fitting the model, we can inspect them:
print(model.intercept_) # about 19881
print(model.coef_) # about [4161, 2596]
These are the same numbers we found by hand.
Conceptually:
model.fit(X, y)
↓
Find β₀, β₁, ..., βₚ
↓
ŷ = β₀ + β₁x₁ + ... + βₚxₚ
↓
Use the equation for prediction
Scikit-learn uses numerical linear-algebra methods internally rather than necessarily evaluating literally. Inverting a matrix directly can be numerically unstable, so practical solvers use more stable factorizations. But mathematically, for ordinary multiple linear regression with an intercept, it is solving the same least-squares problem.
12. Why do we call this “Machine Learning”?
This is perhaps the most interesting part.
The mathematics itself is not new.
Linear regression has a long history in statistics and mathematics.
What makes it part of a machine-learning workflow is the idea that:
- We provide data.
- An algorithm estimates parameters from that data.
- The resulting model is used to make predictions on new data.
For multiple linear regression, the learned parameters happen to be:
For more complicated machine-learning models, there may be thousands, millions, or even billions of parameters.
But the basic idea remains:
Use data to estimate parameters, then use those parameters to produce predictions.
13. The important mental model
When you hear:
“I trained a multiple regression model.”
don’t imagine a mysterious black box.
For multiple linear regression, imagine this:
Training Data
│
▼
Find the best β₀, β₁, ..., βₚ
│
▼
ŷ = β₀ + β₁x₁ + ... + βₚxₚ
│
▼
New inputs x₁ ... xₚ
│
▼
Prediction ŷ
The “learning” part is essentially finding the parameters.
The “prediction” part is simply using those parameters in the equation.
14. The bigger picture
Multiple linear regression is a beautiful example because it shows that going from one input to many changes the bookkeeping, not the idea.
At the surface, we might write:
model.fit(X, y)
and:
model.predict(X_new)
But underneath those convenient APIs is mathematics.
The fundamental process is:
For multiple linear regression:
and then:
That’s really what is happening behind the scenes.
Final takeaway
The next time someone says:
“I used machine learning to predict something using several features.”
remember what is actually happening.
There is a dataset.
We find the plane that best fits the data using a mathematical optimization method such as Ordinary Least Squares.
That gives us the parameters:
Then we put them into:
And that’s the prediction.
The “machine learning” label doesn’t make the mathematics disappear. In this case, the machine is essentially helping us estimate the parameters from data. The prediction itself is just the equation.
Share this article