Introduction
In this tutorial, we will implement Simple Linear Regression using Python and scikit-learn.
The goal is simple:
Given a person’s weight, can we predict their height?
We will use a small dataset containing two columns:
Weight— the input featureHeight— the target we want to predict
The notebook follows the complete basic workflow:
- Load the dataset
- Visualize the relationship between weight and height
- Separate the independent and dependent variables
- Split the data into training and testing sets
- Standardize the input feature
- Train a Linear Regression model
- Understand the learned coefficient and intercept
- Make predictions
- Evaluate the model using MSE, MAE, RMSE, and
- Make a prediction for a new weight
- Examine residuals and basic regression assumptions
The dataset used in the notebook contains 23 observations, and the first few rows are:
| Weight | Height |
|---|---|
| 45 | 120 |
| 58 | 135 |
| 48 | 123 |
| 60 | 145 |
| 70 | 160 |
The transcript describes this as a simple example for understanding how a regression model is trained and used for prediction.
What Is Simple Linear Regression?
Simple Linear Regression is a supervised machine learning algorithm used to predict a numerical value using one input feature.
For example:
- Weight → Height
- Experience → Salary
- Area → House Price
- Temperature → Electricity Consumption
The word simple is important here because we have only one independent variable.
In our example:
Weightis the independent variable, usually represented by .Heightis the dependent variable, usually represented by .
Our goal is to learn a relationship between them.
The basic equation
The equation of a straight line is:
In machine learning, we commonly write the linear regression equation as:
Where:
- = predicted value
- = input feature
- = intercept
- = coefficient or slope
For our problem:
The model’s job is to learn the values of and that produce predictions as close as possible to the observed heights.
1. Load the Required Libraries
The notebook starts by importing the required Python libraries.
import pandas as pd
import matplotlib.pyplot as plt
import numpy as np
%matplotlib inline
Why do we need these libraries?
Pandas
Used for loading and working with the dataset.
import pandas as pd
Matplotlib
Used for creating visualizations such as scatter plots.
import matplotlib.pyplot as plt
NumPy
Used for numerical operations. Later, we use it to calculate RMSE.
import numpy as np
%matplotlib inline tells Jupyter Notebook to display Matplotlib plots directly inside the notebook.
2. Read the Dataset
The dataset is stored in a CSV file named height-weight.csv.
df = pd.read_csv('height-weight.csv')
df.head()
pd.read_csv() loads the CSV file into a Pandas DataFrame.
df.head() displays the first five rows.
The notebook shows:
Weight Height
0 45 120
1 58 135
2 48 123
3 60 145
4 70 160
At this point, we have our dataset loaded into memory.
3. Visualize the Data
Before training a machine learning model, it is useful to understand the relationship between the variables.
We can create a scatter plot:
plt.scatter(df['Weight'], df['Height'])
plt.xlabel("Weight")
plt.ylabel("Height")
A scatter plot represents each observation as a point.
- The horizontal axis represents
Weight. - The vertical axis represents
Height.
This allows us to visually inspect whether there appears to be a relationship between the two variables.
If the points roughly follow a straight-line pattern, a linear regression model may be appropriate.
4. Independent and Dependent Variables
Now we need to tell our machine learning model which column is the input and which column is the output.
X = df[['Weight']] # independent feature
y = df['Height'] # dependent feature
Here:
X = df[['Weight']]
contains the input feature.
And:
y = df['Height']
contains the target we want to predict.
Why is X written with double brackets?
There is an important difference between:
df['Weight']
and:
df[['Weight']]
The first returns a Pandas Series, which is one-dimensional.
The second returns a DataFrame, which is two-dimensional.
Scikit-learn expects feature data in the general form:
Our feature matrix therefore has the shape:
(23, 1)
That means:
- 23 observations
- 1 feature
5. Split the Dataset into Training and Testing Data
A machine learning model should not be evaluated using only the data it was trained on.
We therefore divide the dataset into:
- Training data — used to learn the model
- Testing data — used to evaluate the model on unseen observations
First, import train_test_split:
from sklearn.model_selection import train_test_split
Then split the data:
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.20,
random_state=42
)
What does test_size=0.20 mean?
It means that 20% of the observations are reserved for testing.
The dataset contains 23 observations.
The notebook therefore produces:
X.shape
(23, 1)
and:
X_train.shape, X_test.shape, y_train.shape, y_test.shape
((18, 1), (5, 1), (18,), (5,))
So:
- 18 observations are used for training.
- 5 observations are used for testing.
What is random_state=42?
The train-test split involves randomness.
random_state=42 fixes the random seed so that the same split can be reproduced when the code is run again.
The specific number 42 has no special mathematical meaning here. Any fixed integer could be used.
6. Standardize the Input Feature
The notebook next standardizes the input feature using StandardScaler.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
Then:
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)
Standardization converts values approximately to a scale with:
- mean = 0
- standard deviation = 1
The standardization formula is:
Where:
- = original value
- = mean calculated from the training data
- = standard deviation calculated from the training data
- = standardized value
Why do we use fit_transform() for training data?
X_train = scaler.fit_transform(X_train)
fit calculates the mean and standard deviation from the training data.
transform uses those statistics to standardize the training data.
For the test data, we use:
X_test = scaler.transform(X_test)
We deliberately do not use:
scaler.fit_transform(X_test)
The reason is important.
The test set should represent unseen data. We should not allow information from the test set to influence the preprocessing parameters learned from the training set.
This helps prevent data leakage.
An important correction
Standardization is not required for ordinary linear regression implemented by scikit-learn’s LinearRegression.
The transcript presents scaling as necessary for linear regression because of gradient descent. That is an oversimplification.
For this particular model, LinearRegression can work directly with the original Weight values. Standardization is still valid, and it can be useful when building a consistent preprocessing pipeline or when using algorithms that are sensitive to feature scale.
In this notebook, however, the model is trained on the standardized weight values, so the learned coefficient is the coefficient with respect to the standardized feature.
7. Train the Simple Linear Regression Model
Now we can create the regression model.
from sklearn.linear_model import LinearRegression
regressor = LinearRegression()
At this point, the model object exists, but it has not learned anything yet.
We train it using:
regressor.fit(X_train, y_train)
The .fit() method learns the relationship between the input feature and the target.
Conceptually, the model is trying to find the line:
that best represents the training data.
8. Understand the Coefficient and Intercept
After training, we can inspect the learned parameters.
print("The slope or coefficient of weight is ", regressor.coef_)
print("Intercept:", regressor.intercept_)
The notebook produces:
The slope or coefficient of weight is [17.03440872]
Intercept: 157.5
So the learned model is approximately:
However, remember that x here is the standardized weight, not the original weight in kilograms.
Therefore, it would be incorrect to interpret 17.0344 as:
“Every additional kilogram increases height by 17 cm.”
That interpretation would only make sense if the model had been trained directly on the original weight values.
Instead, 17.0344 means that a one-standard-deviation increase in weight is associated with an estimated increase of about 17.03 units in predicted height for this fitted model.
9. Plot the Regression Line
We can visualize the predictions produced by the model.
plt.scatter(X_train, y_train)
plt.plot(X_train, regressor.predict(X_train), 'r')
The scatter plot shows the training observations.
The red line represents the model’s predictions.
The goal of linear regression is to find a line that provides predictions that are close to the observed target values.
A residual for an observation is:
Where:
- = actual value
- = predicted value
- = residual
The model is fitted using squared errors, rather than simply minimizing the raw sum of positive and negative differences.
This distinction is important because positive and negative residuals could otherwise cancel each other out.
10. Make Predictions on the Test Data
After training the model, we can use it to predict the heights of the test observations.
y_pred_test = regressor.predict(X_test)
The notebook gives:
y_pred_test
[161.08467086,
161.08467086,
129.30415610,
177.45645118,
148.56507414]
The corresponding actual values are:
y_test
177
170
120
182
159
We can compare them:
| Actual Height | Predicted Height | Residual |
|---|---|---|
| 177 | 161.08 | 15.92 |
| 170 | 161.08 | 8.92 |
| 120 | 129.30 | -9.30 |
| 182 | 177.46 | 4.54 |
| 159 | 148.57 | 10.43 |
A positive residual means the model predicted too low.
A negative residual means the model predicted too high.
For example:
So the first test observation has a residual of approximately 15.92.
11. Evaluate the Model
Looking at predictions manually is useful, but we need numerical metrics to evaluate the model.
The notebook calculates:
- Mean Squared Error
- Mean Absolute Error
- Root Mean Squared Error
- Adjusted
Mean Squared Error — MSE
MSE measures the average squared difference between actual and predicted values.
Where:
- = number of observations
- = actual value
- = predicted value
The code is:
from sklearn.metrics import mean_squared_error, mean_absolute_error
mse = mean_squared_error(y_test, y_pred_test)
The notebook produces:
109.77592599051664
A lower MSE generally indicates smaller prediction errors.
However, MSE is expressed in squared units, which makes it harder to interpret directly.
12. Mean Absolute Error — MAE
MAE calculates the average absolute prediction error.
The code is:
mae = mean_absolute_error(y_test, y_pred_test)
The result is:
9.822657814519232
This means that, on average, the predictions differ from the actual heights by approximately 9.82 height units on this test set.
MAE is often easier to understand than MSE because it is expressed in the same units as the target.
13. Root Mean Squared Error — RMSE
RMSE is the square root of MSE.
The notebook calculates it using:
rmse = np.sqrt(mse)
The result is:
10.477400726827081
So the RMSE is approximately:
Like MAE, RMSE is expressed in the same units as the target.
RMSE gives greater influence to larger errors because the errors are squared before taking the average.
14. Score
Now we calculate the coefficient of determination, commonly called .
from sklearn.metrics import r2_score
score = r2_score(y_test, y_pred_test)
The notebook produces:
0.776986986042344
So:
or approximately:
The formula is:
Where:
- = actual value
- = predicted value
- = mean of the actual target values
What does mean?
For this particular test set, the model explains about 77.7% of the variation in the target values relative to a baseline that always predicts the mean target.
It is better to say this than:
“The model is 77.7% accurate.”
is not an accuracy percentage in the same sense as classification accuracy.
Also, because the test set contains only 5 observations, this score should not be treated as a highly reliable estimate of how the model would perform on a much larger unseen dataset.
15. Adjusted
The notebook also calculates adjusted :
1 - (1-score)*(len(y_test)-1)/(len(y_test)-X_test.shape[1]-1)
The result is:
0.7026493147231252
So:
The formula is:
Where:
- = ordinary
- = number of observations
- = number of predictors
In this example:
Adjusted penalizes the model for adding predictors that do not provide enough explanatory value.
This becomes particularly useful when comparing multiple linear regression models with different numbers of predictors.
For a simple one-feature model, its usefulness is limited, especially with such a small test set.
16. Make a Prediction for a New Weight
Now suppose we receive a new observation:
A person weighs 80 kg. What height does the model predict?
Because the model was trained using standardized weights, we must apply the same scaler to the new value.
scaled_weight = scaler.transform([[80]])
scaled_weight
The notebook produces approximately:
[[0.32350772]]
We then pass the standardized value to the trained model:
print(
"The height prediction for weight 80 kg is :",
regressor.predict([scaled_weight[0]])
)
The result is:
The height prediction for weight 80 kg is : [163.01076266]
So the model predicts a height of approximately:
for a weight of 80 kg.
Why must we scale the new value?
The model learned its coefficient using standardized inputs.
Therefore, a new input must go through the same preprocessing transformation before it is given to the model.
A useful rule is:
Whatever preprocessing was applied during training must also be applied to new data before prediction.
17. What Are Residuals?
A residual is the difference between an actual value and the model’s prediction.
The notebook calculates:
residuals = y_test - y_pred_test
residuals
The results are:
15.915329
8.915329
-9.304156
4.543549
10.434926
For example:
The model underestimated that observation by approximately 15.92 units.
Residuals are extremely useful because they help us understand where the model is making mistakes.
18. Check the Residual Distribution
The notebook originally uses:
import seaborn as sns
sns.distplot(residuals, kde=True)
However, distplot() is deprecated in modern versions of Seaborn.
A modern replacement is:
import seaborn as sns
sns.histplot(residuals, kde=True)
plt.xlabel("Residual")
plt.ylabel("Count")
plt.show()
A good regression model often has residuals that are approximately centered around zero, with no obvious problematic pattern.
For some statistical inference procedures, normally distributed residuals are an important assumption. However, normality is not a universal requirement for obtaining predictions from ordinary linear regression.
Also, with only five test observations, it is not meaningful to make a strong conclusion about residual normality from this histogram alone.
19. Residuals vs. Predictions
The notebook also creates:
plt.scatter(y_pred_test, residuals)
This plots:
- predicted values on the x-axis
- residuals on the y-axis
A useful residual plot should ideally show residuals scattered around zero without a systematic pattern.
For example, we generally do not want to see:
- a clear curve
- a funnel shape
- increasing spread as predictions increase
- long runs of positive or negative residuals
A random-looking spread around zero is more consistent with a well-specified linear model.
Important correction
The transcript describes this as looking for a “uniform distribution.”
That is not the correct statistical requirement.
We are not specifically looking for a mathematical uniform distribution.
Instead, we are looking for residuals that are randomly scattered around zero without systematic structure.
This distinction matters because residual diagnostics are about detecting patterns that suggest the model is missing something.
20. Important Regression Assumptions
Residual analysis helps us examine several assumptions associated with linear regression.
1. Linearity
The relationship between the predictor and the expected target should be approximately linear.
In our example, we are assuming that weight and expected height can reasonably be represented by a straight-line relationship.
2. Independent observations/errors
The observations should generally be independent of one another.
This is particularly important with time-series or sequential data, where neighboring observations may be related.
3. Constant variance
The spread of residuals should be reasonably consistent across the range of predicted values.
This property is called homoscedasticity.
A funnel-shaped residual plot can indicate that the error variance changes with the predicted value.
4. Residual normality
For certain statistical inference tasks, residuals are often assumed to be approximately normally distributed.
This is less important for simply generating predictions, and it should not be judged from only five test observations.
5. No strong systematic pattern in residuals
If residuals form a clear curve or another pattern, the linear model may be missing an important relationship.
21. Complete Workflow
At this point, the entire simple linear regression workflow can be summarized as:
Dataset
↓
Explore the data
↓
Visualize the relationship
↓
Separate X and y
↓
Train/Test Split
↓
Fit scaler on training data
↓
Transform training and test data
↓
Create LinearRegression model
↓
Train with X_train and y_train
↓
Predict on X_test
↓
Compare predictions with y_test
↓
Calculate evaluation metrics
↓
Inspect residuals
↓
Use model for new predictions
This is the basic pattern you will see repeatedly in supervised machine learning.
22. What We Learned About the Model
The trained model learned:
Coefficient = 17.03440872
Intercept = 157.5
Because the input was standardized, the model can be represented as:
The model’s test-set results were:
| Metric | Result |
|---|---|
| MSE | 109.776 |
| MAE | 9.823 |
| RMSE | 10.477 |
| 0.777 | |
| Adjusted | 0.703 |
These numbers describe the performance of this particular model on this particular five-observation test set.
They should not be interpreted as a guarantee that the model will perform equally well on new datasets.
23. Limitations of This Example
This example is intentionally small and simple so that the mechanics of linear regression are easy to understand.
There are several limitations.
Very small dataset
The dataset contains only 23 observations.
Only five observations are used for testing.
That makes the evaluation metrics sensitive to individual observations.
Weight alone cannot explain height completely
Real human height depends on many factors.
Using only weight is a simplified demonstration of regression rather than a realistic height prediction system.
Scaling is unnecessary for this particular estimator
LinearRegression can be trained directly on the original weight values.
The notebook uses StandardScaler, but scaling should not be presented as a mandatory step for ordinary linear regression.
The test set is very small
An of approximately 0.777 looks useful, but five test observations are not enough to confidently estimate real-world generalization performance.
Residual conclusions are limited
There are only five test residuals, so we should avoid making strong claims about normality or other statistical assumptions based on these plots.
Key Takeaways
- Simple Linear Regression predicts a numerical target using one input feature.
- The basic model is represented by:
- In this example:
- We split the data into training and testing sets so that we can evaluate the model on observations it did not train on.
- When using preprocessing, fit the scaler only on the training data:
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)
LinearRegression().fit()learns the model parameters.coef_gives the learned coefficient.intercept_gives the intercept.- Predictions are generated with:
regressor.predict(X_test)
- MAE tells us the average absolute prediction error.
- MSE measures average squared error.
- RMSE is the square root of MSE and is expressed in the target’s units.
- measures how much variation in the target is explained by the model relative to a mean-prediction baseline; it is not simply “accuracy.”
- **Adjusted ** accounts for the number of predictors and is more useful when comparing models with different numbers of features.
- Residuals are the differences between actual and predicted values:
- A good residual plot should generally show no obvious systematic pattern rather than specifically following a “uniform distribution.”
- The example predicts a height of approximately 163.01 for a person weighing 80 kg.
- Most importantly, this notebook demonstrates the complete machine learning workflow from data → preprocessing → training → prediction → evaluation → diagnostics.
This simple example provides the foundation for the next step: Multiple Linear Regression, where we use multiple independent features to predict a target.
Share this article