Twelve months of electricity bills
A household in Bengaluru tracks two numbers each month: hours the air conditioner ran per day, and the electricity bill. Plot one against the other and the points sit close to a straight line. Linear regression is the procedure that finds which straight line, and building linear regression from scratch takes about ten lines of numpy.
The model is bill = w × hours + b. Two numbers to find: w, the rupees each extra hour of cooling costs, and b, the bill when the air conditioner never runs. Everything in this chapter is about choosing those two numbers well.
The assumptions a linear model makes about your data are a separate subject with its own treatment. Here the concern is purely mechanical: what the algorithm computes, two ways of computing it, and why they agree.
Infinitely many lines, one cost function
Any pair of w and b draws a line. The cost function scores them so you can prefer one. Mean squared error is the standard choice: for each month, subtract the true bill from what the line predicts, square the difference, and average.
import numpy as np
ac_hours = np.array([0., 1.5, 2., 3.5, 4., 5.5, 6., 7.5, 8., 9.5, 10., 11.5])
bill = np.array([820., 1180., 1310., 1690., 1840., 2290.,
2410., 2810., 2960., 3390., 3520., 3910.])
def mse(w, b):
return float((((w * ac_hours + b) - bill) ** 2).mean())
for w, b in [(200, 900), (270, 800), (273.24, 773.06)]:
print(f"w={w:7.2f} b={b:7.2f} MSE={mse(w, b):10.1f}")
# w= 200.00 b= 900.00 MSE= 152591.7
# w= 270.00 b= 800.00 MSE= 662.5
# w= 273.24 b= 773.06 MSE= 464.9
A poor guess scores 152,592. A close guess scores 662. The best possible pair scores 464.9, and no other combination of two numbers scores lower on this data.
Squaring matters, as how machines learn from data explained: it makes one badly missed month hurt more than several slightly missed ones. That choice is what makes the next step solvable in closed form.
Linear regression from scratch, in two formulas
Because the cost is a smooth bowl in w and b, you can find the bottom exactly rather than searching for it. Set both derivatives to zero and solve, and two formulas drop out.
The slope is the covariance of the two variables divided by the variance of the input. In words: how much the two move together, divided by how much the input moves on its own. The intercept then follows, because the fitted line always passes through the mean of both variables.
x_bar, y_bar = ac_hours.mean(), bill.mean()
slope = (((ac_hours - x_bar) * (bill - y_bar)).sum()
/ ((ac_hours - x_bar) ** 2).sum())
intercept = y_bar - slope * x_bar
print(round(slope, 4), round(intercept, 4))
# 273.2368 773.0551
Each extra hour of air conditioning costs about 273 rupees a month, on top of a fixed 773 rupees. That is a usable finding, and you derived it without a library.
Now confirm it:
from sklearn.linear_model import LinearRegression
model = LinearRegression().fit(ac_hours.reshape(-1, 1), bill)
print(round(float(model.coef_[0]), 4), round(float(model.intercept_), 4))
# 273.2368 773.0551
print(round(float(model.predict([[6.5]])[0]), 1)) # 2549.1
Identical to four decimal places, because LinearRegression solves the same system. Six and a half hours a day predicts a bill of 2549 rupees.
Why anyone bothers with the iterative version
The closed form is exact and fast, so why does gradient descent exist? Because the formula requires inverting a matrix whose size grows with your feature count, which stops being practical around tens of thousands of features, and because most models that are not linear regression have no closed form at all.
# standardise first, which keeps the step size stable
xs = (ac_hours - ac_hours.mean()) / ac_hours.std()
w, b, lr = 0.0, 0.0, 0.1
for step in range(500):
error = (w * xs + b) - bill
w -= lr * 2 * (error * xs).mean()
b -= lr * 2 * error.mean()
print(round(w, 3), round(b, 3)) # 955.515 2344.167
# convert back to the original units
w_original = w / ac_hours.std()
b_original = b - w * ac_hours.mean() / ac_hours.std()
print(round(w_original, 4), round(b_original, 4))
# 273.2368 773.0551
The same two numbers, reached by stepping downhill 500 times instead of solving. The mechanics of the step size and the scaling are covered in gradient descent explained; what matters here is that both routes land on the identical answer, because the cost surface has exactly one bottom.
Note the conversion back. Fitting on standardised inputs gives coefficients in standardised units, so a slope of 955.5 means “rupees per standard deviation of hours” rather than rupees per hour. Forgetting that step is a common source of coefficients that look absurd.
What a single feature cannot do
This model has one input, and that limits it in ways worth naming before you trust it anywhere.
It assumes the relationship is a straight line. If each additional hour costs more than the last, because of tariff slabs, a line will be wrong at both ends while looking reasonable in the middle.
It assumes the only thing that matters is air conditioner hours. Occupancy, outdoor temperature, geyser use and tariff changes are all absent, so their effects are dumped into the error term.
And it extrapolates confidently. Ask it about 40 hours a day and it will answer 11,703 rupees without noticing that a day has 24 hours. A linear model has no concept of the range it was trained on, which is a property you will see fail spectacularly later in this section.
Adding more inputs is the next chapter. The full estimator reference is in the LinearRegression documentation, and the fit and predict contract it follows was introduced in your first model in scikit-learn.
Frequently Asked Questions
How do you calculate linear regression by hand?
The slope is the covariance of the input and target divided by the variance of the input, which in code is ((x - x.mean()) * (y - y.mean())).sum() / ((x - x.mean()). The intercept is
** 2).sum()y.mean() - slope * x.mean(), since the fitted line passes through both means.
Why does linear regression use squared error?
Squaring makes every error positive, penalises large misses more than small ones, and produces a smooth bowl-shaped cost surface with a single minimum that can be solved exactly. Absolute error is more robust to outliers but has no closed-form solution and needs an iterative solver.
Is gradient descent or the normal equation better for linear regression?
The closed form is exact and faster for small problems, which is why scikit-learn uses it. Gradient descent wins when the feature count runs to tens of thousands, when data arrives as a stream, or when you move to models that have no closed-form solution at all.
Key Takeaways
- Derive the slope and intercept directly when you have one feature, because two lines of numpy give the exact answer that scikit-learn returns.
- Read mean squared error as the number the algorithm is minimising, and check it on a few candidate lines to see how sharply it discriminates.
- Standardise your input before running gradient descent, then convert the coefficients back, or they will be in units of standard deviations.
- Expect both the closed form and gradient descent to agree exactly, since the cost surface for linear regression has a single minimum.
- Treat any prediction outside your training range with suspicion, as a linear model will answer confidently for inputs that cannot physically occur.