Linear Regression
- #machine-learning
- #linear-regression
- #supervised-learning
- #regression
- #gradient-descent
- #regularization
- #feature-scaling
- #polynomial-regression
Linear Regression
Part of: Machine Learning Key concepts: Supervised Learning, Cost Function, Gradient Descent, Normal Equation, Feature Scaling, Regularization, Overfitting, Polynomial Regression
Agenda
- What is regression
- Single-variable linear regression
- Cost functions
- Gradient descent and optimization
- Multi-variable linear regression
- Normalization (feature scaling)
- Regularization
- Polynomial regression
What is Regression?
Regression is a supervised learning task where the goal is to predict a continuous numerical output from an input feature vector .
Motivating example — House prices
Square Feet House Price ($1000s) 1400 245 1600 312 1700 279 1875 308 1100 199 2350 405 2450 324 Goal: given a new house's size, predict its price.
Regression vs. classification: regression outputs a real value; classification outputs a discrete class label.
Notation
- — number of training examples
- (or ) — number of input features
- — feature vector of the -th training example
- — target value of the -th example
- — the -th training pair
- — the model's prediction for
- (or ) — the learned model parameters (weights)
Single-Variable Linear Regression
With a single input feature , the hypothesis is a straight line:
- — intercept (bias)
- — slope
The line that "best fits" the data is the one that minimizes some measure of error between the predicted and the actual .
Residuals
The residual for the -th point is:
Performance Metric — Least Squares
We measure the quality of a fit by the sum of squared residuals. Squaring has three effects:
- removes sign so positive and negative errors don't cancel,
- penalises large errors more than small ones,
- produces a smooth, differentiable, convex function — easy to optimize.
Cost function (Mean Squared Error)
The is a convenience: it cancels when we differentiate.
Why "least squares"?Minimizing is equivalent to finding the line with the smallest total squared vertical distance from the data points.
Cost Function Intuition
- Every choice of gives a different line — and hence a different cost.
- Plotting over the parameter space produces a bowl-shaped surface (a paraboloid).
- The cost function for linear regression with MSE is convex — there is a single global minimum and no local minima to get stuck in.
Contour viewIn 2D, the cost surface can be drawn as concentric ellipses (contours). The centre of the smallest ellipse corresponds to the optimal parameters.
Optimization — Gradient Descent
We want to find that minimizes . Two approaches:
- Iterative — gradient descent,
- Analytical — the normal equation (closed form).
Gradient descent — key idea
- At the minimum, the derivative (slope) is zero.
- To move "downhill", take a step against the sign of the gradient.
- The magnitude of the step scales with both the gradient and a learning rate .
Update rule
For each parameter , simultaneously update:
Simultaneous updateAll parameters must be updated from the same snapshot of . Update a temp variable for each, then assign together — never update and then use the new when computing the step for .
Gradients for single-variable linear regression
Learning Rate
The learning rate controls the step size of gradient descent.
| Learning rate | Behaviour |
|---|---|
| Too small | Very slow convergence; can stall in shallow regions or local minima. |
| Too large | Overshoots the minimum; cost can diverge. |
| Good | Steady, monotonic decrease in . |
DiagnosticPlot vs iteration. It should decrease smoothly. If it oscillates or grows → is too large.
Hyper-parameters
- Parameters () — learned by the algorithm from data.
- Hyper-parameters (, number of iterations, regularization strength , batch size, polynomial degree …) — chosen before training, not learned by the algorithm itself.
Hyper-parameters are typically tuned on a validation set.
Try it: step through the graph
type: gradient-descent
Analytical Solution — The Normal Equation
For linear regression, we don't actually need an iterative algorithm — there's a closed-form solution.
Matrix formulation
Stack the data:
- — design matrix, rows are training examples (with a leading column of s for the intercept).
- — target vector.
- — parameters.
The model becomes , and the cost is:
Normal equation
Setting gives:
Gradient descent vs closed form
| Aspect | Gradient descent | Normal equation |
|---|---|---|
| Needs ? | Yes | No |
| Needs iterations? | Yes | No |
| Works with many features | Good — per iteration | Slow — to invert |
| Feature scaling required? | Yes (for fast convergence) | No |
| Non-invertible case? | Still works | Fails (redundant features, ) |
Rule of thumb: use the normal equation for small (say ), gradient descent for larger problems.
Multi-Variable Linear Regression
With features, the hypothesis generalises to:
(where we prepend to absorb the intercept into ).
Vectorised cost and gradient
The update rule is the same, just vectorised:
Variants of Gradient Descent
The cost involves a sum over all training points. Depending on how many points we use per update, we get three flavours:
| Variant | Points per update | Cost curve | Compute per step | Notes |
|---|---|---|---|---|
| Batch GD | All | Smooth | High | True gradient; slow on large datasets. |
| Stochastic GD (SGD) | 1 random point | Very noisy | Low | Fast iterations; never fully settles. |
| Mini-batch GD | Batch of | Slightly noisy | Medium | Sweet spot in practice; uses vectorisation. |
Why stochastic noise can helpWhen the cost function is irregular, the noise in SGD can jump out of local minima — giving it a better chance of finding the global optimum. The flip side: it can never fully settle at the exact minimum. Common fix: use a decaying learning rate so steps shrink over time.
Feature Scaling
What happens in multivariate linear regression if the scales of features differ drastically? You need a very small learning rate, and training becomes slow.
When features have very different scales (e.g. square feet vs bedrooms ), the cost-function contours become long, narrow ellipses. Gradient descent then zig-zags along the steep axis and converges slowly.
Two common scaling methods
-
Min-max scaling (normalisation) — rescales each feature to :
-
Standardisation — zero mean, unit variance:
Choosing between them
| Property | Min-max | Standardisation |
|---|---|---|
| Bounded range | Yes (useful for NNs expecting ) | No |
| Robust to outliers | No (crushes normal values if an outlier dominates) | More robust |
| Affected by feature with outlier ? | Maps range to | Barely affected |
NoteScaling is itself a hyperparameter — there's no universal best method; it depends on the data and algorithm.
Feature scaling is not needed for the normal equation, but strongly recommended for gradient descent.
Overfitting and Underfitting
We want the model to do well on unseen data, not just on the training set — this is the generalisation goal.
| Regime | Training error | Test error | Cause |
|---|---|---|---|
| Underfitting | High | High | Model too simple to capture the pattern. |
| Ideal fit | Low | Low | Complexity matches data-generating process. |
| Overfitting | Very low | High | Model memorised noise; doesn't generalise. |
Model complexity can be controlled by changing the number of terms in the regression function (e.g. the degree of a polynomial, or the number of features used).
Regularization
Any change we apply to the learning algorithm aimed at decreasing generalization error without affecting training error (much).
Intuition: simpler models have smaller absolute weights — or even fewer non-zero weights. Adding a penalty on the size of pushes the optimiser toward simpler solutions that still explain the data, at the cost of slightly higher training error.
Regularised cost
where is the regularization strength (a hyperparameter).
Common penalties
| Name | Penalty | Effect on weights |
|---|---|---|
| L2 — Ridge | Shrinks all weights smoothly toward 0. | |
| L1 — Lasso | Drives some weights exactly to 0 — sparsity / feature selection. | |
| Elastic Net | Combines both. |
Why L1 gives sparsity and L2 doesn'tThe L1 penalty has "corners" at the axes, so the optimum often lands exactly on an axis — a coordinate equals zero. L2's smooth bowl always keeps weights slightly off-axis.
Don't regularise the interceptConventionally, the bias term is excluded from the penalty so that shifting by a constant doesn't get penalised.
Trade-off
There is a trade-off between fitting the training data well and keeping the model simple (lower norm). controls where we sit on this spectrum:
- → ordinary least squares, risk of overfitting.
- → all weights pushed to 0, risk of underfitting.
Ridge in closed form
Ridge still has a closed-form solution:
Adding also fixes invertibility issues when is singular.
Try it: step through the graph
type: bias-variance
Early Stopping
Another form of regularization: stop training before overfitting occurs.
Monitor training and validation loss over iterations:
- Training loss decreases monotonically.
- Validation loss decreases, reaches a minimum, then starts increasing — that's the onset of overfitting.
- Stop at the minimum of the validation loss and keep those weights.
Loss
│
│\ Validation loss
│ \ ╱
│ \_____╱
│ \____
│ ‾‾‾‾‾‾ Training loss
└───────────────────── Iterations
↑
stop here
Polynomial Regression
If a straight line can't capture the shape of the data, we can fit a polynomial while still using linear regression machinery — by treating powers of as new features:
We engineer new features , , …, and run ordinary linear regression on them. The model is linear in the parameters , even though it's nonlinear in .
Caveats
- Higher degree → higher capacity → higher overfitting risk. Use validation to pick .
- Powers blow up the scale of features — feature scaling is essential (e.g. ).
- For multivariate data you can also add interaction terms like .
Summary
- Linear regression fits a linear model by minimising mean squared error — a convex cost.
- Two ways to solve it: gradient descent (iterative, scalable) and the normal equation (closed form, exact but ).
- Feature scaling is crucial for gradient descent; different scales slow convergence.
- Stochastic / mini-batch GD trade smoothness for speed and can escape local minima (relevant for non-convex problems).
- Control overfitting with regularization (Ridge, Lasso, Elastic Net) or early stopping.
- Polynomial regression is linear regression on engineered polynomial features — linear in parameters, not in .
Related notes
- Machine Learning — subject hub
- Gradient Descent
- Cost Function
- Regularization
- Overfitting
- Feature Scaling
- Normal Equation