Supervised learning & regression 15
- Supervised Learning:
Learning a mapping from labelled pairs. ML Lec 2
- Linear Regression:
An OLS model predicting a continuous target as a linear function of the features. DS Lec 4
- OLS:
Ordinary Least Squares — the estimator that minimises the sum of squared residuals. DS Lec 4
- Cost Function:
A scalar measure of model error (e.g. MSE) that training minimises. ML Lec 2
- Gradient Descent:
Iterative first-order optimisation that steps the parameters downhill on the Cost Function. ML Lec 2
- Normal Equation:
The closed-form OLS solution . ML Lec 2
- Polynomial Regression:
A linear model fit on engineered polynomial features. ML Lec 2
- Feature Scaling:
Putting features on a comparable scale (min-max scaling, standardisation). ML Lec 2
- Regularization:
Penalising model complexity — L1 (Lasso), L2 (Ridge), Elastic Net, early stopping. ML Lec 2
- Ridge Regression:
L2-regularised linear regression. DS Lec 4
- Lasso Regression:
L1-regularised regression that performs automatic Feature Selection. DS Lec 4
- Overfitting:
A large generalisation gap — the model fits noise; the heart of the bias-variance trade-off. ML Lec 2
- Bias-Variance Tradeoff:
The trade-off between underfitting (high bias) and overfitting (high variance). DS Lec 4
- RMSE:
Root Mean Squared Error — a regression error metric in the target's units. DS Lec 4
- R-Squared:
The coefficient of determination — the proportion of target variance the model explains. DS Lec 4
Classification 11
- Classification:
A supervised task mapping inputs to discrete labels (binary, multiclass, multilabel). ML Lec 5
- KNN:
K-Nearest Neighbours — a lazy classifier using a majority vote among the closest training points. ML Lec 5
- Decision Tree:
A flowchart classifier that splits on features via if/else rules. DS Lec 3
- Random Forest:
An ensemble of decision trees with bagging plus feature randomness. DS Lec 3
- Logistic Regression:
A probabilistic binary classifier using the Sigmoid Function. DS Lec 3
- Sigmoid Function:
Maps real values to , giving logistic regression its probability output. DS Lec 3
- Gini Impurity:
A node-purity measure used to choose decision-tree splits. DS Lec 3
- Ensemble Methods:
Combining multiple models to lower variance (e.g. bagging in a Random Forest). DS Lec 3
- Distance Function:
The similarity measure KNN relies on (Euclidean, Manhattan, Minkowski). ML Lec 5
- Euclidean Distance:
The standard straight-line distance used by K-Means and KNN. DS Lec 5
- One-Hot Encoding:
Encoding a categorical variable as a set of binary columns. DS Lec 3
Model selection & evaluation 14
- Model Selection:
Choosing the model type and tuning hyperparameters for a given task. ML Lec 5
- Train-Test Split:
Partitioning data into a training set and a held-out test set for honest evaluation. ML Lec 5
- Validation Set:
A third partition used to compare models and tune hyperparameters without touching the test set. ML Lec 5
- Hyperparameter Tuning:
Selecting hyperparameter values (e.g. ) via a validation-set search. ML Lec 5
- Cross-Validation:
K-fold resampling for unbiased model comparison and hyperparameter tuning. ML Lec 7
- Accuracy:
The fraction of correct predictions on the test set. ML Lec 5
- Confusion Matrix:
A table of TP, TN, FP (Type I) and FN (Type II) classification outcomes. ML Lec 7
- Precision:
— minimises false positives. ML Lec 7
- Recall:
— minimises false negatives; also called TPR / sensitivity. ML Lec 7
- Precision and Recall:
The paired evaluation metrics for imbalanced classification — see Precision and Recall. DS Lec 3
- ROC Curve:
TPR vs FPR across all thresholds — measures classifier separability. ML Lec 7
- AUC:
Area under the ROC curve — the probability the model ranks a random positive above a random negative. ML Lec 7
- Data Leakage:
Using test-set information during training, which invalidates the evaluation. ML Lec 7
- Scikit-Learn Pipeline:
Chaining preprocessing and model into one object to prevent Data Leakage. DS Lec 4
Unsupervised learning & dimensionality reduction 10
- Unsupervised Learning:
Finding structure in unlabelled data. DS Lec 5
- K-Means:
A centroid-based clustering algorithm. DS Lec 5
- Elbow Method:
Plotting within-cluster sum of squares vs to pick the number of clusters. DS Lec 5
- Silhouette Score:
Measures how well each point fits its assigned cluster. DS Lec 5
- PCA:
Principal Component Analysis — unsupervised feature extraction via a variance-maximising rotation. ML Lec 7
- Explained Variance:
The proportion of total variance captured by the PCA components. DS Lec 5
- Curse of Dimensionality:
Exponential data sparsity (and degrading distances) in high-dimensional spaces, motivating dimensionality reduction. ML Lec 7
- Customer Segmentation:
A business application of clustering to group customers. DS Lec 5
- Feature Selection:
Choosing a subset of the original features (wrappers and filters). ML Lec 7
- Pearson Correlation:
A linear-dependency measure , used as a filter criterion. ML Lec 7
Exploratory data analysis & data wrangling 13
- Exploratory Data Analysis:
The workflow of summarising, visualising, and sanity-checking a dataset before modelling. DS Lec 2
- Pandas:
The Python library for tabular data manipulation — the DataFrame toolkit. DS Lec 2
- DataFrame:
Pandas' 2-D labelled table (rows × named columns), the core EDA data structure. DS Lec 2
- Groupby:
The split-apply-combine pattern: split rows by a key, apply an aggregation, combine the results. DS Lec 2
- Summary Statistics:
Measures of central tendency (mean, median, mode) and spread (standard deviation, quartiles) describing a column. DS Lec 2
- Distributions:
The shape of a variable's values (symmetric, skewed), read from a histogram. DS Lec 2
- Histograms:
A chart that bins a numeric variable to show its distribution. DS Lec 2
- Boxplots:
A chart showing the median, quartiles, and outliers via the IQR. DS Lec 2
- Missing Values:
Absent entries (
NaN) that must be detected and handled (drop or impute) before modelling. DS Lec 2 - Outliers:
Extreme values — flagged by the IQR rule — that can distort summaries and models. DS Lec 2
- Matplotlib:
The foundational Python plotting library. DS Lec 2
- Seaborn:
A higher-level statistical plotting library built on Matplotlib. DS Lec 2
- Correlation vs Causation:
The caution that a statistical association does not by itself imply a causal effect. DS Lec 2