HW-02
- 1
The 10-Minute ML Explainer. Work with a partner to explain one specific Machine Learning term, showing that you understand the theory and can explain it with a real-world example. Prepare a 10-minute presentation (5 minutes per person), divided into two parts (or focus on one):
- The Theory — What is the concept and why is it a problem?
- The Example — Show a real scenario where this happens and how to fix it.
Format: live presentation. Choose a topic that interests you both (see the suggested list in the body below).
Assignment typeOpen-ended presentation, not a worked problem set. There is no marked solution — the deliverable is a live 10-minute explainer.
How you'll be graded
- Simplicity — could another student understand this easily?
- Accuracy — is the technical explanation correct?
- Visuals — clear drawings or charts instead of long sentences?
- Teamwork — did both partners speak and connect their parts?
Check for understanding (before you start)
- The "So What?" test — if this ML problem happens, what is the actual damage to the project?
- The "Analogy" test — can you explain the concept to a non-engineer with a simple analogy? (e.g. overfitting is like a student memorising a practice test but failing the real exam.)
Suggested topics
1. Training dynamics & model behaviour ("the fit")
- Overfitting — memorising the noise vs. learning the pattern
- Underfitting — why a model might be too simple for the data
- The bias-variance tradeoff — balancing flexibility with stability
- Early stopping — finding the right moment to stop training
- Regularization (L1/L2) — a penalty that stops the model over-complicating
- Data leakage — when the model "cheats" by seeing test answers early
- Hyperparameter tuning — the difference between learning and configuring a model
2. Niche metrics (beyond accuracy)
- Precision vs. recall — the cost of a false alarm vs. a missed event
- F1-score — the best metric when data is unbalanced
- Confusion matrix — a visual map of where the model gets confused
- Log loss — how confident (and how wrong) a prediction is
- Inference latency — when a slow model is a useless model
- False positives vs. false negatives — consequences in medicine vs. security
3. MLOps & real-world operations
- Model drift — the world changes but your model doesn't
- Data drift — incoming data looks different from training data
- Training-serving skew — model behaves differently in the lab vs. production
- Model monitoring — an alarm system for AI
- A/B testing for AI — comparing two models on real users safely
- Feedback loops — a model's own mistakes pollute its future training data
- Feature stores — a consistent data library for all your models
4. Data challenges & handling
- Imbalanced data — the interesting event is 0.1% of the data
- Data augmentation — artificial variety to make the model more robust
- Outlier detection — which points are "special" and which are "errors"
- Label noise — training when human-labelled data is messy
- Cross-validation — proving success isn't just a lucky accident
5. Advanced niche concepts (simplified)
- Transfer learning — take a pre-trained brain and give it a new job
- Explainable AI (XAI) — asking the black box to show its work
- Quantization — shrinking a model to fit on a small device
- Retrieval-augmented generation (RAG) — a knowledge base to prevent hallucinations
- Adversarial attacks — small, invisible changes that trick a model