Mock Paper B — Preprocessing and Evaluation
- #machine-learning
- #past-paper
- #mock-exam
- #loan-pipeline
- #data-leakage
- #class-imbalance
- #confusion-matrix
- #train-test-split
What this paper isA practice mock written to sit alongside the lecturer's sample exam. Same notebook, same three-part structure, same style of question — but a different slice of it. Where the sample paper roams across the whole pipeline, this one drills the preprocessing and evaluation half in depth.
- 📓
loan_pipeline.ipynb— run it before attempting this- 📄 Problem brief + 15 preparation questions
- 📝 Code walkthrough & defect catalogue
What this paper covers
| Territory | Questions |
|---|---|
| Order of operations — fit on train, transform on test | 1, 11 |
| Imputation, fill values, missing-indicators | 2, 8 |
| Standardisation: what it does and does not fix | 1, 2, 11 |
| Split design — grouping, stratification, validation sets | 3, 4, 14, 15 |
| Cross-validation and the spread around an estimate | 3, 15, 22 |
| Class imbalance, baselines, and what accuracy hides | 5, 9, 12, 16 |
| Confusion matrix, recall, precision, ROC-AUC, PR-AUC | 9, 13, 23, 24 |
| Thresholds and cost-sensitive evaluation | 10, 17, 21 |
| Overfitting diagnosed from train-versus-held-out gaps | 7, 18, 22 |
| The truncated-axis chart | 6, 19 |
| Feature provenance and the decision-day test | 20 |
The data-generating process, the model architectures, and deployment governance are touched only in passing — the sample paper and the other mocks cover those.
How to sit it
Three parts, same as the real thing:
- Part 1 (Q1–Q10) — a colleague raises a concern; decide whether it holds. Not every concern is valid, and some are true but trivial. Three of the ten here are not valid as stated.
- Part 2 (Q11–Q20) — the concern is confirmed; choose between four competing plans. The wrong plans are things a competent analyst might genuinely propose.
- Part 3 (Q21–Q25) — the numbers arrive and you have to act on them.
This paper is deliberately harder to guessThe sample paper can be half-answered by picking the longest, most hedged option every time. That tell has been removed here on purpose: options within a question are matched for length, cautious phrasing appears as often in wrong answers as in right ones, and several correct answers are the blunt short ones. You will have to read the code.
Where to look if you get one wrongQ1–Q2 and Q11 turn on the cleaning cell; Q3–Q4 and Q15 on the split; Q5, Q9, Q13 and Q24 on the evaluation. All of them are catalogued in the walkthrough — mark which cell your mistakes cluster around, because that is the cell to reread.
- Q1 — How much of the 98% is the scaler's fault?
A member of your team raises the following concern about the pipeline:
"The scaler was fitted before the split, and that is what produced the 98% accuracy."
Reviewing the code and its outputs yourself — is this concern valid?
- Q2 — What the zero-filled credit scores do after scaling
A member of your team raises the following concern about the pipeline:
"Filling missing credit scores with 0 does not just insert a wrong value — it distorts the whole column once it is standardised."
Reviewing the code and its outputs yourself — is this concern valid?
- Q3 — One split, one number
A member of your team raises the following concern about the pipeline:
"Everything we know about these models rests on a single random 80/20 split, and we have no idea how much that number would move."
Reviewing the code and its outputs yourself — is this concern valid?
- Q4 — No stratification on the label
A member of your team raises the following concern about the pipeline:
"The split is not stratified on the outcome, even though only about one customer in seven defaults."
Reviewing the code and its outputs yourself — is this concern valid?
- Q5 — 'Accuracy is never the right metric'
A member of your team raises the following concern about the pipeline:
"Accuracy is never an appropriate metric for a classification problem, so both reported figures should be discarded."
Reviewing the code and its outputs yourself — is this concern valid?
- Q6 — Reading the comparison chart
A member of your team raises the following concern about the pipeline:
"The bar chart is drawn so as to make the neural network look far better than the Random Forest."
Reviewing the code and its outputs yourself — is this concern valid?
- Q7 — No training score anywhere
A member of your team raises the following concern about the pipeline:
"We are shown a test score for each model and nothing else, so we cannot say whether either one is overfitting."
Reviewing the code and its outputs yourself — is this concern valid?
- Q8 — The mean of a two-humped column
A member of your team raises the following concern about the pipeline:
"The value used to fill missing incomes is the average of a column with two separate peaks in it."
Reviewing the code and its outputs yourself — is this concern valid?
- Q9 — Would a confusion matrix add anything?
A member of your team raises the following concern about the pipeline:
"Asking for a confusion matrix is bureaucracy — it just decomposes a number we already have."
Reviewing the code and its outputs yourself — is this concern valid?
- Q10 — Whose threshold is 0.5?
A member of your team raises the following concern about the pipeline:
"The neural network's outputs are cut at 0.5, which is an arbitrary choice — the Random Forest at least avoids that problem."
Reviewing the code and its outputs yourself — is this concern valid?
- Q11 — Plan: the scaler is still fitted on the test rows
The following concern is real and confirmed. The data team proposes four ways forward:
"We moved the split to the top of the notebook, but the code still calls
fit_transformon the test features as well as the training ones."Which plan do you approve?
- Q12 — Plan: what to do about the 14% base rate
The following concern is real and confirmed. The data team proposes four ways forward:
"About one customer in seven defaults, and nothing in the training procedure acknowledges that the classes are unbalanced."
Which plan do you approve?
- Q13 — Plan: what to report instead of one accuracy figure
The following concern is real and confirmed. The data team proposes four ways forward:
"The only evidence offered for either model is a single accuracy figure computed on one held-out set."
Which plan do you approve?
- Q14 — Plan: nothing decides when training stops
The following concern is real and confirmed. The data team proposes four ways forward:
"The network runs for exactly thirty epochs and nothing at all is watching it while it does."
Which plan do you approve?
- Q15 — Plan: putting an interval around the number
The following concern is real and confirmed. The data team proposes four ways forward:
"We are quoting a single figure to the board with no indication of how much it would move under a different draw."
Which plan do you approve?
- Q16 — Plan: what to compare the model against
The following concern is real and confirmed. The data team proposes four ways forward:
"The notebook compares its models to a coin flip, and nobody has established what the bank's current process would score."
Which plan do you approve?
- Q17 — Plan: where to put the threshold
The following concern is real and confirmed. The data team proposes four ways forward:
"Both models are being scored at a cut-off of one half, which nobody chose and which no cost calculation supports."
Which plan do you approve?
- Q18 — Plan: diagnosing overfitting once the data is clean
The following concern is real and confirmed. The data team proposes four ways forward:
"Once the leaky columns are gone and the split is by customer, we still have no way to tell whether either model is overfitting."
Which plan do you approve?
- Q19 — Plan: the comparison chart
The following concern is real and confirmed. The data team proposes four ways forward:
"The bar chart's vertical axis is set to run from 0.90 to 1.00, which makes a difference of about one point fill the frame."
Which plan do you approve?
- Q20 — Plan: the column nobody dropped
The following concern is real and confirmed. The data team proposes four ways forward:
"Only
defaultandcustomer_idare dropped, somonth— a property of the warehouse export — is being used as a predictor."Which plan do you approve?
- Q21 — The honest numbers come in below the baseline
The model decision:
"The rebuilt pipeline reports, for the Random Forest at the cost-based threshold: 36% of defaulters caught, 29% of flagged applicants genuinely defaulting, and overall accuracy of 71% — down from the 98% first presented, and below the 86% a do-nothing rule achieves."
How should this be read?
- Q22 — Validation recall above test recall
The model decision:
"After tuning, the chosen model reaches 41% recall on validation and 34% on the held-out test set."
What should be made of the seven-point difference?
- Q23 — Choosing between the models on AUC
The model decision:
"A colleague proposes settling the choice on ROC-AUC, on the grounds that it is threshold-free and therefore avoids all the arguments about where to cut."
How should this proposal be treated?
- Q24 — Reading the corrected confusion matrix
The model decision:
"On 1,300 held-out applications containing 182 eventual defaulters, the chosen model catches 62 of them, misses 120, and wrongly rejects 158 applicants who would have repaid."
Which reading of these figures is correct?
- Q25 — The go/no-go call
The model decision:
"Final position: recall 34% and precision 28% at the cost threshold, expected losses roughly 22% below the current rules on the same applications. The approved-only data problem remains unfixed."
What is the defensible call?