Mock Paper C — The Two Models and Methodology
- #machine-learning
- #past-paper
- #mock-exam
- #loan-pipeline
- #model-selection
- #random-forest
What this paper isA practice paper, written in the format of the lecturer's sample exam and built on the same notebook,
loan_pipeline.ipynb. Same structure, same three parts, same 25 questions — different territory.
- 📓
loan_pipeline.ipynb— run it before attempting this- 📄 The problem brief + 15 preparation questions
- 📝 Code walkthrough & defect catalogue
What this paper covers
The sample exam concentrates on the data — where the columns come from, what leaks, how the rows are split, what "default" means. This paper concentrates on the two models and the method used to compare them.
| Theme | Questions |
|---|---|
| Trees: splitting, scale-invariance, Gini impurity, readability | 1, 2, 5, 10 |
| Ensembles: bagging, feature subsetting, tree count | 5, 9, 18 |
| Network: capacity against sample size, activations, the sigmoid output | 3, 4, 11, 20 |
| Hyperparameters and tuning method | 8, 12, 13, 18 |
| Baselines, reproducibility, and how much of a difference to believe | 6, 7, 17, 21, 25 |
| Calibration and the probability scale | 4, 15, 24 |
| Explainability and what makes a model auditable | 10, 19, 22 |
There is deliberately no question here about leakage, imputation order, or metric definitions — those are the sample paper's ground, and covering them again would just be practice at recognising questions you have already seen.
How this paper differs from the sample exam
The sample paper has a tell: its correct option is almost always the longest and most hedged one, which makes it partly guessable without reading the code. This paper is written so that trick does not work.
- Options within a question are matched for length.
- The correct answer is sometimes the shortest and bluntest on the page (Q20, Q23).
- Cautious hedging and cost-admitting clauses appear in wrong options as often as in right ones.
- Every distractor is a mistake a competent analyst might actually make — several are true statements attached to a wrong conclusion (Q1 B, Q5 C, Q23 D), and several reach the right decision by the wrong reasoning (Q16 C, Q20 D, Q25 C).
If you find yourself picking an option because of how it sounds rather than what it claims, that is the habit this paper is built to break.
Before you attempt it
Have three numbers in your head, because a lot of the reasoning hangs off them:
| Number | Where it comes from |
|---|---|
| 7 feature columns | nine columns, minus default and customer_id |
| ~5,200 training rows | 6,557 rows, 80% to training — but only ~1,200 independent customers |
| ~134,000 network weights | three 256-unit layers plus a single-unit output |
The third against the second is Q3, Q20 and Q24. The gap between 5,200 rows and 1,200 customers is why every split in this paper is a grouped split.
- Q1 — Was the forest handicapped by standardisation?
A member of your team raises the following concern about the pipeline:
"Both models were trained on the same standardised matrix, which makes the comparison between them unfair to the Random Forest."
Reviewing the code and its outputs yourself — is this concern valid?
- Q2 — Nobody looked at which features the forest used
A member of your team raises the following concern about the pipeline:
"Nothing in the notebook reports which features the forest actually relied on, even though scikit-learn computes that for free."
Reviewing the code and its outputs yourself — is this concern valid?
- Q3 — A network far larger than the dataset
A member of your team raises the following concern about the pipeline:
"The network is fitting far more weights than the training data has rows."
Reviewing the code and its outputs yourself — is this concern valid?
- Q4 — Does a sigmoid output give you probabilities?
A member of your team raises the following concern about the pipeline:
"The network's sigmoid output means its predictions are already probabilities the bank can price loans against."
Reviewing the code and its outputs yourself — is this concern valid?
- Q5 — Are the hundred trees all the same tree?
A member of your team raises the following concern about the pipeline:
"The forest's hundred trees are effectively a hundred copies of the same tree, since they are all fitted to the same table."
Reviewing the code and its outputs yourself — is this concern valid?
- Q6 — Can anyone else reproduce these results?
A member of your team raises the following concern about the pipeline:
"Nobody else can reproduce these numbers, because both models involve randomness."
Reviewing the code and its outputs yourself — is this concern valid?
- Q7 — Two complex models and no simple one
A member of your team raises the following concern about the pipeline:
"Neither model was ever compared against something simple."
Reviewing the code and its outputs yourself — is this concern valid?
- Q8 — One configuration per model, chosen once
A member of your team raises the following concern about the pipeline:
"Each model was given exactly one configuration, and no alternative was ever tried."
Reviewing the code and its outputs yourself — is this concern valid?
- Q9 — 'A forest cannot overfit'
A member of your team raises the following concern about the pipeline:
"The forest's score needs no scrutiny, because averaging a hundred trees means a Random Forest cannot overfit."
Reviewing the code and its outputs yourself — is this concern valid?
- Q10 — Is a Random Forest fully transparent?
A member of your team raises the following concern about the pipeline:
"Explainability is a solved problem here — the Random Forest is fully transparent, so you can read the rules straight out of it."
Reviewing the code and its outputs yourself — is this concern valid?
- Q11 — Plan: nobody can justify the activation functions
The following concern is real and confirmed. The data team proposes four ways forward:
"The hidden layers use ReLU and the output uses a sigmoid, and nobody on the team can say why."
Which plan do you approve?
- Q12 — Plan: no hyperparameter search was ever run
The following concern is real and confirmed. The data team proposes four ways forward:
"No search over hyperparameters was ever run for either model."
Which plan do you approve?
- Q13 — Plan: thirty epochs, fixed in advance
The following concern is real and confirmed. The data team proposes four ways forward:
"The network trains for exactly thirty epochs with nothing monitoring the run."
Which plan do you approve?
- Q14 — Plan: nothing shows what drives the forest
The following concern is real and confirmed. The data team proposes four ways forward:
"Nothing in the notebook shows which features are driving the forest's predictions."
Which plan do you approve?
- Q15 — Plan: the outputs are being read as probabilities
The following concern is real and confirmed. The data team proposes four ways forward:
"The models' outputs are being treated as genuine probabilities of default, with no check that they are."
Which plan do you approve?
- Q16 — Plan: the recommendation was made on model class
The following concern is real and confirmed. The data team proposes four ways forward:
"The model was chosen on what class of model it belongs to, not on any evidence about its behaviour."
Which plan do you approve?
- Q17 — Plan: is a one-point difference real?
The following concern is real and confirmed. The data team proposes four ways forward:
"Nobody knows whether the difference between the two models' scores is real or an accident of this particular split."
Which plan do you approve?
- Q18 — Plan: a hundred trees, chosen from nowhere
The following concern is real and confirmed. The data team proposes four ways forward:
"The tree count of one hundred was never justified against anything."
Which plan do you approve?
- Q19 — Plan: the regulator wants a reason per rejection
The following concern is real and confirmed. The data team proposes four ways forward:
"The regulator requires a plain-terms reason for each rejection, and neither model produces one today."
Which plan do you approve?
- Q20 — Plan: the team wants more layers
The following concern is real and confirmed. The data team proposes four ways forward:
"The network came in slightly behind the forest, and the team wants to add layers to close the gap."
Which plan do you approve?
- Q21 — What the rerun must hold fixed
The model decision:
"The pipeline has been rebuilt and both models are about to be compared again. What must be held constant for that comparison to mean anything?"
Which specification do you approve?
- Q22 — Reading the model versus explaining it
The model decision:
"The forest can be interrogated directly; the network's decisions can only be explained by fitting a separate method on top of it."
How much weight should that distinction carry?
- Q23 — 'It will scale as more data arrives'
The model decision:
"Someone argues for the neural network on the grounds that it will keep improving as the bank accumulates more lending data."
How should that argument be treated?
- Q24 — Two models, two different probability scales
The model decision:
"On the corrected pipeline, the forest's predicted probabilities cluster between 0.05 and 0.30, while the network's are pushed towards 0 and 1."
What follows from this?
- Q25 — The corrected comparison, with spreads
The model decision:
"The rebuilt evaluation over five customer-grouped splits: Random Forest catches 36% of defaulters, varying by about one point across splits. Neural Network catches 37%, varying by about four points."
What is the defensible call?