Mock Paper A — The Data and How It Was Collected
- #machine-learning
- #past-paper
- #mock-exam
- #loan-pipeline
- #data-leakage
- #selection-bias
What this paper isA practice mock in the format of the real exam: 25 multiple-choice questions about the same deliberately-flawed notebook,
loan_pipeline.ipynb. It is not the lecturer's paper — it is an additional set, written to cover one territory in much more depth than a single paper can.
The territory this paper covers
Every question here is about the data and how it was collected — the half of the exam that is decided before a single model is fitted.
| Theme | Questions |
|---|---|
| What a row represents; the panel structure | 1, 7, 14, 15 |
| The label: definition, dating, observation window | 6, 12 |
| Missing values and why the pattern is informative | 3, 13 |
Mixed units in income, and the histogram that showed it |
2, 8, 18 |
| What is knowable on decision day; point-in-time correctness | 9, 16, 19, 21, 23 |
| Selection bias and the population the file describes | 4, 5, 22, 24, 25 |
| Exploratory work, data dictionaries, provenance | 10, 11, 17, 20 |
Scaler ordering, metric definitions, thresholds and model architecture appear only as distractors here — those belong to the lecturer's paper and to the other mocks.
How this mock differs from the sample paper
The sample paper is, if you read it closely, slightly guessable: the correct option is usually the longest, the most hedged, and the one that admits a cost. Those tells were removed here deliberately.
The shortcuts will not work
- Options within a question are matched for length. Counting words tells you nothing.
- Hedging language appears in wrong options as often as right ones.
- Several correct answers are blunt and short — "drop it", "there is none", "treat 1,200 as the sample size".
- Two questions in Part 1 have a "Not valid" answer, and one more is true but beside the point.
- Every distractor is something a competent analyst might genuinely propose. There are no strawmen to eliminate.
If you find yourself picking an option because of how it is written rather than what it claims, you are doing the thing this paper was built to punish. Go back to the code.
Suggested use
Attempt it closed-book after you have read the walkthrough once, and mark yourself by part:
- Part 1 (1–10) — did you read the code, or did you pattern-match on the concern sounding plausible?
- Part 2 (11–20) — did you fix the cause, or the nearest visible symptom?
- Part 3 (21–25) — did you scope the recommendation, or just answer yes or no?
Losing marks in one part and not the others tells you what to revise. Losing them evenly usually means going back to load_data and reading it line by line.
- Q1 — The column that describes the warehouse, not the applicant
A member of your team raises the following concern about the pipeline:
"The
monthcolumn ends up in the feature matrix, and it describes our filing system rather than the person applying."Reviewing the code and its outputs yourself — is this concern valid?
- Q2 — The histogram nobody read
A member of your team raises the following concern about the pipeline:
"The team plotted an income histogram in the very first cell, then moved straight on to cleaning. That plot was already telling them something."
Reviewing the code and its outputs yourself — is this concern valid?
- Q3 — Who is missing an income figure
A member of your team raises the following concern about the pipeline:
"Income is not missing evenly across applicants, and filling every gap with one average treats a specific group as if they were typical."
Reviewing the code and its outputs yourself — is this concern valid?
- Q4 — 'None of this data is real'
A member of your team raises the following concern about the pipeline:
"This file was written by a generator function in the same notebook that reads it. It is invented data, so none of our findings mean anything."
Reviewing the code and its outputs yourself — is this concern valid?
- Q5 — What the 14% default rate actually measures
A member of your team raises the following concern about the pipeline:
"The brief says about 14% of loans default. That is not the same thing as 14% of applicants being risky."
Reviewing the code and its outputs yourself — is this concern valid?
- Q6 — A label stamped on every snapshot
A member of your team raises the following concern about the pipeline:
"The
defaultflag is written identically onto every one of a customer's rows, including the earliest one."Reviewing the code and its outputs yourself — is this concern valid?
- Q7 — 'We dropped the customer identifier'
A member of your team raises the following concern about the pipeline:
"Removing
customer_idfrom the feature matrix means the repeated-customer structure has been dealt with."Reviewing the code and its outputs yourself — is this concern valid?
- Q8 — Where did `income` come from?
A member of your team raises the following concern about the pipeline:
"Nobody in this room can name the system that
incomewas extracted from, and the notebook does not record it anywhere."Reviewing the code and its outputs yourself — is this concern valid?
- Q9 — Values that look safe
A member of your team raises the following concern about the pipeline:
"
credit_scoreandloan_amountare fine on decision day, but we have not checked that the stored values are the ones that existed then."Reviewing the code and its outputs yourself — is this concern valid?
- Q10 — One plot and nothing else
A member of your team raises the following concern about the pipeline:
"The whole exploratory stage of this notebook is a single histogram. We went straight from loading the file to standardising it."
Reviewing the code and its outputs yourself — is this concern valid?
- Q11 — Plan: no column has a documented definition
The following concern is real and confirmed. The data team proposes four ways forward:
"For not one of the nine columns can we state its source system, its unit, and the date its value refers to."
Which plan do you approve?
- Q12 — Plan: the extract carries no dates at all
The following concern is real and confirmed. The data team proposes four ways forward:
"There is no origination date, no observation date and no extract date anywhere in the nine columns."
Which plan do you approve?
- Q13 — Plan: the gaps are not evenly spread
The following concern is real and confirmed. The data team proposes four ways forward:
"Self-employed applicants are missing an income figure six times as often as salaried ones, and our fix erased that pattern."
Which plan do you approve?
- Q14 — Plan: what to do with `month`
The following concern is real and confirmed. The data team proposes four ways forward:
"The
monthcolumn survived into the feature matrix and both models are using it."Which plan do you approve?
- Q15 — Plan: six thousand rows, twelve hundred customers
The following concern is real and confirmed. The data team proposes four ways forward:
"We have been describing this as a 6,557-row dataset, but there are only 1,200 customers behind those rows."
Which plan do you approve?
- Q16 — Plan: were these the application-day values?
The following concern is real and confirmed. The data team proposes four ways forward:
"We cannot show that the stored
credit_scoreis the value that existed when the application was assessed."Which plan do you approve?
- Q17 — Plan: what to check before modelling anything
The following concern is real and confirmed. The data team proposes four ways forward:
"We ran one histogram and then started fitting. We need an exploratory step that would actually have caught this."
Which plan do you approve?
- Q18 — Plan: telling annual figures from monthly ones
The following concern is real and confirmed. The data team proposes four ways forward:
"Roughly thirty per cent of income values are annual and the rest monthly, and we must decide which is which."
Which plan do you approve?
- Q19 — Plan: the amount in the file is the amount we lent
The following concern is real and confirmed. The data team proposes four ways forward:
"
loan_amountrecords what we granted. An applicant on submission day has asked for something, which may differ."Which plan do you approve?
- Q20 — Plan: how the next extract gets produced
The following concern is real and confirmed. The data team proposes four ways forward:
"
loans.csvwas written by one person and cannot be reproduced. Nobody can say what query produced it."Which plan do you approve?
- Q21 — The model decision: 'we just drop two columns'
The model decision:
"The team accepts the leakage finding and says the fix is a one-line change: drop
avg_days_lateandcollections_flag, re-run, and report the new number. They want to keep the same file otherwise."What do you require before this proceeds?
- Q22 — The model decision: evidence that the file matches production
The model decision:
"You ask the team what evidence they have that
loans.csvdescribes the population of applicants the model will actually face."Which answer would you accept?
- Q23 — The model decision: which columns can explain a rejection
The model decision:
"The regulator requires a plain-terms reason for any rejection. You go through the nine columns asking which of them could appear in such a reason."
What do you conclude?
- Q24 — The model decision: a half-corrected dataset
The model decision:
"The team returns having removed the leaky columns and split by customer. The income units are still unresolved and the sample is still approved loans only. They ask to go live."
What do you decide?
- Q25 — The model decision: who the model may be used on
The model decision:
"The corrected evaluation catches 36–37% of defaulters. The board asks whether the data supports deploying this at all."
What do you tell them?