Mock Paper A (practice) · worked-solution

Mock Paper A — The Data and How It Was Collected

  • #machine-learning
  • #past-paper
  • #mock-exam
  • #loan-pipeline
  • #data-leakage
  • #selection-bias
What this paper is

A practice mock in the format of the real exam: 25 multiple-choice questions about the same deliberately-flawed notebook, loan_pipeline.ipynb. It is not the lecturer's paper — it is an additional set, written to cover one territory in much more depth than a single paper can.

The territory this paper covers

Every question here is about the data and how it was collected — the half of the exam that is decided before a single model is fitted.

Theme Questions
What a row represents; the panel structure 1, 7, 14, 15
The label: definition, dating, observation window 6, 12
Missing values and why the pattern is informative 3, 13
Mixed units in income, and the histogram that showed it 2, 8, 18
What is knowable on decision day; point-in-time correctness 9, 16, 19, 21, 23
Selection bias and the population the file describes 4, 5, 22, 24, 25
Exploratory work, data dictionaries, provenance 10, 11, 17, 20

Scaler ordering, metric definitions, thresholds and model architecture appear only as distractors here — those belong to the lecturer's paper and to the other mocks.

How this mock differs from the sample paper

The sample paper is, if you read it closely, slightly guessable: the correct option is usually the longest, the most hedged, and the one that admits a cost. Those tells were removed here deliberately.

The shortcuts will not work
  • Options within a question are matched for length. Counting words tells you nothing.
  • Hedging language appears in wrong options as often as right ones.
  • Several correct answers are blunt and short — "drop it", "there is none", "treat 1,200 as the sample size".
  • Two questions in Part 1 have a "Not valid" answer, and one more is true but beside the point.
  • Every distractor is something a competent analyst might genuinely propose. There are no strawmen to eliminate.

If you find yourself picking an option because of how it is written rather than what it claims, you are doing the thing this paper was built to punish. Go back to the code.

Suggested use

Attempt it closed-book after you have read the walkthrough once, and mark yourself by part:

  • Part 1 (1–10) — did you read the code, or did you pattern-match on the concern sounding plausible?
  • Part 2 (11–20) — did you fix the cause, or the nearest visible symptom?
  • Part 3 (21–25) — did you scope the recommendation, or just answer yes or no?

Losing marks in one part and not the others tells you what to revise. Losing them evenly usually means going back to load_data and reading it line by line.

  1. Q1 — The column that describes the warehouse, not the applicant

    A member of your team raises the following concern about the pipeline:

    "The month column ends up in the feature matrix, and it describes our filing system rather than the person applying."

    Reviewing the code and its outputs yourself — is this concern valid?

  2. Q2 — The histogram nobody read

    A member of your team raises the following concern about the pipeline:

    "The team plotted an income histogram in the very first cell, then moved straight on to cleaning. That plot was already telling them something."

    Reviewing the code and its outputs yourself — is this concern valid?

  3. Q3 — Who is missing an income figure

    A member of your team raises the following concern about the pipeline:

    "Income is not missing evenly across applicants, and filling every gap with one average treats a specific group as if they were typical."

    Reviewing the code and its outputs yourself — is this concern valid?

  4. Q4 — 'None of this data is real'

    A member of your team raises the following concern about the pipeline:

    "This file was written by a generator function in the same notebook that reads it. It is invented data, so none of our findings mean anything."

    Reviewing the code and its outputs yourself — is this concern valid?

  5. Q5 — What the 14% default rate actually measures

    A member of your team raises the following concern about the pipeline:

    "The brief says about 14% of loans default. That is not the same thing as 14% of applicants being risky."

    Reviewing the code and its outputs yourself — is this concern valid?

  6. Q6 — A label stamped on every snapshot

    A member of your team raises the following concern about the pipeline:

    "The default flag is written identically onto every one of a customer's rows, including the earliest one."

    Reviewing the code and its outputs yourself — is this concern valid?

  7. Q7 — 'We dropped the customer identifier'

    A member of your team raises the following concern about the pipeline:

    "Removing customer_id from the feature matrix means the repeated-customer structure has been dealt with."

    Reviewing the code and its outputs yourself — is this concern valid?

  8. Q8 — Where did `income` come from?

    A member of your team raises the following concern about the pipeline:

    "Nobody in this room can name the system that income was extracted from, and the notebook does not record it anywhere."

    Reviewing the code and its outputs yourself — is this concern valid?

  9. Q9 — Values that look safe

    A member of your team raises the following concern about the pipeline:

    "credit_score and loan_amount are fine on decision day, but we have not checked that the stored values are the ones that existed then."

    Reviewing the code and its outputs yourself — is this concern valid?

  10. Q10 — One plot and nothing else

    A member of your team raises the following concern about the pipeline:

    "The whole exploratory stage of this notebook is a single histogram. We went straight from loading the file to standardising it."

    Reviewing the code and its outputs yourself — is this concern valid?

  11. Q11 — Plan: no column has a documented definition

    The following concern is real and confirmed. The data team proposes four ways forward:

    "For not one of the nine columns can we state its source system, its unit, and the date its value refers to."

    Which plan do you approve?

  12. Q12 — Plan: the extract carries no dates at all

    The following concern is real and confirmed. The data team proposes four ways forward:

    "There is no origination date, no observation date and no extract date anywhere in the nine columns."

    Which plan do you approve?

  13. Q13 — Plan: the gaps are not evenly spread

    The following concern is real and confirmed. The data team proposes four ways forward:

    "Self-employed applicants are missing an income figure six times as often as salaried ones, and our fix erased that pattern."

    Which plan do you approve?

  14. Q14 — Plan: what to do with `month`

    The following concern is real and confirmed. The data team proposes four ways forward:

    "The month column survived into the feature matrix and both models are using it."

    Which plan do you approve?

  15. Q15 — Plan: six thousand rows, twelve hundred customers

    The following concern is real and confirmed. The data team proposes four ways forward:

    "We have been describing this as a 6,557-row dataset, but there are only 1,200 customers behind those rows."

    Which plan do you approve?

  16. Q16 — Plan: were these the application-day values?

    The following concern is real and confirmed. The data team proposes four ways forward:

    "We cannot show that the stored credit_score is the value that existed when the application was assessed."

    Which plan do you approve?

  17. Q17 — Plan: what to check before modelling anything

    The following concern is real and confirmed. The data team proposes four ways forward:

    "We ran one histogram and then started fitting. We need an exploratory step that would actually have caught this."

    Which plan do you approve?

  18. Q18 — Plan: telling annual figures from monthly ones

    The following concern is real and confirmed. The data team proposes four ways forward:

    "Roughly thirty per cent of income values are annual and the rest monthly, and we must decide which is which."

    Which plan do you approve?

  19. Q19 — Plan: the amount in the file is the amount we lent

    The following concern is real and confirmed. The data team proposes four ways forward:

    "loan_amount records what we granted. An applicant on submission day has asked for something, which may differ."

    Which plan do you approve?

  20. Q20 — Plan: how the next extract gets produced

    The following concern is real and confirmed. The data team proposes four ways forward:

    "loans.csv was written by one person and cannot be reproduced. Nobody can say what query produced it."

    Which plan do you approve?

  21. Q21 — The model decision: 'we just drop two columns'

    The model decision:

    "The team accepts the leakage finding and says the fix is a one-line change: drop avg_days_late and collections_flag, re-run, and report the new number. They want to keep the same file otherwise."

    What do you require before this proceeds?

  22. Q22 — The model decision: evidence that the file matches production

    The model decision:

    "You ask the team what evidence they have that loans.csv describes the population of applicants the model will actually face."

    Which answer would you accept?

  23. Q23 — The model decision: which columns can explain a rejection

    The model decision:

    "The regulator requires a plain-terms reason for any rejection. You go through the nine columns asking which of them could appear in such a reason."

    What do you conclude?

  24. Q24 — The model decision: a half-corrected dataset

    The model decision:

    "The team returns having removed the leaky columns and split by customer. The income units are still unresolved and the sample is still approved loans only. They ask to go live."

    What do you decide?

  25. Q25 — The model decision: who the model may be used on

    The model decision:

    "The corrected evaluation catches 36–37% of defaulters. The board asks whether the data supports deploying this at all."

    What do you tell them?