Mock Paper D (practice) · worked-solution

Mock Paper D — Deployment, Governance and Judgement

  • #machine-learning
  • #past-paper
  • #mock-exam
  • #loan-pipeline
  • #model-selection
  • #deployment
What this paper is

A practice mock in the format of the lecturer's sample exam: 25 multiple-choice questions on the same deliberately-flawed notebook, loan_pipeline.ipynb, in the same three-part structure. It is not an official paper and carries no answer key from the course.

The territory this paper covers

The sample paper spends most of its length on the notebook itself — which columns leak, where the split goes wrong, why the scaler is fitted too early. Those defects are settled by the time this paper starts. Mock D picks up at the point where someone says "fine, we've fixed it — now what?"

Theme Questions
Cost matrices, thresholds and expected-loss arithmetic 1, 11, 22
The regulator, explainability and adverse-action reasons 12, 23
Fairness and stability across customer segments 4, 15, 24
Monitoring, drift, and what a metric can tell you in production 2, 6, 14
Feedback loops from a model that filters its own training data 3, 13
Shadow deployment, champion/challenger, controlled experiments 9, 17, 20
Governance: documentation, retraining cadence, kill criteria, escalation 7, 8, 16, 18, 19
What go/no-go actually means, and reading confidence sceptically 5, 10, 21, 25

How this mock differs from the sample paper

The sample paper has a tell: its correct option is almost always the longest and most hedged one, so a student who has not read the code can guess a respectable score. This paper is built so that does not work.

  • Options within a question are matched for length, so you cannot pick by word count.
  • Cautious phrasing ("in part", "for now", "materially") appears in wrong options as often as right ones.
  • The willingness to admit a cost — "accept the losses", "accept the lower score" — is spread across correct and incorrect options deliberately.
  • Some correct answers are the shortest and bluntest option on the page. Some refuse to act. Some are the most specific. The shape varies on purpose.
  • Every distractor is something a competent analyst might genuinely propose. Several are correct in almost every respect and fail on one clause.

If you find yourself narrowing to two options and choosing on tone, you have found a question you do not actually know. Mark it and come back.

How to use it

Work through it closed-book on the notebook first, then check. Count your mistakes by part:

  • Part 1 (1–10) — losing marks here means you are agreeing or disagreeing too readily. Three of these ten concerns are not valid, and two more are true but secondary.
  • Part 2 (11–20) — losing marks here means you are picking the plan that sounds most thorough rather than the one that addresses the mechanism.
  • Part 3 (21–25) — losing marks here means you have not fixed your criteria before reading the numbers.

The last question is the one to reread afterwards. Everything else on this paper is a clause inside its answer.

  1. Q1 — Who chose the cutoff?

    A member of your team raises the following concern about the pipeline:

    "The cutoff at which an application gets rejected was never chosen by anyone — it is whatever the library happened to default to."

    Reviewing the code and its outputs yourself — is this concern valid?

  2. Q2 — The comparison chart

    A member of your team raises the following concern about the pipeline:

    "The comparison chart makes the difference between the two models look far larger than it actually is."

    Reviewing the code and its outputs yourself — is this concern valid?

  3. Q3 — What next year's data will contain

    A member of your team raises the following concern about the pipeline:

    "Once this model is live, next year's training data will consist only of applications it chose to approve."

    Reviewing the code and its outputs yourself — is this concern valid?

  4. Q4 — How the system will treat the self-employed

    A member of your team raises the following concern about the pipeline:

    "Whatever we deploy will behave differently for self-employed applicants than for salaried ones."

    Reviewing the code and its outputs yourself — is this concern valid?

  5. Q5 — 'Ninety percent is the standard bar'

    A member of your team raises the following concern about the pipeline:

    "We should not deploy anything until the model reaches at least 90% accuracy, which is the standard bar for a production system."

    Reviewing the code and its outputs yourself — is this concern valid?

  6. Q6 — Monitoring accuracy month by month

    A member of your team raises the following concern about the pipeline:

    "Tracking the model's accuracy month by month in production will tell us when it has gone stale."

    Reviewing the code and its outputs yourself — is this concern valid?

  7. Q7 — When the branch officer disagrees

    A member of your team raises the following concern about the pipeline:

    "Nobody has said what a branch officer should do when the model rejects an applicant they can see is creditworthy."

    Reviewing the code and its outputs yourself — is this concern valid?

  8. Q8 — No record of what produced the numbers

    A member of your team raises the following concern about the pipeline:

    "There is no record of which data and which version of the code produced the numbers we are being asked to approve."

    Reviewing the code and its outputs yourself — is this concern valid?

  9. Q9 — Switching everything over on day one

    A member of your team raises the following concern about the pipeline:

    "Switching every application over to the model on day one is the fastest way to find out whether it works."

    Reviewing the code and its outputs yourself — is this concern valid?

  10. Q10 — A number reported without uncertainty

    A member of your team raises the following concern about the pipeline:

    "The 98% figure is reported to two significant figures with no sense of how much it would move under a different split."

    Reviewing the code and its outputs yourself — is this concern valid?

  11. Q11 — Plan: deriving the threshold from the costs

    The following concern is real and confirmed. The data team proposes four ways forward:

    "The cutoff sits at 0.5 although the two kinds of error cost the bank roughly ten times different amounts."

    Which plan do you approve?

  12. Q12 — Plan: giving the regulator a reason

    The following concern is real and confirmed. The data team proposes four ways forward:

    "The regulator requires a plain-terms reason for each rejection, and the pipeline outputs only a number between zero and one."

    Which plan do you approve?

  13. Q13 — Plan: breaking the feedback loop

    The following concern is real and confirmed. The data team proposes four ways forward:

    "Once live, the model will only see outcomes for applications it approved, so its blind spots will harden with every retrain."

    Which plan do you approve?

  14. Q14 — Plan: noticing when the population moves

    The following concern is real and confirmed. The data team proposes four ways forward:

    "Nobody would notice if the applicant population shifted away from the population the model was trained on."

    Which plan do you approve?

  15. Q15 — Plan: a segment the model handles badly

    The following concern is real and confirmed. The data team proposes four ways forward:

    "The corrected model performs noticeably worse for self-employed applicants than for salaried ones."

    Which plan do you approve?

  16. Q16 — Plan: how often to retrain

    The following concern is real and confirmed. The data team proposes four ways forward:

    "Nobody has decided how often the model will be retrained, or what has to be true before a new version replaces the live one."

    Which plan do you approve?

  17. Q17 — Plan: how the model reaches live traffic

    The following concern is real and confirmed. The data team proposes four ways forward:

    "There is no plan for how the model would be introduced to live applications."

    Which plan do you approve?

  18. Q18 — Plan: what has to be written down

    The following concern is real and confirmed. The data team proposes four ways forward:

    "Nothing records what the model does, what data it was built on, or where it is known to be weak."

    Which plan do you approve?

  19. Q19 — Plan: how the rollout stops

    The following concern is real and confirmed. The data team proposes four ways forward:

    "The rollout plan describes how the model goes live but says nothing about how it would be stopped."

    Which plan do you approve?

  20. Q20 — Plan: measuring the business impact

    The following concern is real and confirmed. The data team proposes four ways forward:

    "The team plans to measure the model's value by comparing the default rate before and after it goes live."

    Which plan do you approve?

  21. Q21 — What go/no-go is a decision about

    The model decision:

    "The board asks you to bring them a single go/no-go recommendation on the data team's model."

    What should that decision actually be a decision about?

  22. Q22 — Putting numbers on the two errors

    The model decision:

    "Per 1,000 applications, about 140 would default under current rules. At the chosen threshold the model would catch 50 of those 140, and would also reject 200 applicants who would have repaid. A missed defaulter costs roughly 70,000 shekels; a wrongly rejected applicant costs roughly 7,000 in forgone margin."

    What does the arithmetic support?

  23. Q23 — 'You don't have the background to assess this'

    The model decision:

    "You press the team on the evaluation. They reply that you do not have the modelling background to assess it, and that the results speak for themselves."

    What is the defensible response?

  24. Q24 — One headline, two populations

    The model decision:

    "The corrected model catches 41% of defaulters among salaried applicants and 22% among the self-employed, who are about a quarter of the book. The headline figure quoted to the board is 36%."

    What follows?

  25. Q25 — The go/no-go call

    The model decision:

    "Everything is now in place: the leakage removed, the split by customer, the label windowed, the threshold derived from the costs. The Random Forest catches 36% of defaulters and the Neural Network 37%; the forest is explainable and cheap to run. Expected cost per application beats current underwriting. Selection bias on approved-only data remains, and cannot be fixed with the data the bank holds."

    What is the defensible call?