Mock Paper E (practice) · worked-solution

Mock Paper E — Concepts and Fundamentals

What this paper is

A practice mock in the format of the real exam — 25 multiple-choice questions on the same notebook, loan_pipeline.ipynb. It is not the lecturer's paper. It is built directly on the fifteen preparation questions in the problem brief, the ones the brief says "cover everything the exam covers".

The territory this paper covers

Where mocks A–D drill the pipeline, this one drills the concepts underneath it — the theory the exam assumes you already have, tested through the notebook rather than in the abstract.

Prep question Questions here
1 Overfitting and underfitting 1, 2
2 Bias 3
3 Accuracy and baselines 4, 5
4 Regression vs classification 6
5 Deep learning 7
6 Ensembles 8
7 Parameter vs hyperparameter 9
8 Train, validation and test 10, 11
9 Data leakage 12, 13
10 What a row represents 14, 15
11 Missing values 16, 17
12 Order of operations 18
13 Reading the data 19, 20
14 How the data was collected 21, 22
15 Choosing between two models 23, 24, 25

Every one of the fifteen appears. If you can answer all twenty-five, the brief's claim is that you are prepared.

How the options were written

The usual shortcuts will not work
  • The correct option is not the longest. Options are matched for length within each question, and several correct answers are the shortest on offer.
  • Distractors are not strawmen. Every wrong option is something a competent analyst might genuinely say — most of them are true about something, just not about what was asked.
  • "Is that right?" carries no signal. Six questions put a claim in a colleague's mouth. In four the right answer rejects it (1, 6, 10, 22); in two the right answer endorses it (7, 11). Answering "no" on reflex loses marks in both directions.
  • Several distractors are entirely true and entirely beside the point (10D, 18C, 22B, 24A). Recognising relevance is part of the mark.
  • Two distractors reach the right verdict by an invented mechanism (11D, 21C). Getting there for the wrong reason is not getting there.
  • Q3 offers four real problems with this pipeline and asks which one the word refers to. Being right about the notebook and wrong about the vocabulary scores zero.

The near-misses are the point. In Q4, knowing the 86% baseline gets you to option A — which is still wrong, because it concedes twelve points that do not exist. In Q16, spotting the ordering error gets you to D. In Q15, knowing that deduplication is the right idea gets you to C. Each of those is a candidate who has done the reading and stopped one step early.

Suggested use

Work the fifteen prep questions as short answer first — write them out, in your own words, before you look at anything here. Then sit this closed-book and mark yourself by block:

  • Q1–Q11 (concepts) — if you drop marks here, the problem is theory, and re-reading the notebook will not fix it.
  • Q12–Q20 (the concepts applied to this data) — drops here mean you know the definitions but have not read the code closely enough to spot them in the wild.
  • Q21–Q25 (judgement) — drops here are the expensive ones. This is the part of the exam that is actually about owning the decision.
These answers are derived, not official

No answer key exists for this notebook. Every answer here is grounded line by line in the code, the data generator and the brief, with the working shown so you can check it rather than trust it — which the lecturer says is precisely the habit the exam rewards.

  1. Q1 — 'Whatever else is wrong, it isn't overfitting'

    A colleague adds two lines to the notebook and prints the training score alongside the test score:

    text
    Random Forest  — train 0.9924 | test 0.9871
    Neural Network — train 0.9903 | test 0.9817
    

    "Half a point of gap on both models. Whatever else is wrong with this pipeline, overfitting isn't one of the problems."

    Is that conclusion sound?

  2. Q2 — Both numbers poor, no gap

    The pipeline is rebuilt properly: leaked columns dropped, split by customer, imputation fitted on training rows only. One configuration comes back with recall on defaulters of 0.11 on training and 0.10 on test.

    Which response does that pattern call for?

  3. Q3 — Three people say 'bias' and mean three things

    In the go/no-go meeting someone says:

    "The real risk here is bias."

    Three people in the room take that to mean three different things, and all three concerns are genuine.

    Which statement describes high bias in the statistical sense — the sense the term carries in the bias–variance trade-off?

  4. Q4 — 'Far better than a 50/50 coin flip'

    The notebook's closing cell reads:

    "Both models are around 98% - far better than a 50/50 coin flip."

    About 14% of loans in the data end in default. Which reading of that sentence is correct?

  5. Q5 — Asking for the right measures

    The go/no-go pack contains one number per model. You want measures that show how the system actually treats defaulters.

    Which request gets you what you need?

  6. Q6 — 'A sigmoid output makes it regression'

    A colleague reads the last two lines of the network cell:

    python
    layers.Dense(1, activation="sigmoid"),
    ...
    acc_nn = accuracy_score(y_test, (nn.predict(X_test, verbose=0) > 0.5).astype(int))
    

    "The output is a continuous number between 0 and 1, so this is a regression model. The > 0.5 is us converting a regression into a classification by hand."

    Is that right?

  7. Q7 — 'A network was never the right tool here'

    Setting the accuracy figures aside, a colleague makes a claim about the model class itself:

    "On seven tabular columns and 6,500 rows, a neural network was never the right tool for this — and no amount of tuning would have changed that."

    Is that right?

  8. Q8 — What makes a hundred trees different
    python
    rf = RandomForestClassifier(n_estimators=100, random_state=42)
    

    What makes the hundred trees differ from one another, and why does averaging them beat a single tree?

  9. Q9 — 'Model parameters: 3 layers, 30 epochs'

    Two markdown cells describe the settings:

    (model parameters: n_estimators=100, default settings are fine)

    (model parameters: 3 layers, 30 epochs, no tuning needed)

    Which classification of those settings is correct?

  10. Q10 — 'The test rows never go into fit'

    A colleague proposes:

    "Let's try n_estimators at 50, 100, 200 and 500 and keep whichever scores best on the test set. That isn't leakage — the test rows never go into fit."

    Is that right?

  11. Q11 — 'Nothing was tuned, so nothing was lost'

    One member of the team argues that the missing validation set is a technicality in this case:

    "Neither model was tuned. Each was fitted once, at one configuration. A validation set would have had nothing to decide."

    Another disagrees, and says the missing validation set is a real defect regardless. Is the second person right?

  12. Q12 — Applying the decision-day rule

    The general rule: a feature may be used only if the bank knows it about the applicant on the day they apply. X = df.drop(columns=["default", "customer_id"]) leaves seven columns.

    Applying the rule, which of them fail it?

  13. Q13 — Why one column is worth twelve points

    avg_days_late is generated as rng.uniform(2, 35) for defaulters and rng.exponential(2.5) for everyone else, then blurred with rng.normal(0, 2).

    What does that imply about the reported 98%?

  14. Q14 — What one row is

    df.shape prints (6557, 9), and load_data writes int(rng.integers(3, 9)) rows for each of 1,200 customers.

    What does one row represent, and what does that do to train_test_split?

  15. Q15 — Plan: the split that fixes it

    You are writing the remediation plan. Which instruction correctly fixes the grouping problem?

  16. Q16 — Two lines of cleaning
    python
    df["income"] = df["income"].fillna(df["income"].mean())
    df["credit_score"] = df["credit_score"].fillna(0)
    

    Which statement identifies what is wrong with these two lines most completely?

  17. Q17 — Plan: what replaces the two fillna lines

    What should the rebuilt pipeline do about missing income and missing credit_score?

  18. Q18 — Fit, then split
    python
    X = StandardScaler().fit_transform(X)
    ...
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
    

    What is wrong, what is the right order, and what is this class of mistake called?

  19. Q19 — The chart nobody acted on

    The histogram in cell 4 shows two clearly separated peaks, the right one roughly twelve times the left. What is the most likely cause?

  20. Q20 — Plan: fixing the income column

    You are writing the remediation instruction for income. Which one is right?

  21. Q21 — Only the loans that were approved

    The warehouse holds the loans the bank currently has on its books. Every applicant the old rules rejected is absent from the file.

    What is this problem called, and can it be fixed in code?

  22. Q22 — 'Both populations are just people who applied'

    A colleague accepts the selection point but argues it is not blocking:

    "In production the model only ever sees applicants. It was trained on approved loans and it will score new applicants — both are just 'people who applied'. The bias washes out."

    Is that right?

  23. Q23 — On what basis do you choose?

    The pipeline reports about 99% for the Random Forest and 98% for the neural network, then recommends the network "because deep learning is the more advanced technology".

    What is the right basis for choosing between two models?

  24. Q24 — A y-axis from 0.90 to 1.00
    python
    fig.update_yaxes(range=[0.90, 1.00])
    

    Why is the comparison chart misleading?

  25. Q25 — The go/no-go

    The decision is yours. Which position is defensible on what this pipeline actually shows?