Preparation Questions — Loan Pipeline (15 Questions)

  1. 1

    Overfitting and underfitting.

    What is overfitting, and what is underfitting? For each one, describe what you would see when you compare performance on the training data with performance on the test data, and give one way to fix it.

  2. 2

    Bias.

    What does it mean to say a model has high bias? How does this relate to how complex the model is, and how does it relate to underfitting?

  3. 3

    Accuracy and baselines.

    About 14% of the loans in our data end in default. What accuracy would you get by simply predicting "no default" for every customer? Given that, why is accuracy a weak way to judge this model, and which measures would show you how it actually handles defaulters?

  4. 4

    Regression vs classification.

    What is the difference between regression and classification? Give one banking example of each, and say which of the two the loan-approval problem is.

  5. 5

    Deep learning.

    What is deep learning, and when is a neural network the better choice? For a table of about 6,500 rows and 7 columns, would you expect a neural network or a tree-based model to do better, and why?

  6. 6

    Ensembles.

    What is an ensemble of models? Explain simply why combining many decision trees usually works better than a single tree, and what makes the individual trees in a Random Forest different from one another.

  7. 7

    Parameter vs hyperparameter.

    What is the difference between a parameter and a hyperparameter? Give two examples of each from loan_pipeline.ipynb, and say where hyperparameters should be chosen.

  8. 8

    Train, validation and test.

    What is each of the three sets — training, validation and test — for? What goes wrong if you choose your settings by testing them on the test set, and what goes wrong if there is no validation set at all, as in our pipeline?

  9. 9

    Data leakage.

    What is data leakage? Give the general definition, then name the columns in loans.csv that leak and explain exactly why each one cannot be used to decide on a new applicant.

  10. 10

    What a row represents.

    What does one row of loans.csv represent, and how many rows does a single customer have? Explain why splitting the data randomly by row is a problem here, and what kind of split solves it.

  11. 11

    Missing values.

    In the pipeline, missing income is filled with the mean of the whole table and missing credit_score is filled with 0. Give three separate problems with those two lines, and say what should be done instead.

  12. 12

    Order of operations.

    Why is it a mistake to fit a StandardScaler, or to compute an imputation value, before splitting the data? What is the correct order, and what is this kind of mistake called?

  13. 13

    Reading the data.

    The income histogram produced by the pipeline shows two separate peaks, one about twelve times the other. What is the most likely cause, and what should be done about it before any modelling continues?

  14. 14

    How the data was collected.

    Our training data contains only loans that the bank approved in the past. What is this problem called, why does it matter when the model goes into production, and can it be fixed by changing the code?

  15. 15

    Choosing between two models.

    The pipeline reports about 99% for the Random Forest and 98% for the neural network, then recommends the neural network because "deep learning is the more advanced technology". List the criteria you would actually use to choose between two models, and explain why the comparison chart — whose y-axis runs from 0.90 to 1.00 — is misleading.

What this is

The fifteen preparation questions from the lecturer's problem brief, worked in full. The brief says they "cover everything the exam covers; if you can answer all fifteen clearly, you are prepared."

They are short-answer, not multiple-choice — they cover the theory the 25-question sample exam assumes you already have. Work these first, then sit the sample exam.

📄 The problem brief + prep questions · 📓 loan_pipeline.ipynb · 🔍 Code walkthrough & defect catalogue

These answers are derived, not official

No answer key was distributed with the brief. Every answer here is grounded line-by-line in the notebook and the brief, and the working is shown in full so you can check it rather than trust it — which is precisely the habit the lecturer says the exam is designed to reward:

"read the answer critically rather than copying it … an answer you accepted without checking is exactly the habit the exam is designed to catch."

How these map to the exam

Prep question Sample-exam questions it feeds
1–2 Overfitting, bias background for Q6, Q18
3 Accuracy and baselines Q9, Q19, Q25
4 Regression vs classification background for Q24
5 Deep learning Q21, Q23, Q25
6 Ensembles Q7, Q23
7 Parameter vs hyperparameter Q8, Q18
8 Train/validation/test Q6, Q18
9 Data leakage Q1, Q7, Q12, Q13
10 What a row represents Q2, Q11, Q16
11 Missing values Q14
12 Order of operations Q13
13 Reading the data Q5, Q17
14 How data was collected Q4, Q20
15 Choosing between models Q21, Q22, Q24, Q25

Every one of the 25 exam questions traces back to at least one of these fifteen. If a prep question is shaky, the exam questions in its row are where you will lose marks.

A note on two small inconsistencies

Worth knowing so they do not throw you in the exam:

  • 99% vs 98%. Prep question 15 says the Random Forest scored "about 99%", while the notebook's closing cell says both models are "around 98%". Either way the point stands — the forest scored at least as high as the network, and was not recommended.
  • The label rule. Sample-exam questions 3 and 15 quote a label defined as "missed 3+ payments at some point". That phrasing appears in neither the brief nor the notebook, where default is generated from a latent risk score. Treat the no-fixed-window criticism as given by the exam rather than something you could derive from the code.