Applied Econometrics · Dr. Aluma Dembo

Exam Question Playbook — Any Experiment

The reusable skeleton for Dembo's applied exam: decode the experiment's variables and the standard questions (and their answers) fall out. Variable-triage table, question archetypes A–K, and a decision tree.

Part of: Econometrics Purpose: Dembo hands out the experiment/paper a few days before the exam. The experiment changes every sitting; the questions almost never do. This note is the reusable skeleton — decode the variables, and the questions (and their answers) fall out automatically. Derived from: PP_01-Emotions & Risky Choice (Practice Exam) · PP_02-Holdup Game & Bargaining (2023 Moed A) · PP_03-Time Preferences & Discounting (2023 Moed B) Key concepts: Linear Probability Model, Causal Diagram, Endogeneity, Omitted Variable Bias, Fixed Effects, Instrumental Variables, Difference-in-Differences, Probit Model, Marginal Effects


Big picture: the exam is always the same five moves

Every past paper runs the same arc. Whatever the costume (lotteries, bargaining, discounting), the questions test:

  1. What model is this? → read the outcome variable (binary ⇒ LPM/probit/logit; continuous ⇒ OLS).
  2. What does each coefficient mean? → sign, size, units, significance.
  3. Why is OLS wrong here? → an omitted common cause / endogenous regressor ⇒ OVB, drawn as a DAG.
  4. How do we fix it? → fixed effects (time-invariant confounder), IV (any endogenous regressor), or a natural experiment (DiD/RDD).
  5. Did the fix work / is it valid? → relevance vs validity, parallel trends, overid tests, etc.
The single habit that wins marks

Classify every variable before you read the questions. Outcome type and regressor type decide the model, the standard errors, and which "fix" the paper is steering toward. Do the triage table below first, on scrap paper, every time.


Step 0 — Variable triage (do this the moment you get the paper)

For every variable, fill in four cells. This single table pre-answers ~60% of the exam.

Variable Outcome or regressor? Type (binary / continuous) Randomised by researcher, or observed? Time-varying or fixed within unit?
(fill in) … … … …

What each column unlocks:

  • Binary outcome ⇒ every lm/feols/anova on it is an LPM; coefficients are in percentage points; you must use heteroskedasticity-robust SEs (LPM error is mechanically heteroskedastic). Continuous outcome ⇒ ordinary OLS.
  • Observed (not randomised) regressor ⇒ suspect endogeneity ⇒ the question is fishing for OVB / IV. Randomised/researcher-set regressor ⇒ clean, exogenous, no bias story.
  • Fixed within unit (a personality trait, a village, intrinsic preferences) ⇒ if it's the confounder and it's unobserved, the fix is fixed effects (it demeans away). Time-varying confounder ⇒ FE won't save you; you need IV or a design.
  • Panel structure (same unit over many rounds) ⇒ opens FE and requires clustered SEs (cluster at the unit level).

How triage maps to the three past papers — the same four-column table, filled in. Swap in the new paper's variable names and the answers fall out.

PP1 · Emotions & Risky Choice — panel: 243 subjects × 15 rounds

Variable Outcome or regressor? Type Randomised or observed? Time-varying or fixed within unit?
chose.B outcome binary observed time-varying
diff.E regressor continuous researcher-set (lottery design) time-varying (by round)
emotion, energy regressor continuous (1–7) observed (self-report) time-varying
celsius, sleep regressor (IV candidates) continuous observed (natural variation) time-varying
risk.preferences omitted confounder — unobserved fixed within unit

Reading: binary outcome ⇒ LPM (cluster SEs by subject — it's a panel). emotion/energy are observed and driven by the fixed, unobserved risk.preferences ⇒ OVB ⇒ fixed effects (demeans the time-invariant trait) or IV (celsius, sleep). The Wed-final exam shock ⇒ DiD.

PP2 · Holdup Game — cross-section: 103 Buyer–Seller pairs (one game each)

Variable Outcome or regressor? Type Randomised or observed? Time-varying or fixed within unit?
control / promises / threats regressor (dummies) categorical randomised (assigned) — (cross-section)
invest outcome (Stage 1) binary observed —
offer outcome (Stage 2) continuous (0–100) observed NA unless invest=1
accept outcome (Stage 3) binary observed NA unless invest=1

Reading: invest/accept binary ⇒ LPM / logit; offer continuous ⇒ OLS. Treatments randomised ⇒ clean (treatment coefficients are causal). But offer/accept exist only when invest=1 ⇒ endogenous selection, and the sequential stages make earlier choices endogenous regressors for later ones. One game per pair ⇒ no panel / no FE.

PP3 · Time Preferences — panel: 178 subjects × 15 rounds (2,670 obs)

Variable Outcome or regressor? Type Randomised or observed? Time-varying or fixed within unit?
choice outcome continuous observed time-varying
today.always, delay.always outcome binary observed time-varying
A.payment, A.delay regressor continuous researcher-set (each round) time-varying
income regressor continuous observed (2002 survey) fixed within unit
subject's fixed traits omitted confounder — unobserved fixed within unit

Reading: choice continuous ⇒ OLS; today.always binary ⇒ LPM (robust SEs). A.delay/A.payment researcher-set ⇒ exogenous (clean). income is an observed survey covariate ⇒ endogenous ⇒ IV. Panel + unobserved fixed traits ⇒ OVB ⇒ fixed effects + cluster by subject.


The question archetypes (with ready-to-adapt answers)

Each block: when it shows up → what she asks → how to answer → traps. Swap in the new paper's variable names.

A · "What model / how do you interpret this coefficient?"

Trigger: any regression table. Asks: interpret β^j\hat\beta_j, sign + magnitude + significance.

  • LPM (binary outcome): β^j\hat\beta_j = change in the probability of y=1y=1, in percentage points, for a one-unit rise in xjx_j (others held fixed). A dummy regressor reads "vs the omitted base group."
  • Continuous OLS: β^j\hat\beta_j = change in yy (in yy's units) per one-unit xjx_j.
  • Log forms (if a variable is logged): log-dep ⇒ coefficient ×100 = % change in yy; log-indep ⇒ /100 = effect of a 1% change in xx; log-log ⇒ elasticity.
  • Significance: compare ∣t∣|t| to 1.96 (5%) or read p<0.05p<0.05. State the threshold explicitly.

🔧 t=β^j/SE(β^j)t = \hat\beta_j / \text{SE}(\hat\beta_j), and a 95% CI is β^j±1.96 SE\hat\beta_j \pm 1.96\,\text{SE}.

Traps

Don't say "increases yy by β\beta" for a binary outcome — say "increases the probability by β\beta percentage points." Always name the base group for a dummy. State "holding the other regressors fixed."

B · "Run a joint test" — anova() / F-test

Trigger: two nested models, or "test whether these variables jointly matter." Asks: state H0H_0, conclude, say what it implies.

📝 H0:βa=βb=0H_0:\beta_a=\beta_b=0 vs H1:H_1: at least one ≠0\neq 0. Large FF ⇒ small pp ⇒ reject ⇒ the dropped variables jointly matter ⇒ leaving them out (the restricted model) pushes them into the error term, so the restricted model is under-specified / its error has a systematic component.

🔧 F=(RSSr−RSSu)/qRSSu/(n−k−1)F = \dfrac{(\text{RSS}_r - \text{RSS}_u)/q}{\text{RSS}_u/(n-k-1)}, qq = number of restrictions. See F-test.

C · "Draw the causal diagram" — DAG with a confounder

Trigger: "a colleague worries that…", "update the diagram". Asks: add the confounder and explain the backdoor path.

📝 Put the unobserved common cause CC with arrows into both a regressor XX and the outcome YY. That X←C→YX \leftarrow C \rightarrow Y structure is the backdoor path that biases β^X\hat\beta_X. Name it, draw arrows, say "CC is unobserved so it sits in the error."

graph TD
    X["X (regressor of interest)"] --> Y["Y (outcome)"]
    C["C (unobserved confounder)"] --> X
    C --> Y
    u["u (error)"] --> Y
    class X,Y,C internal-link;
In the exam you draw this by hand — practise the X←C→YX \leftarrow C \rightarrow Y fork until it's muscle memory. Sequential/staged games (like the holdup game) are a special case: earlier choices cause later ones, so the stage order is the DAG.

D · "What does this mean for the OLS estimate?" — OVB / endogeneity

Trigger: follows the DAG question. Asks: is β^\hat\beta causal? Which direction is the bias?

📝 Because CC is omitted, Cov(X,u)≠0\text{Cov}(X,u)\neq 0 ⇒ endogeneity ⇒ β^X\hat\beta_X is biased and not causal; it conflates X→YX\to Y with CC's effect. Sign of bias = sign(effect of CC on YY) × sign(corr of CC with XX): both same sign ⇒ upward bias; opposite ⇒ downward. See Omitted Variable Bias.

E · "Why fixed effects here?" — panel confounder fix

Trigger: panel data + a time-invariant unobserved confounder (a trait, a village, a firm). Asks: why FE beats OLS, interpret FE coefficients.

📝 An individual fixed effect aia_i is a per-unit intercept that absorbs everything constant within the unit — including the unobserved confounder. The within estimator demeans each unit, so anything fixed within it (Ci−Cˉi=0C_i - \bar C_i = 0) is annihilated. Identification now comes from within variation ("how the same unit changes round to round"). Coefficients are interpreted relative to the unit's own average.

Traps

Put the FE at the level where the confounder is constant (subject | ID, not day). FE only kills time-invariant confounders — a time-varying one survives. Cluster SEs at the FE level. FE can't estimate the effect of a variable that never changes within a unit (it's collinear with aia_i).

F · "Interpret the probit/logit" — and the LPM-vs-probit contrast

Trigger: glm(... family = binomial). Asks: interpret coefficients; contrast with LPM.

📝 In probit/logit you can read sign and significance only — the coefficient is not the marginal effect. The effect on the probability is non-constant: βj ϕ(x′β)\beta_j\,\phi(\mathbf{x}'\beta) (probit), steep near P=0.5P=0.5, flat at the tails. To get a magnitude you evaluate ϕ(⋅)\phi(\cdot) at a chosen point (e.g. means). Estimated by MLE, not OLS.

Easy-marks contrast

LPM coefficient = the marginal effect (constant, in pp) but can predict P∉[0,1]P\notin[0,1]. Probit/logit is bounded in [0,1][0,1] but the coefficient ≠ marginal effect. That trade-off is a recurring question.

G · "Propose / evaluate an instrument" — IV & 2SLS

Trigger: an endogenous regressor (observed, correlated with uu). Asks: write the IV, judge relevance and validity.

📝 You need ≥ 1 instrument per endogenous regressor. Two conditions:

  • Relevance: Cov(z,X)≠0\text{Cov}(z,X)\neq 0 — testable (first-stage correlation / F-stat; F < 10 ⇒ weak).
  • Validity / exclusion: Cov(z,u)=0\text{Cov}(z,u)=0 — not testable from data (u is unobserved); you must argue it (z affects Y only through X). A correlation table can show relevance but never proves validity.

🔧 fixest: feols(y ~ exog | endog ~ z1 + z2, data, se = "hetero"). First stage: regress each endogenous XX on all instruments + exogenous regressors; second stage uses fitted X^\hat X.

Extra IV diagnostics she likes

Sargan/overid test (only when #instruments > #endogenous): H0H_0 = instruments valid; reject ⇒ at least one invalid. Wu-Hausman: H0H_0 = regressor exogenous (OLS fine); reject ⇒ endogeneity confirmed, use IV.

H · "Find the natural experiment" — Difference-in-Differences

Trigger: a shock/event hitting some units but not others, with before/after data. Asks: define groups & periods, compute DiD, state the assumption.

📝 Treatment = units hit by the shock; Control = units not hit. Pre = before; Post = after. The estimate nets out the common trend:

🔧 δ^=(YˉT,post−YˉT,pre)−(YˉC,post−YˉC,pre)\hat\delta = (\bar Y_{T,post}-\bar Y_{T,pre}) - (\bar Y_{C,post}-\bar Y_{C,pre}) — the difference of the two changes. As a regression (gives a standard error): Y=β0+δ0 post+β1 treat+δ1(post×treat)+uY = \beta_0 + \delta_0\,\text{post} + \beta_1\,\text{treat} + \delta_1(\text{post}\times\text{treat}) + u, where the interaction δ1\delta_1 = the DiD estimate.

Identifying assumption

Parallel trends: absent the shock, both groups would have moved by the same amount (untestable directly). Also check no other shock hit either group at the same time, and groups are otherwise comparable. The estimand is the ATT.

I · "Treatment switches at a threshold" — RDD (less common, be ready)

Trigger: treatment determined by a cutoff in a running variable (age, score, vote share). Asks: sharp vs fuzzy, the jump, bandwidth.

📝 Sharp: treatment goes 0→1 cleanly at the cutoff; effect = jump in YY at the cutoff (β2\beta_2 on the treatment dummy). Fuzzy: crossing only shifts the probability ⇒ estimate by IV (Wald estimator = jump in Y / jump in treatment prob). Continuity: absent treatment YY would pass smoothly through the cutoff. Checks: McCrary (no sorting), placebo (covariates don't jump). Smaller bandwidth = less bias, more variance.

J · "Outcome only observed for some units" — sample selection

Trigger: a variable is NA unless an earlier binary outcome = 1 (e.g. offer exists only if invest=1). Asks: is the sub-sample regression biased?

📝 If who is in the sample depends on unobservables that also affect the outcome, that's endogenous selection ⇒ sample-selection bias — the conditional regression isn't causal. Conceptually fixed by a Heckman two-step (inverse Mills ratio in stage 2; significant λ ⇒ selection confirmed).

K · "Which standard errors?" — the SE question

📝 LPM (binary outcome) ⇒ always robust (error variance depends on XX). Panel / repeated obs per unit ⇒ clustered at the unit level (serial correlation within unit). Time series ⇒ HAC/Newey-West. Default homoskedastic SEs are almost never right in this course.


Decision tree — from the new paper to the answer

graph TD
    A["Pick a regression in the paper"] --> B{"Outcome binary?"}
	    B -->|Yes| C["LPM (robust SEs) or probit/logit -> coef = pp / sign-only"]
    B -->|No| D["OLS -> coef in y's units"]
    C --> E{"Regressor of interest randomised?"}
    D --> E
    E -->|Yes| F["Exogenous - clean causal read"]
    E -->|No| G{"Confounder time-invariant within unit?"}
    G -->|Yes, and panel| H["Fixed Effects (+ clustered SEs)"]
    G -->|No / not panel| I{"Got a valid instrument?"}
    I -->|Yes| J["IV / 2SLS - argue relevance + validity"]
    I -->|No| K{"Is there a shock or a cutoff?"}
    K -->|Shock + before/after| L["Difference-in-Differences (parallel trends)"]
    K -->|Threshold cutoff| M["RDD (sharp/fuzzy, continuity)"]
    class A,C,D,F,H,J,L,M internal-link;

Universal mark-grabbers (the grader's checklist)

  • Always write the error term uitu_{it} and use the right subscripts (ii for unit, tt for round) — panels lose a mark for either omission.
  • State H0H_0 and H1H_1 explicitly before concluding any test.
  • Name the base/omitted group for every dummy.
  • For binary outcomes say "probability… percentage points," never just "yy increases."
  • Distinguish testable vs assumed: relevance/overid are testable; validity/parallel-trends/exclusion are assumptions you argue.
  • When you give a "fix," say what bias it removes and why (e.g. "FE demeans the time-invariant trait, breaking Cov(X,u)\text{Cov}(X,u)").
  • Round-trip every claim to the research question: end with "…so this answers RQ(x): yes/no, because…".

How to use this note before & during the exam

  1. When the paper drops: do Step 0 triage on every variable. Then walk the decision tree once per regression the paper shows.
  2. Pre-exam drill: for each archetype A–K, write the one-paragraph answer using the new paper's variable names. If you can fill all of them, you've pre-written the exam.
  3. Predict her questions: she almost always asks (i) interpret a coefficient, (ii) draw a confounder DAG, (iii) diagnose OVB/endogeneity, (iv) propose FE or IV, (v) judge the fix, (vi) find a DiD shock. Map each to a part of the new paper in advance.