Applied Econometrics · Dr. Aluma Dembo

Exam Formula Sheet (Lec 1–10)

One-page-per-topic exam reference — every key term, formula, interpretation, and what-to-do tip from Lectures 1–10. Built for the closed-book final.

Part of: Econometrics · Glossary: _Econometrics Concepts Scope: the whole Sem-2 causal-inference toolkit. Each section = one lecture: key terms → formulas → how to read them → exam tips. Default everywhere: 5% significance, two-sided; reject if ∣t∣>1.96|t|>1.96 (or p<0.05p<0.05). Standard errors are robust/clustered unless told otherwise.


0. Method picker — "which tool does this question want?"

You see in the question… Reach for Lecture
Binary outcome y∈{0,1}y\in\{0,1\}, want simple/linear answer LPM (OLS + robust SE) L2
Binary outcome, want valid probabilities / non-constant effect Logit / Probit L3
Regressor correlated with the error (OVB, reverse causality, simultaneity) IV / 2SLS L4, L6
Outcome only observed for a self-selected subsample Heckman two-step L5
Price & quantity (supply/demand) jointly determined Simultaneous eqns + IV L6
Data over time, errors correlated across periods HAC / serial corr. fix L6
Both variables trend upward over time detrend / add tt L7
Effect of one dated event on an outcome Event study L7
Same units over time + unobserved fixed traits Fixed effects (demeaning) L8
Treatment switches at a threshold of some variable RDD L9
Treatment vs control, before vs after a policy DiD L10
First move on every applied question

(1) Is the outcome binary? (2) Is the key regressor exogenous, or is there a confounder / reverse causality / selection? (3) What's the data structure — cross-section, time series, or panel? Those three answers pin down the method.


1. Foundations recap (assumed from Intro Metrics)

Classical assumptions — A1–A5

A1 linear in parameters · A2 random sampling · A3 no perfect collinearity · A4 zero conditional mean E[u∣X]=0\mathbb{E}[u\mid X]=0 (the causal one) · A5 homoskedasticity Var(u∣X)=σ2\text{Var}(u\mid X)=\sigma^2. Under A1–A5, OLS is BLUE.

Course convention (varies by source): A3 = independent errors / A4 = E[u∣x]=0\mathbb E[u\mid x]=0 / A5 = homoskedastic. Read the question's numbering — the content is what matters.

t-test (single coefficient)

t^=β^j−βj0se(β^j) ∼ t n−k−1H0:βj=0⇒t^=β^jse(β^j)\hat t=\frac{\hat\beta_j-\beta_j^{0}}{\text{se}(\hat\beta_j)}\ \sim\ t_{\,n-k-1}\qquad H_0:\beta_j=0\Rightarrow \hat t=\frac{\hat\beta_j}{\text{se}(\hat\beta_j)}

Reject H0H_0 if ∣t^∣>tcrit≈1.96|\hat t|>t_{crit}\approx1.96 (large nn, 5%).

F-test (joint / nested models) — F-test

F=(RSSR−RSSUR)/qRSSUR/(n−k−1) ∼ Fq, n−k−1F=\frac{(\text{RSS}_R-\text{RSS}_{UR})/q}{\text{RSS}_{UR}/(n-k-1)}\ \sim\ F_{q,\,n-k-1}

qq = number of restrictions (regressors dropped). Large FF / small pp ⇒ reject that the dropped variables are jointly zero. In R: anova(unrestricted, restricted).

Confidence interval

β^j ± tcrit⋅se(β^j)(≈β^j±1.96 se)\hat\beta_j\ \pm\ t_{crit}\cdot\text{se}(\hat\beta_j)\quad(\approx\hat\beta_j\pm1.96\,\text{se})

Functional-form interpretation (quick table)

Model Coefficient β1\beta_1 means…
level–level y=β0+β1xy=\beta_0+\beta_1x Δx=1⇒Δy=β1\Delta x=1\Rightarrow \Delta y=\beta_1 (units)
log–level log⁡y=β0+β1x\log y=\beta_0+\beta_1x Δx=1⇒y\Delta x=1\Rightarrow y changes ≈100β1%\approx100\beta_1\%
level–log y=β0+β1log⁡xy=\beta_0+\beta_1\log x xx up 1% ⇒Δy≈β1/100\Rightarrow \Delta y\approx\beta_1/100
log–log β1\beta_1 = elasticity (% per %)

2. Lecture 1 — Causal inference & treatment effects

Goal: read β^1\hat\beta_1 as causal, not just correlational. Requires ruling out confounders, reverse causality, selection.

  • Data types: cross-section · time series · panel · repeated cross-section.
  • Exogenous (X) → endogenous (Y): the arrow is causal only if X varies independently of everything else affecting Y.
  • Continuous vs dummy regressor: with few discrete values, use dummies (one per level, minus a reference). Dummies are flexible (each level free); a continuous slope forces one straight line — test it with a joint F-test on the dummies.
  • Panel breaks A3/A4: repeated obs per unit ⇒ errors correlated within unit ⇒ classical SEs too small. Quick fix: subsample one obs/unit. Proper fix: fixed effects + clustered SEs (L8).
Exam tip

"Is this coefficient causal?" → name the specific confounder / reverse-causality / selection story, state Cov(x,u)≠0\text{Cov}(x,u)\neq0, then say which tool (control, FE, IV, RDD, DiD) shuts that path.


3. Lecture 2 — Linear Probability Model (LPM)

Binary y∈{0,1}y\in\{0,1\} run through OLS. Under A2, the fit is a probability:

E[y∣x]=P(y=1∣x)=β0+β1x\mathbb{E}[y\mid x]=P(y=1\mid x)=\beta_0+\beta_1x
  • Coefficient = marginal effect in percentage points (constant). β1=∂P(y=1)/∂x\beta_1=\partial P(y=1)/\partial x. Dummy coef = difference in P(y=1)P(y=1) vs base group.
  • Drawback 1 — unbounded: y^\hat y can fall below 0 or exceed 1 (nonsense probabilities). → motivates logit/probit.
  • Drawback 2 — built-in heteroskedasticity (always):
Var(u∣x)=p(x) [1−p(x)](max 0.25 at p=0.5)\text{Var}(u\mid x)=p(x)\,[1-p(x)]\quad(\text{max }0.25\text{ at }p=0.5)
  • Fix for drawback 2: robust (HC) standard errors — mandatory for any LPM. β^\hat\beta stays unbiased; only SEs were wrong. R: feols(y ~ x, se="hetero").
Exam tip

If the outcome is 0/1, every lm/feols is an LPM — interpret coefficients in pp and expect robust/clustered SEs. Saying "y^\hat y can leave [0,1]" + "errors are heteroskedastic by construction" earns the two standard marks.


4. Lecture 3 — Logit & Probit (discrete choice)

Wrap the linear index in a CDF GG so P∈(0,1)P\in(0,1):

P(y=1∣x)=G(x′β),x′β=β0+β1x1+⋯+βkxkP(y=1\mid\mathbf x)=G(\mathbf x'\boldsymbol\beta),\qquad \mathbf x'\boldsymbol\beta=\beta_0+\beta_1x_1+\dots+\beta_kx_k
GG (CDF) density gg R
Logit ex′β1+ex′β\dfrac{e^{\mathbf x'\beta}}{1+e^{\mathbf x'\beta}} dlogis glm(...,family=binomial(link="logit"))
Probit Φ(x′β)\Phi(\mathbf x'\beta) dnorm glm(...,family=binomial(link="probit"))

GG is increasing, S-shaped, → 0/1 at the tails, steepest at x′β=0\mathbf x'\beta=0 (P=0.5P=0.5), symmetric G(z)=1−G(−z)G(z)=1-G(-z).

Marginal effects (the key formula)

∂P(y=1)∂xj=g(x′β)⋅βj\frac{\partial P(y=1)}{\partial x_j}=g(\mathbf x'\boldsymbol\beta)\cdot\beta_j

Non-constant — depends where you sit on the curve. Evaluate at sample means g(xˉ′β^)⋅β^jg(\bar{\mathbf x}'\hat{\boldsymbol\beta})\cdot\hat\beta_j for "the average person."

  • Raw coef → sign & significance only, NOT magnitude. Never compare raw coef magnitudes across LPM/logit/probit; do compare marginal effects.
  • Estimation = MLE (model is non-linear, OLS invalid). Log-likelihood:
log⁡L(β)=∑i[yilog⁡G(xi′β)+(1−yi)log⁡(1−G(xi′β))]\log\mathcal L(\boldsymbol\beta)=\sum_i\big[y_i\log G(\mathbf x_i'\boldsymbol\beta)+(1-y_i)\log(1-G(\mathbf x_i'\boldsymbol\beta))\big]

Joint testing = Likelihood Ratio test (replaces F)

LR=2[log⁡LUR−log⁡LR] ∼ χq2LR=2\big[\log\mathcal L_{UR}-\log\mathcal L_{R}\big]\ \sim\ \chi^2_{q}

qq = # restrictions. Reject if LR>χcrit2(q)LR>\chi^2_{crit}(q). R: lrtest(ur, r).

Exam tip

Recipe for a marginal effect: (1) get β^\hat\beta; (2) compute xˉ′β^\bar{\mathbf x}'\hat\beta at the means; (3) multiply by g(⋅)g(\cdot) — dnorm (probit) / dlogis (logit). If only sign asked, just read the coefficient sign + pp/zz.


5. Lecture 4 — Instrumental Variables & 2SLS

Problem: endogeneity Cov(x,u)≠0\text{Cov}(x,u)\neq0 (OVB, reverse causality, simultaneity) ⇒ OLS biased & inconsistent (more data doesn't help).

Fix: an instrument zz that moves xx but not yy directly. Two conditions:

Condition Formula Testable?
Relevance Cov(z,x)≠0\text{Cov}(z,x)\neq0 ✅ first-stage F > 10
Validity (exclusion) Cov(z,u)=0\text{Cov}(z,u)=0 ❌ argue conceptually

2SLS

  • First stage: x=γ0+γ1z+(controls)+ε⇒x^x=\gamma_0+\gamma_1 z+(\text{controls})+\varepsilon\Rightarrow\hat x
  • Second stage: y=β0+β1x^+(controls)+uy=\beta_0+\beta_1\hat x+(\text{controls})+u
  • Simple IV estimator: β^1IV=Cov(y,z)Cov(x,z)\hat\beta_1^{IV}=\dfrac{\text{Cov}(y,z)}{\text{Cov}(x,z)}

R (correct SEs): feols(y ~ controls | x ~ z). Never run two manual lms (wrong second-stage SEs).

  • # instruments ≥ # endogenous regressors (need ≥2 instruments for 2 endogenous vars).
  • Weak instrument (F<10): tiny denominator amplifies bias — can be worse than OLS.
  • IV estimates LATE, not ATE — the effect for compliers (those whose xx moves with zz).
  • IV is less precise than OLS (larger SEs) — the price of consistency.
  • Sargan/overid test (only if overidentified): H0H_0 = all instruments valid. Reject ⇒ ≥1 invalid.
  • Wu–Hausman: H0H_0 = regressor exogenous (OLS fine). Reject ⇒ IV needed.
Exam tip

To defend an instrument, write both conditions explicitly with the covariance, show relevance via first-stage F, then give a one-sentence story for why zz can't reach yy except through xx — and a sentence on how it could fail. Validity can never be "proven from a table."


6. Lecture 5 — Sample selection & Heckman

Consistency vs unbiasedness

  • Unbiased: E[β^]=β\mathbb E[\hat\beta]=\beta (right on average, any nn).
  • Consistent: β^→β\hat\beta\to\beta as n→∞n\to\infty.
  • OLS with endogenous xx: biased AND inconsistent. 2SLS: biased in small samples but consistent.

Sample selection — you only observe yy when si=1s_i=1

Type Selection depends on OLS
Random nothing ✅ unbiased
Exogenous xx only (observed) ✅ unbiased
Endogenous uu (unobservables affecting yy) ❌ biased

Mincer running example

log⁡wage=β0+β1educ+β2exper+β3exper2+u\log\text{wage}=\beta_0+\beta_1\text{educ}+\beta_2\text{exper}+\beta_3\text{exper}^2+u

Heckman two-step ("Heckit")

  • Step 1 (full sample, probit): P(si=1∣z)=Φ(zi′γ)P(s_i=1\mid z)=\Phi(z_i'\gamma). Need an exclusion restriction — a variable in zz that affects selection but not the outcome (e.g. kids under 6 → labour-force participation, not the wage). Compute inverse Mills ratio:
λ^i=ϕ(zi′γ^)Φ(zi′γ^)\hat\lambda_i=\frac{\phi(z_i'\hat\gamma)}{\Phi(z_i'\hat\gamma)}
  • Step 2 (selected sample, OLS): add λ^i\hat\lambda_i as a regressor:
log⁡wagei=β0+β1educi+⋯+ρ λ^i+error\log\text{wage}_i=\beta_0+\beta_1\text{educ}_i+\dots+\rho\,\hat\lambda_i+\text{error}
  • Test for selection bias: H0:ρ=0H_0:\rho=0. Reject ⇒ endogenous selection present, correction matters.
Exam tip

Heckman mirrors 2SLS: step 1 builds a correction term, step 2 adds it as a control. Always state the exclusion-restriction variable and what testing ρ=0\rho=0 tells you.


7. Lecture 6 — Simultaneous equations & time series

Simultaneous equations (supply & demand)

Price & quantity jointly determined ⇒ Cov(P,u)≠0\text{Cov}(P,u)\neq0 in both equations ⇒ OLS biased. Fix: a shifter of one curve instruments price to trace the other.

  • Weather shifts supply → identifies the demand curve.
  • Day of week shifts demand → identifies the supply curve.

Serial correlation

Cov(ut,ut−1)≠0\text{Cov}(u_t,u_{t-1})\neq0. OLS stays unbiased & consistent but not efficient and SEs are wrong (too small) ⇒ invalid tt/FF.

  • Test: regress u^t=ρu^t−1+et\hat u_t=\rho\hat u_{t-1}+e_t; H0:ρ=0H_0:\rho=0.
  • Fix: HAC (Newey–West) SEs — coefficients unchanged, SEs widened. R: vcov="newey_west".

Time-series OLS assumptions

  • TS.2 Strict exogeneity E[ut∣X]=0\mathbb E[u_t\mid\mathbf X]=0 for all tt (past, present, future) → unbiasedness.
  • Contemporaneous exogeneity E[ut∣xt]=0\mathbb E[u_t\mid x_t]=0 → consistency only.
  • TS.3 homoskedastic · TS.4 no serial corr · TS.5 normal.

Model zoo

Model Equation Use
Static yt=β0+β1xt+uty_t=\beta_0+\beta_1x_t+u_t instant effect only
Distributed lag yt=α0+δ0xt+δ1xt−1+⋯+δqxt−q+uty_t=\alpha_0+\delta_0x_t+\delta_1x_{t-1}+\dots+\delta_qx_{t-q}+u_t lagged effects
AR yt=α+ϕ1yt−1+⋯+ety_t=\alpha+\phi_1y_{t-1}+\dots+e_t persistence
ADL AR + DL together general
  • DL: δ0\delta_0 = impact effect; ∑kδk\sum_{k}\delta_k = long-run multiplier. Lags are collinear ⇒ individual δk\delta_k imprecise but the sum can be well-identified.
  • AR: yt−1y_{t-1} correlates with et−1e_{t-1} ⇒ strict exogeneity fails ⇒ consistent but biased in finite samples.
  • Seasonality: add seasonal dummies (11 months or 3 quarters, one omitted as reference). Skip if data already seasonally adjusted.

  • Linear: yt=α0+α1t+ety_t=\alpha_0+\alpha_1 t+e_t — α1\alpha_1 = change per period.
  • Exponential: log⁡yt=α0+α1t+et\log y_t=\alpha_0+\alpha_1 t+e_t — α1≈\alpha_1\approx growth rate per period.

Spurious regression

Two trending series ⇒ huge R2R^2 + "significant" coef even if unrelated (OLS picks up shared drift). Two equivalent fixes (same β^1\hat\beta_1 by Frisch–Waugh–Lovell):

  1. Add tt as a regressor: yt=β0+β1xt+αt+uty_t=\beta_0+\beta_1x_t+\alpha t+u_t.
  2. Detrend both: regress yy and xx on tt, keep residuals y¨,x¨\ddot y,\ddot x, then regress y¨\ddot y on x¨\ddot x.

Event study

  • Estimation period (pre-event): fit a "normal" model, e.g. Rtstock=β0+β1Rtmarket+utR_t^{\text{stock}}=\beta_0+\beta_1R_t^{\text{market}}+u_t.
  • Abnormal return: ARt=Rtactual−R^tpredictedAR_t=R_t^{\text{actual}}-\hat R_t^{\text{predicted}}.
  • Dummy version: add day.before, day.of, day.after dummies; their coefficients γ^\hat\gamma are the ARs. Significant γ^0\hat\gamma_0 = event effect; significant γ^−1\hat\gamma_{-1} = anticipation/leak.
  • Use trading days, not calendar days. Watch for confounding events near the window.

9. Lecture 8 — Fixed effects in panel data

Panel = same units over time ⇒ compare a unit to itself, controlling for all its time-invariant traits.

yit=β0+β1xit+ai+uit\boxed{y_{it}=\beta_0+\beta_1x_{it}+a_i+u_{it}}

aia_i = individual fixed effect — absorbs everything stable about ii (observed or not).

  • Between variation (across-unit means) = confounded. Within variation (unit vs its own mean) = clean.
  • Within estimator / demeaning: subtract each unit's mean, y¨it=yit−yˉi\ddot y_{it}=y_{it}-\bar y_i, then OLS:
β^1within=∑i∑tx¨ity¨it∑i∑tx¨it2\hat\beta_1^{\text{within}}=\frac{\sum_i\sum_t\ddot x_{it}\ddot y_{it}}{\sum_i\sum_t\ddot x_{it}^2}

Equivalent to adding a dummy per unit (FWL theorem). Time-invariant variables drop out (can't be estimated).

  • Time FE: dummy per period (absorbs shocks common to all units). Two-way FE: unit + time. Interacted (city×month) absorbs cell-specific effects.
  • R: feols(y ~ x | unit) · feols(y ~ x | city + month) · feols(y ~ x | city^month).

Standard-error chooser

SE type When R
Robust (Huber–White, HC) heteroskedasticity vcov="hetero"
HAC (Newey–West) time series, serial corr vcov="newey_west"
Clustered panel / FE cluster=~unit
Exam tip

"Why FE over OLS?" → name the time-invariant confounder, show aia_i absorbs it by demeaning (anything constant within a unit subtracts to 0), and cluster SEs at the FE level. Put the FE at the level where the confounder is constant.


10. Lecture 9 — Regression Discontinuity (RDD)

Treatment switches at a cutoff of a running variable. Units just on either side are comparable ⇒ the jump in the outcome at the cutoff = the effect.

  • Continuity: absent treatment, the outcome would pass smoothly through the cutoff (credible because units can't perfectly sort).
  • Gives a LATE at the threshold: strong internal, weak external validity.

Sharp RDD regression (center the running variable!)

Y=β0+β1(R−c)+β2Treated+β3[Treated⋅(R−c)]+uY=\beta_0+\beta_1(\text{R}-c)+\beta_2\text{Treated}+\beta_3\big[\text{Treated}\cdot(\text{R}-c)\big]+u

Treatment effect = β^2\hat\beta_2 (coefficient on the Treated dummy = the gap at the cutoff). Centering at cc makes β0\beta_0 and β0+β2\beta_0+\beta_2 the two lines' heights at the threshold. Quadratic version adds (R−c)2(\text R-c)^2 terms — jump is still the coef on Treated.

Fuzzy RDD = IV at the cutoff

Cutoff only shifts the probability of treatment ⇒ "above cutoff" instruments "treated." Wald estimator:

effect=jump in outcomejump in treatment probability=reduced formfirst stage\text{effect}=\frac{\text{jump in outcome}}{\text{jump in treatment probability}}=\frac{\text{reduced form}}{\text{first stage}}

Validity checks

  • McCrary density test: is the count of units smooth at the cutoff? A jump ⇒ manipulation/sorting.
  • Placebo test: do predetermined covariates jump at the cutoff? They shouldn't.
  • Bandwidth & functional form: narrow window so a line fits locally; check estimate is stable across bandwidths. Avoid high-order polynomials. R: rdrobust(y, x, c=0), rdbwselect.
Exam tip

Read the jump off the Treated coefficient. Always state continuity, that it's a local effect, and the two diagnostic tests (McCrary = data density; placebo = covariates). For fuzzy, report first-stage strength.


11. Lecture 10 — Difference-in-Differences (DiD)

Two groups (treatment/control) × two periods (before/after). Works on repeated cross-sections (same individuals not needed). Estimand = ATT.

Estimator (difference of differences)

δ^1=(yˉ2,T−yˉ1,T)⏟treat change−(yˉ2,C−yˉ1,C)⏟control change=(yˉ2,T−yˉ2,C)⏟after gap−(yˉ1,T−yˉ1,C)⏟before gap\hat\delta_1=\underbrace{(\bar y_{2,T}-\bar y_{1,T})}_{\text{treat change}}-\underbrace{(\bar y_{2,C}-\bar y_{1,C})}_{\text{control change}}=\underbrace{(\bar y_{2,T}-\bar y_{2,C})}_{\text{after gap}}-\underbrace{(\bar y_{1,T}-\bar y_{1,C})}_{\text{before gap}}

The control's change = the counterfactual common trend you subtract off.

As a regression (gives SEs + controls)

yi=β0+δ0 d2i+β1 dTi+δ1(d2i⋅dTi)+uiy_i=\beta_0+\delta_0\,d2_i+\beta_1\,dT_i+\delta_1(d2_i\cdot dT_i)+u_i

d2d2 = after dummy, dTdT = treatment dummy. DiD effect = δ1\delta_1 (the interaction). Maps to the 2×2 table: β0\beta_0 = control-before, δ0\delta_0 = time trend, β1\beta_1 = fixed group gap, δ1\delta_1 = effect. With many groups/periods this is two-way FE.

Validity

  1. Parallel trends (the big one) — absent treatment, both groups move together. Can't prove; inspect pre-trends.
  2. No coincident shock to the control group.
  3. Groups comparable. Adding controls/FE ⇒ conditional parallel trends.
Exam tip

Compute δ1\delta_1 from the 2×2 means (either difference order — same answer). If significance asked, it's the interaction coef in the regression. Always name parallel trends as the identifying assumption.


12. Standard-error & test quick-reference

Situation Use Test it with
Heteroskedasticity (incl. any LPM) Robust HC SEs (heteroskedasticity is assumed)
Time-series serial correlation HAC / Newey–West u^t\hat u_t on u^t−1\hat u_{t-1}, H0:ρ=0H_0:\rho=0
Panel / fixed effects Clustered SEs (at FE level) —
Joint significance (OLS/LPM) F-test anova()
Joint significance (logit/probit) LR test χq2\chi^2_q lrtest()
Endogeneity present? — Wu–Hausman
All instruments valid? (overid) — Sargan/Hansen

Critical values (5%, large nn): t,z≈1.96t,z\approx1.96 · χ12=3.84\chi^2_1=3.84, χ22=5.99\chi^2_2=5.99, χ32=7.81\chi^2_3=7.81, χ42=9.49\chi^2_4=9.49 · weak-instrument first-stage F > 10.


13. Exam question playbook

"Interpret this coefficient." Identify the model first. LPM → pp change in P(y=1)P(y=1). Log-dep → ≈100β%\approx100\beta\%. Logit/probit raw → sign & significance only (magnitude needs g(⋅)βg(\cdot)\beta). State significance: ∣t∣>1.96|t|>1.96 or p<0.05p<0.05.

"Is β^\hat\beta causal / what's wrong with OLS here?" Name the mechanism — confounder (OVB), reverse causality, simultaneity, or selection — write Cov(x,u)≠0\text{Cov}(x,u)\neq0, then prescribe the fix (control / FE / IV / Heckman / RDD / DiD).

"Draw / update the causal diagram." Put the confounder as a node with arrows into both the regressor and the outcome; that backdoor path is the bias. Unobserved confounder sits in uu.

"Evaluate this instrument." Relevance Cov(z,x)≠0\text{Cov}(z,x)\neq0 (first-stage F>10, testable) + validity Cov(z,u)=0\text{Cov}(z,u)=0 (argue, untestable). Give one failure story.

"Compute the DiD / RDD effect." DiD = difference of the two changes (= interaction δ1\delta_1). RDD = coefficient on the Treated dummy (centered running var). Fuzzy RDD = outcome jump ÷ treatment-prob jump.

"Which SEs and why?" Heteroskedasticity → robust; time series → Newey–West; panel/FE → clustered. Default homoskedastic SEs are almost never right in this course.

"State the identifying assumption." IV → exclusion restriction. FE → confounder is time-invariant. RDD → continuity (no sorting). DiD → parallel trends. Heckman → exclusion variable in selection eqn.


The 6 estimators in one breath

LPM = OLS on 0/1 (pp, robust SE). Logit/Probit = CDF squashes to (0,1), ME = g(⋅)βg(\cdot)\beta, MLE + LR test. IV/2SLS = borrow exogenous variation, relevance + validity, LATE. Heckman = probit selection → add IMR. Fixed effects = demean to kill time-invariant confounders, cluster SEs. RDD = jump at a cutoff (continuity). DiD = difference of differences (parallel trends).