Foundations, OLS & inference 14

  • OLS Estimation:

    Ordinary Least Squares: chooses coefficients that minimise the sum of squared residuals. Unbiased and BLUE under the classical assumptions. Lec 1

  • Hypothesis Testing:

    Framework for testing claims about parameters via a test statistic, a null/alternative, and a significance level (5% throughout). Lec 1

  • F-test:

    Joint-significance test comparing nested models; built from the residual sum of squares of the restricted vs unrestricted model. Large FF (small pp) ⇒ reject that the dropped regressors are jointly zero. PP1 Q1a

  • Classical Assumptions A1-A5:

    Gauss–Markov conditions: linearity, random sampling, no perfect collinearity, zero conditional mean E[u∣X]=0E[u\mid X]=0, and homoskedasticity. Under them OLS is BLUE. Lec 2

  • Heteroskedasticity:

    Error variance depends on xx (Var(u∣x)(u\mid x) not constant); violates A5, so classical SEs are wrong (usually too small). Lec 2

  • Robust Standard Errors:

    Heteroskedasticity-consistent ("sandwich"/HC) standard errors; give correct inference when errors are heteroskedastic — mandatory for the LPM. Lec 2

  • Consistency:

    An estimator converges in probability to the true value as n→∞n\to\infty. Weaker than unbiasedness. Lec 5

  • Endogeneity:

    A regressor is correlated with the error, Cov(x,u)≠0(x,u)\neq0, making OLS biased and inconsistent. Sources: omitted variables, simultaneity, measurement error. Lec 4

  • Omitted Variable Bias:

    Bias from leaving out a variable that both affects yy and correlates with an included regressor; the omitted effect loads onto the included coefficient. PS4

  • Causal Inference:

    Estimating the effect of a cause — the counterfactual change in yy from changing xx — as opposed to mere correlation. Lec 1

  • Causal Diagram:

    Directed acyclic graph (DAG) of assumed causal relationships; used to spot confounders and backdoor paths. PS2

  • Dummy Variables:

    Binary 0/1 indicator for a category; its coefficient is the mean difference versus the omitted base group. Lec 2

  • Game Theory:

    Study of strategic decision-making; here the backdrop for binary-choice experiments. Lec 2

  • Ultimatum Game:

    Proposer offers a split, responder accepts/rejects; rejecting "unfair" offers violates pure self-interest. The motivating LPM example (Andersen et al. 2011). Lec 2

Treatment effects & potential outcomes 4

  • Treatment Effect:

    Difference between a unit's outcome with vs without treatment, (Y∣T=1)−(Y∣T=0)(Y\mid T=1)-(Y\mid T=0); never observable for one unit (the fundamental problem of causal inference). Lec 10

  • Counterfactual:

    The unobserved potential outcome — what would have happened under the other treatment status. The central missing quantity in causal inference. Lec 10

  • Average Treatment Effect on the Treated:

    ATT: the mean treatment effect among treated units; the estimand that difference-in-differences recovers. Lec 10

  • Local Average Treatment Effect:

    LATE: IV identifies the effect only for "compliers" — units whose treatment status responds to the instrument. Lec 4

Binary outcomes — LPM, logit & probit 7

  • Linear Probability Model:

    OLS on a binary outcome; the fitted value is P(y=1∣x)P(y=1\mid x) and coefficients are marginal effects in percentage points. Drawbacks: predictions can leave [0,1][0,1], and errors are inherently heteroskedastic. Lec 2

  • Binary Outcomes:

    Outcome coded 0/1 (accept/reject, chose B/A); modelled with the LPM, logit, or probit. Lec 2

  • Logit Model:

    Binary model using the logistic CDF to keep P(y=1)P(y=1) in (0,1)(0,1); estimated by maximum likelihood. Lec 3

  • Probit Model:

    Binary model using the normal CDF Φ(⋅)\Phi(\cdot); coefficients give sign and significance, not the marginal effect. Lec 3

  • Marginal Effects:

    The partial derivative ∂P(y=1)/∂x\partial P(y=1)/\partial x. In the LPM it equals the coefficient (constant); in logit/probit it is β⋅density(x′β)\beta\cdot\text{density}(\mathbf{x}'\beta) and varies with xx. Lec 3

  • Maximum Likelihood Estimation:

    Estimates parameters by maximising the likelihood of the observed data; used for logit/probit. Consistent and asymptotically normal. Lec 3

  • Likelihood Ratio Test:

    Tests restrictions in MLE models by comparing log-likelihoods of restricted vs unrestricted models — the MLE analogue of the F-test. Lec 3

Instrumental variables 9

  • Instrumental Variables:

    IV: uses an instrument zz to isolate exogenous variation in an endogenous regressor, restoring consistency when Cov(x,u)≠0(x,u)\neq0. Lec 4

  • Two Stage Least Squares:

    2SLS: regress xx on zz (first stage), then yy on the fitted x^\hat x (second stage); the practical IV estimator. Lec 4

  • First Stage:

    Regression of the endogenous regressor on the instrument(s) and exogenous controls; its strength gauges relevance. Lec 4

  • Second Stage:

    Regression of the outcome on the first-stage fitted values, giving the IV/2SLS estimate. Lec 4

  • Instrument Relevance:

    The instrument must be correlated with the endogenous regressor, Cov(z,x)≠0(z,x)\neq0 (testable, e.g. first-stage FF). Lec 4

  • Instrument Validity:

    The exclusion restriction: the instrument affects yy only through xx, Cov(z,u)=0(z,u)=0 (untestable; argued conceptually). Lec 4

  • Weak Instruments:

    Instruments only weakly correlated with xx (first-stage F<10F<10); produce biased, imprecise IV estimates. Lec 4

  • Wu-Hausman Test:

    Tests for endogeneity; rejection ⇒ OLS is inconsistent and IV is needed. Lec 8

  • Overidentifying Restrictions Test:

    Sargan/Hansen test, available only when instruments outnumber endogenous regressors (overidentified); tests the joint null that all instruments are valid (uncorrelated with the error). Rejection ⇒ at least one instrument violates the exclusion restriction (assuming ≥1 is valid). PP3 Q2c

Sample selection (Heckman) 5

  • Sample Selection Bias:

    Bias from non-random inclusion in the estimation sample when selection depends on unobservables that affect yy, Cov(selection, u)≠0u)\neq0. Lec 5

  • Heckman Selection Model:

    Two-step estimator: model selection (probit), then add the inverse Mills ratio to the outcome regression to correct selection bias. Lec 5

  • Inverse Mills Ratio:

    Correction term λ\lambda added in Heckman step 2; a significant λ\lambda confirms selection bias. Lec 5

  • Endogenous Selection:

    When who appears in the sample depends on unobservables that also drive the outcome. Lec 5

  • Mincer Wage Equation:

    Standard log-wage model: log⁡(wage)=f(educ,exper,exper2)\log(\text{wage})=f(\text{educ},\text{exper},\text{exper}^2); a common setting for selection corrections. Lec 5

Simultaneous equations & time series 10

  • Simultaneous Equations Model:

    System where variables are jointly determined (e.g. supply & demand), creating simultaneity bias. Lec 6

  • Serial Correlation:

    Errors correlated over time, Cov(ut,ut−1)≠0(u_t,u_{t-1})\neq0; biases classical SEs in time series. Lec 6

  • HAC Standard Errors:

    Newey–West standard errors, robust to heteroskedasticity and autocorrelation. Lec 6

  • Time Series:

    Observations indexed by time, where the past can influence the future. Lec 6

  • Static Model:

    xx affects yy only contemporaneously (no lags). Lec 6

  • Strict Exogeneity:

    E[ut∣X]=0E[u_t\mid X]=0 for all time periods (stronger than contemporaneous exogeneity). Lec 6

  • Distributed Lag Model:

    yty_t depends on current and past values of xx; the lag coefficients trace the dynamic response. Lec 6

  • Autoregressive Model:

    yty_t depends on its own lagged values; consistent but not unbiased. Lec 6

  • Seasonality:

    Regular calendar-driven patterns in a series, controlled with seasonal dummy variables. Lec 6

  • Seasonal Controls:

    Seasonal dummy variables added to absorb predictable calendar effects. Lec 6

Panel data & fixed effects 11

  • Panel Data:

    Repeated observations on the same units over time. Lec 8

  • Fixed Effects:

    Unit/group dummies (intercepts) that absorb all time-invariant characteristics of that unit. Lec 8

  • Within Estimator:

    OLS on demeaned data; identifies β\beta from within-unit variation only. Lec 8

  • Demeaning:

    Subtracting each unit's own mean from every variable; annihilates anything constant within the unit (so fixed effects vanish). Lec 8

  • Between Variation:

    Variation in unit averages across units; contaminated by cross-unit confounders. Lec 8

  • Within Variation:

    Variation within a unit over time; the variation fixed effects exploit. Lec 8

  • Individual Fixed Effect:

    aia_i: a per-individual intercept capturing all of that individual's time-invariant traits. Lec 8

  • Time Fixed Effects:

    Per-period dummies absorbing shocks common to all units in that period. Lec 8

  • Two-Way Fixed Effects:

    Unit and time fixed effects together; the panel form of difference-in-differences. Lec 8

  • Clustered Standard Errors:

    SEs allowing arbitrary correlation within groups; cluster at the level of the fixed effect. Lec 8

  • Pooled OLS:

    OLS on all panel rows ignoring the ii/tt structure; biased when fixed unit traits correlate with the regressor. Lec 8

Regression discontinuity 10

  • Regression Discontinuity:

    Exploits a treatment cutoff in a running variable; the jump in the outcome at the cutoff is the effect. Lec 9

  • Running Variable:

    The forcing variable whose value relative to the cutoff determines treatment. Lec 9

  • Cutoff:

    Threshold of the running variable that switches treatment on/off. Lec 9

  • Bandwidth:

    Window around the cutoff used for estimation; a bias–variance trade-off. Lec 9

  • Continuity Assumption:

    Absent treatment, the outcome would vary smoothly through the cutoff. Lec 9

  • Sharp RDD:

    Treatment probability jumps cleanly 0→10\to1 at the cutoff. Lec 9

  • Fuzzy RDD:

    Crossing the cutoff only changes the probability of treatment; estimated by IV (the Wald estimator). Lec 9

  • McCrary Density Test:

    Checks that the density of the running variable is smooth at the cutoff (no manipulation/sorting). Lec 9

  • Placebo Test:

    Checks that predetermined covariates (or pre-periods) show no jump/effect where none should exist. Lec 9

  • Wald Estimator:

    (Jump in outcome) ÷ (jump in treatment probability); the fuzzy-RDD / IV ratio. Lec 9

Difference-in-differences 5