Exam Formula Sheet (Lec 1–10)
One-page-per-topic exam reference — every key term, formula, interpretation, and what-to-do tip from Lectures 1–10. Built for the closed-book final.
- #econometrics
- #exam-prep
- #formula-sheet
- #reference
- #cheat-sheet
Part of: Econometrics · Glossary: _Econometrics Concepts Scope: the whole Sem-2 causal-inference toolkit. Each section = one lecture: key terms → formulas → how to read them → exam tips. Default everywhere: 5% significance, two-sided; reject if (or ). Standard errors are robust/clustered unless told otherwise.
0. Method picker — "which tool does this question want?"
| You see in the question… | Reach for | Lecture |
|---|---|---|
| Binary outcome , want simple/linear answer | LPM (OLS + robust SE) | L2 |
| Binary outcome, want valid probabilities / non-constant effect | Logit / Probit | L3 |
| Regressor correlated with the error (OVB, reverse causality, simultaneity) | IV / 2SLS | L4, L6 |
| Outcome only observed for a self-selected subsample | Heckman two-step | L5 |
| Price & quantity (supply/demand) jointly determined | Simultaneous eqns + IV | L6 |
| Data over time, errors correlated across periods | HAC / serial corr. fix | L6 |
| Both variables trend upward over time | detrend / add | L7 |
| Effect of one dated event on an outcome | Event study | L7 |
| Same units over time + unobserved fixed traits | Fixed effects (demeaning) | L8 |
| Treatment switches at a threshold of some variable | RDD | L9 |
| Treatment vs control, before vs after a policy | DiD | L10 |
First move on every applied question(1) Is the outcome binary? (2) Is the key regressor exogenous, or is there a confounder / reverse causality / selection? (3) What's the data structure — cross-section, time series, or panel? Those three answers pin down the method.
1. Foundations recap (assumed from Intro Metrics)
Classical assumptions — A1–A5
A1 linear in parameters · A2 random sampling · A3 no perfect collinearity · A4 zero conditional mean (the causal one) · A5 homoskedasticity . Under A1–A5, OLS is BLUE.
Course convention (varies by source): A3 = independent errors / A4 = / A5 = homoskedastic. Read the question's numbering — the content is what matters.
t-test (single coefficient)
Reject if (large , 5%).
F-test (joint / nested models) — F-test
= number of restrictions (regressors dropped). Large / small ⇒ reject that the dropped variables are jointly zero. In R: anova(unrestricted, restricted).
Confidence interval
Functional-form interpretation (quick table)
| Model | Coefficient means… |
|---|---|
| level–level | (units) |
| log–level | changes |
| level–log | up 1% |
| log–log | = elasticity (% per %) |
2. Lecture 1 — Causal inference & treatment effects
Goal: read as causal, not just correlational. Requires ruling out confounders, reverse causality, selection.
- Data types: cross-section · time series · panel · repeated cross-section.
- Exogenous (X) → endogenous (Y): the arrow is causal only if X varies independently of everything else affecting Y.
- Continuous vs dummy regressor: with few discrete values, use dummies (one per level, minus a reference). Dummies are flexible (each level free); a continuous slope forces one straight line — test it with a joint F-test on the dummies.
- Panel breaks A3/A4: repeated obs per unit ⇒ errors correlated within unit ⇒ classical SEs too small. Quick fix: subsample one obs/unit. Proper fix: fixed effects + clustered SEs (L8).
Exam tip"Is this coefficient causal?" → name the specific confounder / reverse-causality / selection story, state , then say which tool (control, FE, IV, RDD, DiD) shuts that path.
3. Lecture 2 — Linear Probability Model (LPM)
Binary run through OLS. Under A2, the fit is a probability:
- Coefficient = marginal effect in percentage points (constant). . Dummy coef = difference in vs base group.
- Drawback 1 — unbounded: can fall below 0 or exceed 1 (nonsense probabilities). → motivates logit/probit.
- Drawback 2 — built-in heteroskedasticity (always):
- Fix for drawback 2: robust (HC) standard errors — mandatory for any LPM. stays unbiased; only SEs were wrong. R:
feols(y ~ x, se="hetero").
Exam tipIf the outcome is 0/1, every
lm/feolsis an LPM — interpret coefficients in pp and expect robust/clustered SEs. Saying " can leave [0,1]" + "errors are heteroskedastic by construction" earns the two standard marks.
4. Lecture 3 — Logit & Probit (discrete choice)
Wrap the linear index in a CDF so :
| (CDF) | density | R | |
|---|---|---|---|
| Logit | dlogis |
glm(...,family=binomial(link="logit")) |
|
| Probit | dnorm |
glm(...,family=binomial(link="probit")) |
is increasing, S-shaped, → 0/1 at the tails, steepest at (), symmetric .
Marginal effects (the key formula)
Non-constant — depends where you sit on the curve. Evaluate at sample means for "the average person."
- Raw coef → sign & significance only, NOT magnitude. Never compare raw coef magnitudes across LPM/logit/probit; do compare marginal effects.
- Estimation = MLE (model is non-linear, OLS invalid). Log-likelihood:
Joint testing = Likelihood Ratio test (replaces F)
= # restrictions. Reject if . R: lrtest(ur, r).
Exam tipRecipe for a marginal effect: (1) get ; (2) compute at the means; (3) multiply by —
dnorm(probit) /dlogis(logit). If only sign asked, just read the coefficient sign + /.
5. Lecture 4 — Instrumental Variables & 2SLS
Problem: endogeneity (OVB, reverse causality, simultaneity) ⇒ OLS biased & inconsistent (more data doesn't help).
Fix: an instrument that moves but not directly. Two conditions:
| Condition | Formula | Testable? |
|---|---|---|
| Relevance | ✅ first-stage F > 10 | |
| Validity (exclusion) | ❌ argue conceptually |
2SLS
- First stage:
- Second stage:
- Simple IV estimator:
R (correct SEs): feols(y ~ controls | x ~ z). Never run two manual lms (wrong second-stage SEs).
- # instruments ≥ # endogenous regressors (need ≥2 instruments for 2 endogenous vars).
- Weak instrument (F<10): tiny denominator amplifies bias — can be worse than OLS.
- IV estimates LATE, not ATE — the effect for compliers (those whose moves with ).
- IV is less precise than OLS (larger SEs) — the price of consistency.
- Sargan/overid test (only if overidentified): = all instruments valid. Reject ⇒ ≥1 invalid.
- Wu–Hausman: = regressor exogenous (OLS fine). Reject ⇒ IV needed.
Exam tipTo defend an instrument, write both conditions explicitly with the covariance, show relevance via first-stage F, then give a one-sentence story for why can't reach except through — and a sentence on how it could fail. Validity can never be "proven from a table."
6. Lecture 5 — Sample selection & Heckman
Consistency vs unbiasedness
- Unbiased: (right on average, any ).
- Consistent: as .
- OLS with endogenous : biased AND inconsistent. 2SLS: biased in small samples but consistent.
Sample selection — you only observe when
| Type | Selection depends on | OLS |
|---|---|---|
| Random | nothing | ✅ unbiased |
| Exogenous | only (observed) | ✅ unbiased |
| Endogenous | (unobservables affecting ) | ❌ biased |
Mincer running example
Heckman two-step ("Heckit")
- Step 1 (full sample, probit): . Need an exclusion restriction — a variable in that affects selection but not the outcome (e.g. kids under 6 → labour-force participation, not the wage). Compute inverse Mills ratio:
- Step 2 (selected sample, OLS): add as a regressor:
- Test for selection bias: . Reject ⇒ endogenous selection present, correction matters.
Exam tipHeckman mirrors 2SLS: step 1 builds a correction term, step 2 adds it as a control. Always state the exclusion-restriction variable and what testing tells you.
7. Lecture 6 — Simultaneous equations & time series
Simultaneous equations (supply & demand)
Price & quantity jointly determined ⇒ in both equations ⇒ OLS biased. Fix: a shifter of one curve instruments price to trace the other.
- Weather shifts supply → identifies the demand curve.
- Day of week shifts demand → identifies the supply curve.
Serial correlation
. OLS stays unbiased & consistent but not efficient and SEs are wrong (too small) ⇒ invalid /.
- Test: regress ; .
- Fix: HAC (Newey–West) SEs — coefficients unchanged, SEs widened. R:
vcov="newey_west".
Time-series OLS assumptions
- TS.2 Strict exogeneity for all (past, present, future) → unbiasedness.
- Contemporaneous exogeneity → consistency only.
- TS.3 homoskedastic · TS.4 no serial corr · TS.5 normal.
Model zoo
| Model | Equation | Use |
|---|---|---|
| Static | instant effect only | |
| Distributed lag | lagged effects | |
| AR | persistence | |
| ADL | AR + DL together | general |
- DL: = impact effect; = long-run multiplier. Lags are collinear ⇒ individual imprecise but the sum can be well-identified.
- AR: correlates with ⇒ strict exogeneity fails ⇒ consistent but biased in finite samples.
- Seasonality: add seasonal dummies (11 months or 3 quarters, one omitted as reference). Skip if data already seasonally adjusted.
8. Lecture 7 — Time trends & event studies
Deterministic trends
- Linear: — = change per period.
- Exponential: — growth rate per period.
Spurious regression
Two trending series ⇒ huge + "significant" coef even if unrelated (OLS picks up shared drift). Two equivalent fixes (same by Frisch–Waugh–Lovell):
- Add as a regressor: .
- Detrend both: regress and on , keep residuals , then regress on .
Event study
- Estimation period (pre-event): fit a "normal" model, e.g. .
- Abnormal return: .
- Dummy version: add
day.before,day.of,day.afterdummies; their coefficients are the ARs. Significant = event effect; significant = anticipation/leak. - Use trading days, not calendar days. Watch for confounding events near the window.
9. Lecture 8 — Fixed effects in panel data
Panel = same units over time ⇒ compare a unit to itself, controlling for all its time-invariant traits.
= individual fixed effect — absorbs everything stable about (observed or not).
- Between variation (across-unit means) = confounded. Within variation (unit vs its own mean) = clean.
- Within estimator / demeaning: subtract each unit's mean, , then OLS:
Equivalent to adding a dummy per unit (FWL theorem). Time-invariant variables drop out (can't be estimated).
- Time FE: dummy per period (absorbs shocks common to all units). Two-way FE: unit + time. Interacted (city×month) absorbs cell-specific effects.
- R:
feols(y ~ x | unit)·feols(y ~ x | city + month)·feols(y ~ x | city^month).
Standard-error chooser
| SE type | When | R |
|---|---|---|
| Robust (Huber–White, HC) | heteroskedasticity | vcov="hetero" |
| HAC (Newey–West) | time series, serial corr | vcov="newey_west" |
| Clustered | panel / FE | cluster=~unit |
Exam tip"Why FE over OLS?" → name the time-invariant confounder, show absorbs it by demeaning (anything constant within a unit subtracts to 0), and cluster SEs at the FE level. Put the FE at the level where the confounder is constant.
10. Lecture 9 — Regression Discontinuity (RDD)
Treatment switches at a cutoff of a running variable. Units just on either side are comparable ⇒ the jump in the outcome at the cutoff = the effect.
- Continuity: absent treatment, the outcome would pass smoothly through the cutoff (credible because units can't perfectly sort).
- Gives a LATE at the threshold: strong internal, weak external validity.
Sharp RDD regression (center the running variable!)
Treatment effect = (coefficient on the Treated dummy = the gap at the cutoff). Centering at makes and the two lines' heights at the threshold. Quadratic version adds terms — jump is still the coef on Treated.
Fuzzy RDD = IV at the cutoff
Cutoff only shifts the probability of treatment ⇒ "above cutoff" instruments "treated." Wald estimator:
Validity checks
- McCrary density test: is the count of units smooth at the cutoff? A jump ⇒ manipulation/sorting.
- Placebo test: do predetermined covariates jump at the cutoff? They shouldn't.
- Bandwidth & functional form: narrow window so a line fits locally; check estimate is stable across bandwidths. Avoid high-order polynomials. R:
rdrobust(y, x, c=0),rdbwselect.
Exam tipRead the jump off the Treated coefficient. Always state continuity, that it's a local effect, and the two diagnostic tests (McCrary = data density; placebo = covariates). For fuzzy, report first-stage strength.
11. Lecture 10 — Difference-in-Differences (DiD)
Two groups (treatment/control) × two periods (before/after). Works on repeated cross-sections (same individuals not needed). Estimand = ATT.
Estimator (difference of differences)
The control's change = the counterfactual common trend you subtract off.
As a regression (gives SEs + controls)
= after dummy, = treatment dummy. DiD effect = (the interaction). Maps to the 2×2 table: = control-before, = time trend, = fixed group gap, = effect. With many groups/periods this is two-way FE.
Validity
- Parallel trends (the big one) — absent treatment, both groups move together. Can't prove; inspect pre-trends.
- No coincident shock to the control group.
- Groups comparable. Adding controls/FE ⇒ conditional parallel trends.
Exam tipCompute from the 2×2 means (either difference order — same answer). If significance asked, it's the interaction coef in the regression. Always name parallel trends as the identifying assumption.
12. Standard-error & test quick-reference
| Situation | Use | Test it with |
|---|---|---|
| Heteroskedasticity (incl. any LPM) | Robust HC SEs | (heteroskedasticity is assumed) |
| Time-series serial correlation | HAC / Newey–West | on , |
| Panel / fixed effects | Clustered SEs (at FE level) | — |
| Joint significance (OLS/LPM) | F-test | anova() |
| Joint significance (logit/probit) | LR test | lrtest() |
| Endogeneity present? | — | Wu–Hausman |
| All instruments valid? (overid) | — | Sargan/Hansen |
Critical values (5%, large ): · , , , · weak-instrument first-stage F > 10.
13. Exam question playbook
"Interpret this coefficient." Identify the model first. LPM → pp change in . Log-dep → . Logit/probit raw → sign & significance only (magnitude needs ). State significance: or .
"Is causal / what's wrong with OLS here?" Name the mechanism — confounder (OVB), reverse causality, simultaneity, or selection — write , then prescribe the fix (control / FE / IV / Heckman / RDD / DiD).
"Draw / update the causal diagram." Put the confounder as a node with arrows into both the regressor and the outcome; that backdoor path is the bias. Unobserved confounder sits in .
"Evaluate this instrument." Relevance (first-stage F>10, testable) + validity (argue, untestable). Give one failure story.
"Compute the DiD / RDD effect." DiD = difference of the two changes (= interaction ). RDD = coefficient on the Treated dummy (centered running var). Fuzzy RDD = outcome jump ÷ treatment-prob jump.
"Which SEs and why?" Heteroskedasticity → robust; time series → Newey–West; panel/FE → clustered. Default homoskedastic SEs are almost never right in this course.
"State the identifying assumption." IV → exclusion restriction. FE → confounder is time-invariant. RDD → continuity (no sorting). DiD → parallel trends. Heckman → exclusion variable in selection eqn.
The 6 estimators in one breathLPM = OLS on 0/1 (pp, robust SE). Logit/Probit = CDF squashes to (0,1), ME = , MLE + LR test. IV/2SLS = borrow exogenous variation, relevance + validity, LATE. Heckman = probit selection → add IMR. Fixed effects = demean to kill time-invariant confounders, cluster SEs. RDD = jump at a cutoff (continuity). DiD = difference of differences (parallel trends).