Linear Probability Model (LPM)
- #econometrics
- #LPM
- #binary-outcomes
- #game-theory
- #ultimatum-game
- #heteroskedasticity
- #robust-standard-errors
Linear Probability Model (LPM)
Part of: Econometrics Lecture 02 — Applied Econometrics, Dr. Aluma Dembo Previous: Lec_01-Introduction & Treatment Effects | Next: Lec_03-Logit & Probit Models Key concepts: Binary Outcomes, LPM, Heteroskedasticity, Robust Standard Errors, Ultimatum Game
Motivation — Binary Outcomes
Often the outcome of interest is binary: accept/reject, buy/don't buy, employed/unemployed, vote yes/no. We code it as and ask: what's the probability , and how does it depend on ?
Today's agenda
- Motivating example — the ultimatum game (Andersen et al. 2011)
- Define the Linear Probability Model
- Show under A2
- Interpret coefficients
- Identify LPM's two big drawbacks — and how to patch one of them
Worked Example: The Ultimatum Game
The SetupTwo players, one-shot, anonymous:
- Proposer receives a pot and offers share to the Responder.
- Responder sees the offer and either accepts (proposer keeps , responder gets ) or rejects (both get 0).
Subgame-perfect Nash prediction: proposer offers the smallest positive ; responder accepts anything . Real humans reject "unfair" offers — a classic violation of pure self-interest.
Andersen, Ertaç, Gneezy, Hoffman & List (2011)
Ran the ultimatum game in rural India with stakes varied from 20 Rs up to 20,000 Rs (roughly a year's income) to test whether "irrational" rejection persists when money actually matters.
| Variable | Meaning |
|---|---|
| Did the responder accept? | |
| Share of the pie offered | |
| Dummies for 200 / 2,000 / 20,000 Rs (ref = 20 Rs) |
Research question: Does the probability of acceptance rise with the offer, and does it change when the stakes are large?
The Linear Probability Model
For a binary , run OLS directly:
The magic: under assumption A2 (), the fitted value is a conditional probability.
Proof — is an estimate of
Take conditional expectations of both sides:
Since is binary:
Therefore:
So estimates the conditional probability of acceptance.
Interpreting LPM Coefficients
Continuous regressor
A one-unit increase in changes the probability of by — in percentage points, not percent.
Dummy regressor
If , then = difference in between the treatment and reference group.
Andersen et al. results
Interpretation:
- : a 10 percentage-point larger offer raises acceptance probability by 5.83 pp.
- : at the 20,000 Rs stakes, acceptance probability is 39.2 pp higher than in the 20 Rs reference group — people are far more accepting when rejecting is very costly.
Key findingHigh stakes do change behaviour — responders are much more willing to accept low offers when the amount at stake is meaningful. The "fairness" penalty shrinks as money grows.
LPM Drawback #1 — Unbounded Predictions
Probabilities must lie in . OLS has no such constraint: can be negative or greater than 1.
In Andersen et al., the max fitted value is 1.1943 — a "probability" of 119%, which is nonsense.
type: lpm-problems
mode: unbounded
Why this happensLPM fits a straight line through a cloud of 0s and 1s. Extend the line far enough and it leaves the unit interval. The more extreme gets, the more often this happens. In the shaded regions the model predicts probabilities below 0 or above 1 — exactly the 119% fitted value Andersen et al. hit.
This is the motivation for Lec_03-Logit & Probit Models (next lecture), which squash into via the logistic / normal CDF.
LPM Drawback #2 — Built-in Heteroskedasticity
Proof —
Since , is also binary, taking:
- with probability
- with probability
Then (by A2), and:
What this violatesSince depends on through , assumption A4 (homoskedasticity) is always violated in LPM. The classical OLS standard errors are wrong — usually too small.
The shape of the variance
| 0.1 | 0.09 |
| 0.3 | 0.21 |
| 0.5 | 0.25 (max) |
| 0.7 | 0.21 |
| 0.9 | 0.09 |
Variance is largest when (maximum uncertainty) and shrinks toward the extremes.
type: lpm-problems
mode: variance
Reading the frownThe error variance traces an inverted parabola in — it has to, because a coin near 50/50 is the least predictable while one near 0 or 1 is nearly certain. Because the height of this curve changes with , the homoskedasticity assumption (A4) is broken by construction in every LPM, which is why robust standard errors are non-negotiable.
Fix: Robust (Heteroskedasticity-Consistent) Standard Errors
from OLS is still unbiased — it's just the standard errors that are wrong. Compute HC/"sandwich" standard errors instead:
- In R base:
lmtest::coeftest(m, vcov = sandwich::vcovHC(m, type = "HC1")) - Cleaner:
fixest::feols(accept ~ percent_offer + factor(stakes), data = df, se = "hetero")
Rule of thumbAlways report robust SEs for LPM. It costs you nothing and repairs the one assumption the model definitionally breaks.
Summary — What Lecture 2 Teaches
- Binary outcomes the OLS fitted value is a conditional probability (under A2).
- LPM coefficients are marginal effects on , measured in percentage points.
- Empirical lesson: Andersen et al. show fairness preferences survive small stakes but weaken sharply at life-changing amounts.
- Drawback 1 (fatal-ish): can leave — motivates logit/probit.
- Drawback 2 (fixable): errors are always heteroskedastic with — use robust SEs.
Related Notes
- Previous: Lec_01-Introduction & Treatment Effects
- Next: Lec_03-Logit & Probit Models — fixes the problem
- Foundational: OLS Estimation, Classical Assumptions A1-A5, Hypothesis Testing
- Related: Heteroskedasticity, Robust Standard Errors, Game Theory