Back to the on-screen lesson ·
Model a yes-or-no outcome with a straight line or an S-curve, read LPM slopes in percentage points, and compute logit probabilities, marginal effects and odds.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you will be able to read a linear probability model, compute a logit probability and marginal effect, and work with odds.
You can read a regression with dummies on the right-hand side, and from the second lesson you know that a regression predicts a conditional mean. From statistics you know that the mean of a variable that is either zero or one is the share of ones: the probability.
Many economic outcomes are yes or no. Is a mortgage denied? Does a person work? Does a customer renew? Does a firm export? When the dependent variable is a dummy, a regression predicts a probability, and this lesson shows the three standard ways to model it and how to read each.
| Term | What it means |
|---|---|
| Binary outcome | A dependent variable that equals 1 or 0. |
| Linear probability model | OLS with a binary $y$: $P(y = 1 \mid x) = \beta_0 + \beta_1 x$. |
| Logit | $P = e^{z}/(1 + e^{z})$ with index $z = \beta_0 + \beta_1 x$. |
| Probit | $P = \Phi(z)$, the standard normal cumulative distribution of the index. |
| Index | The linear combination $z = \beta_0 + \beta_1 x$ inside a logit or probit. |
| Marginal effect | The change in probability for a small change in $x$: $\beta_1 G'(z)$. |
| Odds | $p/(1 - p)$; in a logit, $e^{\beta_1}$ is the factor by which the odds change per unit of $x$. |
With a binary outcome, $E[y \mid x] = P(y = 1 \mid x)$. The linear probability model (LPM) fits that probability with OLS:
$$P(y = 1 \mid x) = \beta_0 + \beta_1 x.$$
Its slope is directly a change in probability — multiply by 100 for percentage points — and everything from earlier lessons carries over. Its drawbacks: a straight line can predict probabilities below zero or above one, and it assumes the effect of $x$ is the same whether the probability starts at $0.05$ or $0.5$. Its errors are also heteroskedastic by construction, so robust standard errors are required.
Logit and probit pass the index $z = \beta_0 + \beta_1 x$ through an S-shaped function $G$ that stays between zero and one:
$$P(y = 1 \mid x) = G(z), \qquad \text{logit: } G(z) = \frac{e^{z}}{1 + e^{z}}.$$
The coefficients are no longer probability changes. The marginal effect is $\beta_1 G'(z)$, which for logit equals $\beta_1 p(1 - p)$: largest when $p$ is near one half, small near zero or one.
Another way: action
Plot mortgage denial (0 or 1) against the payment-to-income ratio. The points form two horizontal lines. A straight line through them eventually leaves the strip between 0 and 1; an S-curve bends to stay inside it.
Another way: steps
Given its flaws, why do so many economists use the LPM? Because in most samples it gives average marginal effects very close to logit and probit, and its coefficients need no translation. A coefficient of $0.18$ on a dummy means eighteen percentage points, full stop.
The LPM also works smoothly with the research designs later in this course — instrumental variables, fixed effects, differences-in-differences — where nonlinear models become awkward. Its main danger is prediction far from the data, where the straight line leaves the unit interval. For estimating the average effect of a policy on the probability of an outcome, near the middle of the data, it is often the clearest choice, and it should always be paired with robust standard errors.
Differentiate the logistic function: $G(z) = e^{z}/(1 + e^{z})$ has derivative $G'(z) = G(z)(1 - G(z)) = p(1 - p)$. By the chain rule, the change in probability for a small change in $x$ is $\beta_1 G'(z) = \beta_1 p(1 - p)$.
This product is largest at $p = 0.5$, where $p(1 - p) = 0.25$, and shrinks toward zero at the extremes. That is the S-curve at work: a nudge to someone on the fence changes their probability a lot; the same nudge to someone who is almost certain either way changes it very little. A single logit coefficient therefore implies different marginal effects for different people, and a report must say at which probability, or averaged over whom, an effect is computed.
Probit behaves the same way with the normal curve in place of the logistic one; its coefficients are smaller by a factor of about $1.6$, but its marginal effects are nearly identical to logit's.
The odds of an event are $p/(1 - p)$: a probability of $0.8$ is odds of $4$ to $1$. In a logit model the log of the odds is exactly the index: $\log\big(p/(1 - p)\big) = \beta_0 + \beta_1 x$. So each unit of $x$ adds $\beta_1$ to the log odds, which multiplies the odds by $e^{\beta_1}$.
Medical and public-health research often reports these odds ratios. They are easy to misread as risk ratios. When an event is rare, odds and probabilities are close, and doubling the odds roughly doubles the risk. When it is common they diverge: doubling odds of $4$ to $8$ moves the probability from $0.8$ to about $0.89$, not to $1.6$. Converting back to probabilities before interpreting is the safe habit.
OLS does not fit logit and probit models. They are estimated by maximum likelihood: choose the coefficients that make the observed pattern of ones and zeros as probable as possible. For each person the model assigns a probability $p_i$ to the outcome that happened — $p_i$ if they were denied, $1 - p_i$ if not — and maximum likelihood picks the coefficients that make the product of those probabilities largest.
Software does this numerically, and the output looks like a regression table: coefficients, standard errors, z statistics. Tests work as before, with $1.96$ as the 5 percent cutoff. There is no $R^2$ in the usual sense; programs report a "pseudo $R^2$" or the share of outcomes predicted correctly, and neither should be compared with an OLS $R^2$.
What maximum likelihood does not change is the identification question. A logit coefficient on a treatment is as biased as an LPM coefficient if the treatment is correlated with omitted determinants of the outcome. The choice between LPM, logit and probit is about the shape of the probability curve, not about causality. A careful study reports the same effect from more than one of the three models and shows that the conclusion does not hinge on the choice. That kind of robustness check is cheap to run and reassures a skeptical reader. It should be routine in any published table.
Because a logit marginal effect differs from person to person, a single summary is needed for a report. Two are common. The marginal effect at the mean evaluates $\beta_1 p(1 - p)$ at the probability of an average person. The average marginal effect computes $\beta_1 p_i(1 - p_i)$ for every person in the sample and averages them.
The average marginal effect is usually preferred: it answers "by how much would the share of ones change on average if everyone's $x$ rose by one unit?", which is the question a policy maker asks. It is also the number that is most directly comparable with an LPM coefficient, and in most applications the two are close. When they are not — typically because many people sit near zero or one — the nonlinear model is telling you something the straight line cannot, and it is worth reporting effects separately for groups with different starting probabilities.
A logit model of homeownership has an income coefficient of $0.5$ per ten thousand dollars. For one household, $e^{z} = 4$.
Three checks.
And an interpretation check: a dummy coefficient in a denial regression is a gap holding the included variables fixed. Whether it reflects discrimination depends on what has been left out.
In 1990 the Federal Reserve Bank of Boston collected detailed data on thousands of mortgage applications to test whether lenders treated black and white applicants differently. The raw denial rates were about $28$ percent for black applicants and $9$ percent for white applicants.
A linear probability model with the payment-to-income ratio and a dummy for black applicants gives roughly $\hat P(deny) = -0.09 + 0.56\,PI + 0.18\,black$. The $0.56$ says a rise in $PI$ of $0.1$ — payments taking ten percent more of income — raises the denial probability by about $5.6$ points. The $0.18$ says that, holding $PI$ fixed, black applicants were about $18$ points more likely to be denied. For a black applicant with $PI = 0.3$, the predicted probability is $-0.09 + 0.168 + 0.18 = 0.258$.
Logit and probit versions give almost the same effects near these probabilities. The researchers then added many more controls — credit history, loan-to-value ratio, employment — and the gap shrank to roughly $8$ points but did not vanish. The study became central to fair-lending enforcement, and it is a model for how a binary-outcome regression is read: in percentage points, holding stated variables fixed, with the omitted-variable question kept in view.
The most common mistake is reading a logit or probit coefficient as the change in probability. It is the change in the index; the probability changes by the coefficient times $G'(z)$, which depends on where the person starts.
A second mistake is treating an odds ratio as a risk ratio. They are close only for rare events.
A third is dismissing the LPM because it can predict outside $[0, 1]$. For average effects near the middle of the data it is usually close to logit and far easier to read.
Read the LPM slope.
$\beta_1 = 0.04$
Per year of schooling.
Read the change in x.
$\Delta x = 3$
Three more years.
Multiply the two.
$0.04 \times 3 = 0.12$
Change in probability.
Convert to points.
$12 \text{ percentage points}$
Times 100.
Check the range.
$\text{baseline } 0.7 \to 0.82$
Still inside $[0, 1]$.
Write the logistic function.
$p = e^{z}/(1 + e^{z})$
The S-curve.
Read the index value.
$e^{z} = 3$
For one person.
Compute the probability.
$3 \div 4 = 0.75$
Three chances in four.
Compute the odds.
$0.75 \div 0.25 = 3$
Equal to $e^{z}$.
Apply an odds ratio of 2.
$\text{odds } 3 \to 6$
One more unit of $x$.
Convert back to probability.
$6 \div 7 \approx 0.86$
Not $1.5$.
Read the logit coefficient.
$\beta_1 = 0.8$
Per unit of $x$.
Take a person on the fence.
$p = 0.5$
Where the S-curve is steepest.
Compute the slope factor.
$0.5 \times 0.5 = 0.25$
Its maximum.
Multiply by the coefficient.
$0.8 \times 0.25 = 0.2$
Twenty points.
Take a nearly certain person.
$p = 0.9$
Near the top of the curve.
Compute the slope factor.
$0.9 \times 0.1 = 0.09$
Much smaller.
Multiply by the coefficient.
$0.8 \times 0.09 = 0.072$
Seven points: same coefficient.
Compute the probability.
$1.5 \div 2.5 = 0.6$
The logistic function.
Compute the slope factor.
$0.6 \times 0.4 = 0.24$
p times one minus p.
Multiply by the coefficient.
A linear probability model of whether a worker is employed gives a coefficient of $0.08$ on years of schooling. By how many percentage points does the predicted probability of employment rise with $4$ more years of schooling?
Answer: percentage points
Complete the worked solution: a logit model predicts a probability of $0.6$ that a customer renews a subscription, and each extra year as a customer multiplies the odds of renewing by $3$. The new probability is $kp / (1 - p + kp)$. Find the chance of not renewing, the scaled term $kp$, and the denominator.
Compute the current chance of not renewing.
$1 - 0.6 =$ a
The complement of the probability.
Multiply the probability by the factor.
$3 \times 0.6 =$ b
The numerator of the new probability.
Add the two to form the denominator.
$\text{complement} + \text{scaled} =$ c
New probability = scaled ÷ this sum.
A linear probability model of loan approval is $\hat P = 0.3 + 0.06\,score$, with the credit score in tens of points above a threshold. Fill in the predicted probability at $score = 5$ and at $score = 10$.
At score 5: lo. At score 10: hi.
For one applicant, a logit model's index satisfies $e^{z} = 3$. What is the predicted probability?
Answer: probability
A logit coefficient is $0.5$ and, for one person, $e^{z} = 4$. Fill in the probability, $p(1 - p)$ and the marginal effect.
| value | |
|---|---|
| predicted probability | |
| p times (1 − p) | |
| marginal effect |
Match each property to the model it describes.
| linear probability model | logit or probit | |
|---|---|---|
| a coefficient is directly a change in probability | ||
| predicted probabilities always lie between 0 and 1 | ||
| can predict a probability of 1.15 | ||
| the effect of x is largest when p is near 0.5 |
A linear probability model of mortgage denial in Boston gives $\hat P(deny) = -0.09 + 0.7\,PI + 0.18\,black$, where $PI$ is the ratio of loan payments to income. Fill in the change in predicted denial probability when $PI$ rises by $0.3$, and the predicted probability for a black applicant with $PI = 0.3$.
Change in denial probability: dp. Predicted probability: p.
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A logit coefficient is $0.5$ and, for one person, $e^{z} = 4$. Fill in the probability, $p(1 - p)$ and the marginal effect.
| value | |
|---|---|
| predicted probability | |
| p times (1 − p) | |
| marginal effect |
You can model a yes-or-no outcome. Explain to someone why one logit coefficient implies different probability changes for different people.
17. Your turn: logit coefficient 0.5 and e^z = 1.5., step 3
$0.5 \times 0.24 = 0.12$
Twelve percentage points.