Back to the on-screen lesson ·

Binary outcomes: linear probability, logit and probit

Model a yes-or-no outcome with a straight line or an S-curve, read LPM slopes in percentage points, and compute logit probabilities, marginal effects and odds.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you will be able to read a linear probability model, compute a logit probability and marginal effect, and work with odds.

2. What you already have

You can read a regression with dummies on the right-hand side, and from the second lesson you know that a regression predicts a conditional mean. From statistics you know that the mean of a variable that is either zero or one is the share of ones: the probability.

Many economic outcomes are yes or no. Is a mortgage denied? Does a person work? Does a customer renew? Does a firm export? When the dependent variable is a dummy, a regression predicts a probability, and this lesson shows the three standard ways to model it and how to read each.

3. Terms to use precisely

TermWhat it means
Binary outcomeA dependent variable that equals 1 or 0.
Linear probability modelOLS with a binary $y$: $P(y = 1 \mid x) = \beta_0 + \beta_1 x$.
Logit$P = e^{z}/(1 + e^{z})$ with index $z = \beta_0 + \beta_1 x$.
Probit$P = \Phi(z)$, the standard normal cumulative distribution of the index.
IndexThe linear combination $z = \beta_0 + \beta_1 x$ inside a logit or probit.
Marginal effectThe change in probability for a small change in $x$: $\beta_1 G'(z)$.
Odds$p/(1 - p)$; in a logit, $e^{\beta_1}$ is the factor by which the odds change per unit of $x$.

4. A straight line or an S-curve for a probability

With a binary outcome, $E[y \mid x] = P(y = 1 \mid x)$. The linear probability model (LPM) fits that probability with OLS:

$$P(y = 1 \mid x) = \beta_0 + \beta_1 x.$$

Its slope is directly a change in probability — multiply by 100 for percentage points — and everything from earlier lessons carries over. Its drawbacks: a straight line can predict probabilities below zero or above one, and it assumes the effect of $x$ is the same whether the probability starts at $0.05$ or $0.5$. Its errors are also heteroskedastic by construction, so robust standard errors are required.

Logit and probit pass the index $z = \beta_0 + \beta_1 x$ through an S-shaped function $G$ that stays between zero and one:

$$P(y = 1 \mid x) = G(z), \qquad \text{logit: } G(z) = \frac{e^{z}}{1 + e^{z}}.$$

The coefficients are no longer probability changes. The marginal effect is $\beta_1 G'(z)$, which for logit equals $\beta_1 p(1 - p)$: largest when $p$ is near one half, small near zero or one.

Another way: action

Plot mortgage denial (0 or 1) against the payment-to-income ratio. The points form two horizontal lines. A straight line through them eventually leaves the strip between 0 and 1; an S-curve bends to stay inside it.

Another way: steps

  1. LPM: read $\beta_1 \Delta x$ as a change in probability.
  2. Logit: compute the index $z$, then $p = e^{z}/(1 + e^{z})$.
  3. Marginal effect: $\beta_1 p(1 - p)$ at the $p$ of interest.
  4. Odds: $p/(1 - p)$; each unit of $x$ multiplies them by $e^{\beta_1}$.
  5. Report effects in percentage points at a stated starting probability.

5. Why the LPM is still widely used

Given its flaws, why do so many economists use the LPM? Because in most samples it gives average marginal effects very close to logit and probit, and its coefficients need no translation. A coefficient of $0.18$ on a dummy means eighteen percentage points, full stop.

The LPM also works smoothly with the research designs later in this course — instrumental variables, fixed effects, differences-in-differences — where nonlinear models become awkward. Its main danger is prediction far from the data, where the straight line leaves the unit interval. For estimating the average effect of a policy on the probability of an outcome, near the middle of the data, it is often the clearest choice, and it should always be paired with robust standard errors.

6. Why the logit marginal effect is beta times p times one minus p

Differentiate the logistic function: $G(z) = e^{z}/(1 + e^{z})$ has derivative $G'(z) = G(z)(1 - G(z)) = p(1 - p)$. By the chain rule, the change in probability for a small change in $x$ is $\beta_1 G'(z) = \beta_1 p(1 - p)$.

This product is largest at $p = 0.5$, where $p(1 - p) = 0.25$, and shrinks toward zero at the extremes. That is the S-curve at work: a nudge to someone on the fence changes their probability a lot; the same nudge to someone who is almost certain either way changes it very little. A single logit coefficient therefore implies different marginal effects for different people, and a report must say at which probability, or averaged over whom, an effect is computed.

Probit behaves the same way with the normal curve in place of the logistic one; its coefficients are smaller by a factor of about $1.6$, but its marginal effects are nearly identical to logit's.

7. Odds and odds ratios

The odds of an event are $p/(1 - p)$: a probability of $0.8$ is odds of $4$ to $1$. In a logit model the log of the odds is exactly the index: $\log\big(p/(1 - p)\big) = \beta_0 + \beta_1 x$. So each unit of $x$ adds $\beta_1$ to the log odds, which multiplies the odds by $e^{\beta_1}$.

Medical and public-health research often reports these odds ratios. They are easy to misread as risk ratios. When an event is rare, odds and probabilities are close, and doubling the odds roughly doubles the risk. When it is common they diverge: doubling odds of $4$ to $8$ moves the probability from $0.8$ to about $0.89$, not to $1.6$. Converting back to probabilities before interpreting is the safe habit.

8. How logit and probit are estimated

OLS does not fit logit and probit models. They are estimated by maximum likelihood: choose the coefficients that make the observed pattern of ones and zeros as probable as possible. For each person the model assigns a probability $p_i$ to the outcome that happened — $p_i$ if they were denied, $1 - p_i$ if not — and maximum likelihood picks the coefficients that make the product of those probabilities largest.

Software does this numerically, and the output looks like a regression table: coefficients, standard errors, z statistics. Tests work as before, with $1.96$ as the 5 percent cutoff. There is no $R^2$ in the usual sense; programs report a "pseudo $R^2$" or the share of outcomes predicted correctly, and neither should be compared with an OLS $R^2$.

What maximum likelihood does not change is the identification question. A logit coefficient on a treatment is as biased as an LPM coefficient if the treatment is correlated with omitted determinants of the outcome. The choice between LPM, logit and probit is about the shape of the probability curve, not about causality. A careful study reports the same effect from more than one of the three models and shows that the conclusion does not hinge on the choice. That kind of robustness check is cheap to run and reassures a skeptical reader. It should be routine in any published table.

9. Average marginal effects

Because a logit marginal effect differs from person to person, a single summary is needed for a report. Two are common. The marginal effect at the mean evaluates $\beta_1 p(1 - p)$ at the probability of an average person. The average marginal effect computes $\beta_1 p_i(1 - p_i)$ for every person in the sample and averages them.

The average marginal effect is usually preferred: it answers "by how much would the share of ones change on average if everyone's $x$ rose by one unit?", which is the question a policy maker asks. It is also the number that is most directly comparable with an LPM coefficient, and in most applications the two are close. When they are not — typically because many people sit near zero or one — the nonlinear model is telling you something the straight line cannot, and it is worth reporting effects separately for groups with different starting probabilities.

10. Working a logit effect, step by step

A logit model of homeownership has an income coefficient of $0.5$ per ten thousand dollars. For one household, $e^{z} = 4$.

  1. Probability. $p = 4/5 = 0.8$.
  2. Slope factor. $p(1 - p) = 0.8 \times 0.2 = 0.16$.
  3. Marginal effect. $0.5 \times 0.16 = 0.08$.
  4. In points. Eight percentage points per ten thousand dollars.
  5. Compare. For a household on the fence with $p = 0.5$, the effect is $0.5 \times 0.25 = 0.125$: over twelve points. The same coefficient, a larger effect.

11. How to check a binary-outcome answer

Three checks.

  1. Is every probability between zero and one? From logit or probit it must be; from an LPM it may not be, and that is worth reporting.
  2. Is the logit marginal effect smaller than the coefficient? Since $p(1 - p) \le 0.25$, it is at most a quarter of $\beta_1$.
  3. Are the units percentage points or probability? A change of $0.08$ is eight points; write which.

And an interpretation check: a dummy coefficient in a denial regression is a gap holding the included variables fixed. Whether it reflects discrimination depends on what has been left out.

12. In the world: mortgage denials in Boston

In 1990 the Federal Reserve Bank of Boston collected detailed data on thousands of mortgage applications to test whether lenders treated black and white applicants differently. The raw denial rates were about $28$ percent for black applicants and $9$ percent for white applicants.

A linear probability model with the payment-to-income ratio and a dummy for black applicants gives roughly $\hat P(deny) = -0.09 + 0.56\,PI + 0.18\,black$. The $0.56$ says a rise in $PI$ of $0.1$ — payments taking ten percent more of income — raises the denial probability by about $5.6$ points. The $0.18$ says that, holding $PI$ fixed, black applicants were about $18$ points more likely to be denied. For a black applicant with $PI = 0.3$, the predicted probability is $-0.09 + 0.168 + 0.18 = 0.258$.

Logit and probit versions give almost the same effects near these probabilities. The researchers then added many more controls — credit history, loan-to-value ratio, employment — and the gap shrank to roughly $8$ points but did not vanish. The study became central to fair-lending enforcement, and it is a model for how a binary-outcome regression is read: in percentage points, holding stated variables fixed, with the omitted-variable question kept in view.

13. A logit coefficient is not a change in probability

The most common mistake is reading a logit or probit coefficient as the change in probability. It is the change in the index; the probability changes by the coefficient times $G'(z)$, which depends on where the person starts.

A second mistake is treating an odds ratio as a risk ratio. They are close only for rare events.

A third is dismissing the LPM because it can predict outside $[0, 1]$. For average effects near the middle of the data it is usually close to logit and far easier to read.

14. An LPM change in probability

  1. Read the LPM slope.

    $\beta_1 = 0.04$

    Per year of schooling.

  2. Read the change in x.

    $\Delta x = 3$

    Three more years.

  3. Multiply the two.

    $0.04 \times 3 = 0.12$

    Change in probability.

  4. Convert to points.

    $12 \text{ percentage points}$

    Times 100.

  5. Check the range.

    $\text{baseline } 0.7 \to 0.82$

    Still inside $[0, 1]$.

15. A logit probability and its odds

  1. Write the logistic function.

    $p = e^{z}/(1 + e^{z})$

    The S-curve.

  2. Read the index value.

    $e^{z} = 3$

    For one person.

  3. Compute the probability.

    $3 \div 4 = 0.75$

    Three chances in four.

  4. Compute the odds.

    $0.75 \div 0.25 = 3$

    Equal to $e^{z}$.

  5. Apply an odds ratio of 2.

    $\text{odds } 3 \to 6$

    One more unit of $x$.

  6. Convert back to probability.

    $6 \div 7 \approx 0.86$

    Not $1.5$.

16. Marginal effects at two starting points

  1. Read the logit coefficient.

    $\beta_1 = 0.8$

    Per unit of $x$.

  2. Take a person on the fence.

    $p = 0.5$

    Where the S-curve is steepest.

  3. Compute the slope factor.

    $0.5 \times 0.5 = 0.25$

    Its maximum.

  4. Multiply by the coefficient.

    $0.8 \times 0.25 = 0.2$

    Twenty points.

  5. Take a nearly certain person.

    $p = 0.9$

    Near the top of the curve.

  6. Compute the slope factor.

    $0.9 \times 0.1 = 0.09$

    Much smaller.

  7. Multiply by the coefficient.

    $0.8 \times 0.09 = 0.072$

    Seven points: same coefficient.

17. Your turn: logit coefficient 0.5 and e^z = 1.5.

  1. Compute the probability.

    $1.5 \div 2.5 = 0.6$

    The logistic function.

  2. Compute the slope factor.

    $0.6 \times 0.4 = 0.24$

    p times one minus p.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Multiply by the coefficient.

18. Guided practice

A linear probability model of whether a worker is employed gives a coefficient of $0.08$ on years of schooling. By how many percentage points does the predicted probability of employment rise with $4$ more years of schooling?

Answer: percentage points

19. Guided practice

Complete the worked solution: a logit model predicts a probability of $0.6$ that a customer renews a subscription, and each extra year as a customer multiplies the odds of renewing by $3$. The new probability is $kp / (1 - p + kp)$. Find the chance of not renewing, the scaled term $kp$, and the denominator.

  1. Compute the current chance of not renewing.

    $1 - 0.6 =$ a

    The complement of the probability.

  2. Multiply the probability by the factor.

    $3 \times 0.6 =$ b

    The numerator of the new probability.

  3. Add the two to form the denominator.

    $\text{complement} + \text{scaled} =$ c

    New probability = scaled ÷ this sum.

20. Guided practice

A linear probability model of loan approval is $\hat P = 0.3 + 0.06\,score$, with the credit score in tens of points above a threshold. Fill in the predicted probability at $score = 5$ and at $score = 10$.

At score 5: lo. At score 10: hi.

21. Practice

For one applicant, a logit model's index satisfies $e^{z} = 3$. What is the predicted probability?

Answer: probability

22. Practice

A logit coefficient is $0.5$ and, for one person, $e^{z} = 4$. Fill in the probability, $p(1 - p)$ and the marginal effect.

value
predicted probability
p times (1 − p)
marginal effect

23. Practice

Match each property to the model it describes.

linear probability modellogit or probit
a coefficient is directly a change in probability
predicted probabilities always lie between 0 and 1
can predict a probability of 1.15
the effect of x is largest when p is near 0.5

24. Somewhere new

A linear probability model of mortgage denial in Boston gives $\hat P(deny) = -0.09 + 0.7\,PI + 0.18\,black$, where $PI$ is the ratio of loan payments to income. Fill in the change in predicted denial probability when $PI$ rises by $0.3$, and the predicted probability for a black applicant with $PI = 0.3$.

Change in denial probability: dp. Predicted probability: p.

25. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

26. Test question

A logit coefficient is $0.5$ and, for one person, $e^{z} = 4$. Fill in the probability, $p(1 - p)$ and the marginal effect.

value
predicted probability
p times (1 − p)
marginal effect

27. What you can do now

You can model a yes-or-no outcome. Explain to someone why one logit coefficient implies different probability changes for different people.

Working for the steps left to you

17. Your turn: logit coefficient 0.5 and e^z = 1.5., step 3

$0.5 \times 0.24 = 0.12$

Twelve percentage points.