PP5001 · Week 10 · Martinmas 2026
Y_{it}=\alpha_i+\lambda_t+\delta D_{it}+u_{it}.
With different adoption dates, which comparisons determine \widehat\delta?
Policies often begin in different places at different times:
A cohort is a group of units first treated in the same calendar period.
Timing groups, from the original lecture slides.
For discussion: Identify the early-treated, late-treated and never-treated groups. When can each group supply untreated observations?
With early, late and never-treated groups, TWFE combines:
In comparison 4, the comparison group’s outcomes already include treatment effects.
For the basic TWFE regression in a balanced panel,
\widehat\delta^{TWFE}=\sum_j s_j\widehat\delta_j^{2\times2}, \qquad s_j\ge0,\qquad \sum_j s_j=1.
Suppose untreated outcomes follow parallel trends and there is no anticipation.
If treatment has the same constant effect for every treated cohort and period, an already-treated group’s treatment effect cancels when its outcome is differenced.
The difficulty arises when that effect changes over the comparison period.
Suppose every group’s untreated outcome would remain at 10.
| Group | Period 0 | Period 1 | Period 2 |
|---|---|---|---|
| Early treated | 10 | 12 | 16 |
| Late treated | 10 | 10 | 12 |
| Never treated | 10 | 10 | 10 |
Early adopters have an effect of 2 initially and 6 later. Late adopters have an effect of 2 when they adopt.
For discussion: From period 1 to period 2, whose observed outcome changes most? What would happen if we used that group as the control for late adopters?
From period 0 to period 1:
\widehat\delta_{Early,Late}=(12-10)-(10-10)=2.
From period 1 to period 2, using never-treated units:
\widehat\delta_{Late,Never}=(12-10)-(10-10)=2.
Each calculation recovers the newly treated group’s effect.
Compare late with early adopters from period 1 to period 2:
\widehat\delta_{Late,Early}=(12-10)-(16-12)=-2.
The late group’s treatment effect is +2, but this comparison subtracts the +4 change in the early group’s treatment effect.
For discussion: Why does a negative comparison arise even though every treatment effect is positive?
Under parallel trends in untreated outcomes,
DiD_{Late,Early} =\text{late group's new effect} -\text{change in early group's effect}.
Growing effects in the early cohort reduce this component of TWFE.
The weights on the component 2\times2 comparisons are positive.
But an already-treated comparison subtracts changes in treatment effects. When TWFE is expressed in terms of underlying cohort-period effects, some of those effects can receive negative weight.
A TWFE coefficient can consequently lie outside the range of the underlying treatment effects.
In a conventional TWFE event study with staggered adoption:
Even an apparent pre-treatment coefficient can reflect post-treatment effects elsewhere.
The approach we will use has three steps:
Callaway and Sant’Anna (2021) organise their method around these steps.
Let G_i=g mean that unit i first receives treatment in period g.
ATT(g,t)=E\left[Y_{it}^{g}-Y_{it}^{\infty}\mid G_i=g\right],\qquad t\ge g.
ATT(2004,2006)
means the effect in calendar year 2006, for units first treated in 2004.
Their event time is
k=t-g=2006-2004=2.
For discussion: Which effect would describe the 2006 cohort in its first treatment year? Which would describe that same cohort one year later?
Use g-1, the last pre-treatment period, as the baseline.
With never-treated units as controls, assume
E\left[Y_{it}^{\infty}-Y_{i,g-1}^{\infty}\mid G_i=g\right] =E\left[Y_{it}^{\infty}-Y_{i,g-1}^{\infty}\mid G_i=\infty\right].
Also assume no anticipation at the baseline.
Under those assumptions,
\begin{aligned} ATT(g,t)={}&E\left[Y_{it}-Y_{i,g-1}\mid G_i=g\right]\\ &-E\left[Y_{it}-Y_{i,g-1}\mid G_i=\infty\right]. \end{aligned}
The sample estimator replaces each expectation with its corresponding sample mean.
We already know how to calculate this comparison.
Never treated: remain untreated throughout the observation period.
Not yet treated: remain untreated through the outcome period t.
For the latter, use units with G_i>t, including never-treated units where available, and maintain the corresponding parallel-trends and no-anticipation assumptions.
Sometimes parallel trends is plausible only among units with similar baseline characteristics X_i.
Estimate comparisons conditional on those characteristics, then average over the treated cohort’s characteristics.
Overlap is needed: the comparison group must contain suitable counterparts for the treated cohort.
Once we have ATT(g,t), we can average:
These averages can differ because they combine different effects with different weights.
At event time k\ge0, combine available cohorts at t=g+k:
\theta_k=\sum_{g\in\mathcal G_k}w_{g,k}\,ATT(g,g+k), \qquad w_{g,k}\ge0,\qquad \sum_{g\in\mathcal G_k}w_{g,k}=1.
\mathcal G_k contains cohorts observed at event time k with a valid comparison group. Cohort-size weights give larger cohorts more influence.
Suppose data end in 2007.
| First treated | Last observed event time |
|---|---|
| 2004 | 3 |
| 2006 | 1 |
| 2007 | 0 |
An event-time graph can change because effects evolve and because the contributing cohorts change.
A balanced-cohort graph keeps the same cohorts at every displayed exposure length.
Callaway and Sant’Anna study county-level teenage employment during 2001–2007.
For discussion: Explain the policy, data and source of identifying variation. What differs from Card and Krueger’s comparison?
Callaway and Sant’Anna (2021), Table 2, p. 217.
For discussion: Describe the differences between treated and untreated counties. What do these differences imply for the parallel-trends argument and the choice of controls?
Callaway and Sant’Anna (2021), Figure 1(a), p. 218.
For discussion: Identify the cohorts and axes. Interpret one post-treatment estimate. What concerns arise from the pre-treatment evidence?
Callaway and Sant’Anna (2021), Figure 1(b), p. 218.
For discussion: What changes after adjustment for baseline characteristics? How convincing is the identifying assumption now?
Callaway and Sant’Anna (2021), Table 3(a), p. 219.
For discussion: Focus on TWFE and the group-specific effects. Compare the TWFE estimate with the overall group-specific effect in the rightmost column. How similar are they?
Callaway and Sant’Anna (2021), Table 3(b), p. 219.
For discussion: Compare TWFE with the overall group-specific effect in the rightmost column. What changes relative to panel (a) after covariate adjustment? Would your policy conclusion depend on the estimator?
For discussion: Which of these choices would matter most if you were advising a government about another minimum-wage increase?
The basic DiD logic remains: use an untreated group’s change to construct a counterfactual.
With staggered adoption, make the comparison separately for each cohort and period, then average deliberately.
Next week: constructing a comparison for a single treated country or region using synthetic control.
Callaway, B., and P. H. C. Sant’Anna (2021). Difference-in-Differences with Multiple Time Periods. Journal of Econometrics, 225(2), 200–230.
Goodman-Bacon, A. (2021). Difference-in-Differences with Variation in Treatment Timing. Journal of Econometrics, 225(2), 254–277.
Sun, L., and S. Abraham (2021). Estimating Dynamic Treatment Effects in Event Studies with Heterogeneous Treatment Effects. Journal of Econometrics, 225(2), 175–199.