Week 10 exercise: Modern difference-in-differences

PP5001 · Martinmas 2026 · Ungraded practice

Use the authors’ public teaching subset to estimate TWFE and construct group-time treatment effects from changes in means. This exercise is ungraded. Add code and explanations, then render to HTML or PDF.

Download the setup notebook · Download the data

Setup

Save mpdta.csv beside your notebook. The authors’ dataset documentation describes this subset: 500 counties observed annually in 2003–2007. The published application has a larger sample and a longer period. Your estimates therefore describe this teaching sample.

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import statsmodels.formula.api as smf

data = pd.read_csv("mpdta.csv").rename(columns={"first.treat": "first_treat"})
Variable Meaning
countyreal County identifier
year Calendar year
lemp Log teenage employment
lpop Log population in thousands
first_treat First treatment year, with 0 for never-treated counties
treat Whether the county ever receives treatment during the sample

1. Treatment status and event time

(a) Count counties in each treatment cohort and the never-treated group. Report adoption years and available post-treatment years.

(b) Create current treatment status:

data["D"] = ((data["first_treat"] > 0) &
             (data["year"] >= data["first_treat"])).astype(int)

Why would using treat instead of D in the TWFE regression be a mistake?

(c) Draw the mean of lemp against calendar year for each cohort and the never-treated group. Use different markers or line styles as well as colours. Describe the comparisons suggested by this graph.

2. TWFE

Estimate the regression below. Report the treatment coefficient, its standard error and a 95% confidence interval.

twfe = smf.ols(
    "lemp ~ D + C(countyreal) + C(year)", data=data
).fit(
    cov_type="cluster",                       # Cluster the standard errors.
    cov_kwds={"groups": data["countyreal"]}    # County identifies each cluster.
)

Observations from the same county in different years may share common shocks. Clustered standard errors allow for this dependence when measuring uncertainty about the estimates. In the code, cov_type="cluster" requests clustered standard errors and groups specifies which observations belong to the same county. This changes the standard errors and confidence intervals. The estimated coefficients stay the same.

Interpret the coefficient in percentage terms. What concerns arise if treatment effects vary by cohort or exposure length?

3. Construct two group-time comparisons

Focus on the cohort first treated in 2004 and use never-treated counties as its comparison group. Its baseline is 2003.

Reshape the outcome so each county has one row:

outcomes = data.pivot(index="countyreal", columns="year", values="lemp")
cohorts = data.groupby("countyreal")["first_treat"].first()
comparison = outcomes.join(cohorts)
comparison = comparison.loc[comparison["first_treat"].isin([0, 2004])].copy()
comparison["cohort2004"] = (comparison["first_treat"] == 2004).astype(int)

(a) Calculate the change from 2003 to 2004 for each county. Compute the difference between the treated and comparison groups’ mean changes. This estimates \(ATT(2004,2004)\) under parallel trends and no anticipation.

(b) Repeat using the change from 2003 to 2007. This estimates \(ATT(2004,2007)\).

(c) Estimate each comparison by regressing the county-level outcome change on cohort2004, using HC1 standard errors. Estimate the two regressions separately. Check that the coefficients equal your differences in mean changes.

(d) Present the two estimates and confidence intervals together. Interpret the two different exposure lengths. Is either estimate intended to measure the same average as the TWFE coefficient?

4. Who can supply the comparison?

Could counties first treated in 2006 be used as untreated controls for \(ATT(2004,2004)\)? Could they be used for \(ATT(2004,2007)\)? Explain the assumptions and why the answers differ.

5. Read the published evidence

Use Callaway and Sant’Anna’s Table 3 and Figure 1.

(a) Explain why the event-study average and the balanced-group event-study average differ. Which cohorts contribute at longer exposure lengths?

(b) Compare the conditional TWFE estimate with the overall group-specific estimate. Explain why the two procedures can yield different answers.

(c) Identify evidence that makes you cautious about a causal interpretation. Would changing the estimator settle those concerns?

6. Critique Claude

Ask Claude:

A minimum-wage study uses Callaway and Sant’Anna’s estimator instead of two-way fixed effects. Does this make its estimated employment effects credible for policy? Explain briefly.

Evaluate the response using the published application and your own calculations. Identify one claim that needs qualification and explain why.

Sources

Callaway, B., and P. H. C. Sant’Anna (2021). Difference-in-Differences with Multiple Time Periods. Journal of Econometrics, 225(2), 200–230.

Data: mpdta from the authors’ did package. The supplied CSV preserves the variables and observations in the package’s R data file.