---
title: "Week 8 exercise: Difference-in-Differences"
subtitle: "PP5001 · Martinmas 2026 · Ungraded practice"
format:
  html:
    embed-resources: true
    code-fold: false
    toc: true
execute:
  eval: false
jupyter: python3
---

Replicate the basic employment comparison in Card and Krueger (1994). Calculate the difference-in-differences from group means, draw the counterfactual, and estimate the same effect by regression. This post-class exercise is **ungraded**. Add your code and written answers below, then render your notebook to HTML or PDF.

## Data and setup

Download the [New Jersey–Pennsylvania data archive from David Card's website](https://eml.berkeley.edu/~card/data_sets/njmin.zip). Unzip it and put `public.dat` in the same folder as this notebook. The archive includes a `codebook` explaining the variables.

The data contain 410 restaurants interviewed around New Jersey's April 1992 minimum-wage increase. Each row is a restaurant. Variables ending in `2` refer to the second interview. The code below imports the original text file and gives its columns the names in the codebook. A dot in the original file represents a missing value.

```{python}
import pandas as pd
import matplotlib.pyplot as plt
import statsmodels.formula.api as smf

columns = [
    "sheet", "chain", "co_owned", "state", "southj", "centralj",
    "northj", "pa1", "pa2", "shore", "ncalls", "empft", "emppt",
    "nmgrs", "wage_st", "inctime", "firstinc", "bonus", "pctaff",
    "meals", "open", "hrsopen", "psoda", "pfry", "pentree", "nregs",
    "nregs11", "type2", "status2", "date2", "ncalls2", "empft2",
    "emppt2", "nmgrs2", "wage_st2", "inctime2", "firstin2",
    "special2", "meals2", "open2r", "hrsopen2", "psoda2", "pfry2",
    "pentree2", "nregs2", "nregs112"
]

raw = pd.read_csv("public.dat", sep=r"\s+", header=None,
                  names=columns, na_values=".")
```

| Variable | Meaning |
|:--|:--|
| `sheet` | Restaurant identifier |
| `state` | 1 for New Jersey, 0 for Pennsylvania |
| `empft`, `empft2` | Full-time employees before and after |
| `emppt`, `emppt2` | Part-time employees before and after |
| `nmgrs`, `nmgrs2` | Managers and assistant managers before and after |
| `status2` | Second-interview status. Code 3 denotes permanent closure. |

## 1. The policy and the comparison

**(a)** Identify the treatment and control groups, the policy change, and the two observation periods. Why might eastern Pennsylvania provide a useful comparison for New Jersey?

**(b)** State the parallel-trends assumption in terms of employment at these restaurants. Explain why a difference in employment levels before the policy would not, by itself, invalidate this assumption.

## 2. Four means and two differences

**(a)** Construct full-time-equivalent employment in each period as full-time employees plus managers plus half the number of part-time employees:

```{python}
raw["fte_before"] = raw["empft"] + raw["nmgrs"] + 0.5 * raw["emppt"]
raw["fte_after"] = raw["empft2"] + raw["nmgrs2"] + 0.5 * raw["emppt2"]
```

Permanently closed restaurants already have zero employment in the second interview. Retain those zeros. Leave missing values as missing.

Using all available observations in each period, make a table showing mean employment and the number of observations for each state and period. Compare the means with **Table 3, rows 1–2, columns (i)–(ii)**. Calculate the employment change in each state and the difference-in-differences. Compare it with **row 3, column (iii)**.

```{python}
# Your table and calculations.
```

**(b)** Now retain restaurants with employment observed in both periods. This gives the same restaurants before and after:

```{python}
data = raw.dropna(subset=["fte_before", "fte_after"]).copy()
```

Report the number of restaurants in each state. Recalculate the four means and the difference-in-differences. Compare the changes with **Table 3, row 4, columns (i)–(iii)**. Explain why this calculation can differ from part (a).

```{python}
# Your balanced-sample table and calculations.
```

Use this same sample for Questions 3–4.

## 3. Draw the counterfactual

Plot the two states' mean employment before and after the policy, joining each state's points with a line. Label the axes and distinguish the states using different markers as well as colours.

Add a dashed line from New Jersey's initial mean to its predicted employment after the policy **if its employment had changed by the same amount as Pennsylvania's**. Label the gap between this counterfactual and New Jersey's observed post-policy mean.

Explain how the graph represents the parallel-trends assumption. With only one pre-policy observation per restaurant, what can these data tell us about trends before the increase?

```{python}
# Your graph and interpretation.
```

## 4. Estimate the same effect by regression

**(a)** Reshape the balanced sample so that there are two rows per restaurant, one before and one after:

```{python}
before = data[["sheet", "state", "fte_before"]].rename(
    columns={"fte_before": "fte"}
)
before["post"] = 0

after = data[["sheet", "state", "fte_after"]].rename(
    columns={"fte_after": "fte"}
)
after["post"] = 1

long = pd.concat([before, after], ignore_index=True)
```

Estimate

$$
Y_{it}=\beta_0+\beta_1 NJ_i+\beta_2 Post_t
       +\beta_3\left(NJ_i\times Post_t\right)+u_{it}.
$$

In Python, `"fte ~ state * post"` includes the two indicators and their interaction. Report all four coefficients and explain how each relates to your four means. Verify that the interaction coefficient equals the difference-in-differences you calculated.

The two observations for a restaurant may have correlated disturbances. Use standard errors clustered by restaurant:

```{python}
did = smf.ols("fte ~ state * post", data=long).fit(
    cov_type="cluster", cov_kwds={"groups": long["sheet"]}
)
```

These standard errors allow arbitrary correlation within a restaurant and heteroskedasticity across restaurants. They still assume independence across restaurants. State-wide economic shocks are therefore an important limitation when interpreting uncertainty in a comparison of two states.

**(b)** Calculate each restaurant's change in employment and regress that change on an intercept and the New Jersey indicator. Use HC1 standard errors. Verify that the New Jersey coefficient equals the interaction coefficient in part (a).

```{python}
# Calculate the change and estimate the regression.
```

**(c)** Present the two treatment-effect estimates, standard errors, confidence intervals, and sample sizes in a labelled table. Distinguish the number of rows from the number of restaurants. Explain why the point estimates agree. Small differences in the standard errors can arise from the finite-sample corrections used by the two procedures.

## 5. Interpret the evidence

Write a short policy conclusion. Explain the effect's units and the population and period to which it applies. Identify one plausible threat to parallel trends and discuss what further evidence would help assess it.

Then ask Claude:

> In Card and Krueger's New Jersey–Pennsylvania minimum-wage study, the difference-in-differences estimate is positive. Does this establish that raising the minimum wage increases employment? Explain briefly.

Evaluate its response using your results and the identifying assumption. Quote one claim you would retain or revise and explain your reasoning.

## 6. E-ZPass and air pollution

Read **Currie and Walker (2011), Table 7, columns (1)–(2)** and the accompanying discussion. These regressions examine pollution around the introduction of E-ZPass.

**(a)** Identify the treated and comparison monitors. What is the outcome in each column?

**(b)** Interpret the treatment-by-post coefficient in column (1), including its units. Why is the second pollutant a useful comparison?

**(c)** There is only one monitor within 2 km of a toll plaza. Explain why many daily observations do not provide the same evidence as many independently treated locations.

**(d)** How do these results support the proposed explanation for the infant-health effects? What additional assumptions would be required to estimate the causal effect of pollution on infant health using E-ZPass as an instrument?

## Readings

Card, D., and A. B. Krueger (1994). [Minimum Wages and Employment: A Case Study of the Fast-Food Industry in New Jersey and Pennsylvania](https://eml.berkeley.edu/~card/papers/njmin-aer.pdf). *American Economic Review*, 84(4), 772–793.

Currie, J., and R. Walker (2011). [Traffic Congestion and Infant Health: Evidence from E-ZPass](https://doi.org/10.1257/app.3.1.65). *American Economic Journal: Applied Economics*, 3(1), 65–90.
