---
title: "Week 9 exercise: TWFE and event studies"
subtitle: "PP5001 · Martinmas 2026 · Ungraded practice"
format:
  html:
    embed-resources: true
    code-fold: false
    toc: true
execute:
  eval: false
---

Use Cesur, Güneş, Tekin and Ulker's (2023) province-level data to estimate the effect of Turkey's Family Medicine Program on teenage fertility and produce an event-study graph. This exercise is **ungraded**. Add your code and answers, then render to HTML or PDF.

**[Download the Quarto setup notebook](../resources/week-09/twfe-setup.qmd)**

## Data and setup

Use `data_main_analysis.dta` from the replication package supplied in Moodle. Put it beside your notebook. The file has 1,458 observations: 81 provinces observed annually from 2001 to 2018.

```{python}
import pandas as pd
import matplotlib.pyplot as plt
import statsmodels.formula.api as smf

data = pd.read_stata("data_main_analysis.dta", convert_categoricals=False)
```

| Variable | Meaning |
|:--|:--|
| `br_women_15to19K` | Births per 1,000 women aged 15–19 |
| `nuts3` | Province identifier |
| `year` | Calendar year |
| `FMP` | 1 when the program operates in the province, 0 otherwise |
| `rby` | Region-by-year identifier |
| `yeartrend` | Linear time trend |
| `meanwomenpop_ages15to19` | Province's mean population of women aged 15–19, used as a regression weight |

## 1. The policy and identification

**(a)** Briefly explain the policy, how its introduction varies across provinces, and the outcome. What comparison identifies the effect in a TWFE regression?

**(b)** State the parallel-trends assumption in this application. What might make it fail? What do province effects, region-by-year effects, and province-specific trends each account for?

## 2. Replicate Table 2, Panel A, column (2)

The authors include time-varying controls, region-by-year effects, province effects and province-specific linear trends. The following string specifies those terms exactly as in their replication code:

```{python}
controls = (
    "log_unemp_rate + missing_log_unemp_rate + log_percapita_gdp"
    " + log_mvpk + log_percent_high_school + log_percent_college"
    " + logpg_mps + C(rby) + C(nuts3) + C(nuts3):yeartrend"
)
```

The controls measure unemployment, income, vehicles per capita, educational attainment and the governing party's seat share. `missing_log_unemp_rate` identifies observations with missing unemployment information. `C(rby)` includes region-by-year indicators. `C(nuts3):yeartrend` allows each province its own linear trend.

**(a)** Estimate the weighted regression below, then report the FMP coefficient, its province-clustered standard error, a 95% confidence interval and the number of observations.

```{python}
policy = smf.wls(
    "br_women_15to19K ~ FMP + " + controls,
    data=data, weights=data["meanwomenpop_ages15to19"]
).fit(cov_type="cluster", cov_kwds={"groups": data["nuts3"]})

# Create a labelled results table.
```

**(b)** Compare your estimate with the published −1.585. Interpret its units and magnitude. Small differences in standard errors can arise from how software counts redundant fixed-effect terms in its finite-sample correction.

## 3. Construct an event-study graph

The data already contain the authors' event indicators:

| Variables | Event time |
|:--|:--|
| `FMP_min7orbefore` | Seven or more years before introduction |
| `FMP_minus6`, `FMP_minus5`, `FMP_minus4` | Six, five and four years before |
| `FMP_0` | Year of introduction |
| `FMP_1`, `FMP_2`, `FMP_3`, `FMP_4` | One through four years after |
| `FMP_5up` | Five or more years after |

For Figure 4, the authors omit years −3, −2 and −1. Their specification combines universal eventual adoption with province-specific linear trends and imposes these additional reference-period restrictions. Use their specification here. Its normalization differs from our introductory example with only year −1 omitted.

**(a)** Replace `FMP` in the regression with the ten indicators above, keeping the same controls, weights and clustered standard errors. Estimate this regression separately from Question 2.

```{python}
# Estimate the event-study regression.
```

**(b)** Create a graph of the ten estimated coefficients with 95% confidence intervals. Put event time on the horizontal axis. Label the endpoint bins “≤−7” and “≥5”. Mark zero on the vertical axis and the program's introduction on the horizontal axis. Indicate that −3 through −1 are omitted reference years. Compare the shape with **Figure 4, Panel B**.

To extract selected coefficients and confidence intervals from a fitted model named `event_study`:

```{python}
terms = [
    "FMP_min7orbefore", "FMP_minus6", "FMP_minus5", "FMP_minus4",
    "FMP_0", "FMP_1", "FMP_2", "FMP_3", "FMP_4", "FMP_5up"
]
coefficients = event_study.params.loc[terms]
intervals = event_study.conf_int().loc[terms]

# Use these values to construct your graph.
```

**(c)** Describe the pre-treatment coefficients and the evolution of the post-treatment coefficients. What evidence do the pre-treatment coefficients provide about the design? How precise is that evidence?

**(d)** Explain why calendar-year controls remain in an event-study regression. What changes when the graph's horizontal axis uses event time?

## 4. Policy interpretation

Write a short conclusion drawing on your estimate and graph. Explain the identifying assumption, uncertainty and the population to which the results apply. All provinces eventually adopt. What comparison groups are available early in the rollout, and which are available once every province has adopted? We will examine the implications for estimation in Week 10.

## Reference

Cesur, R., P. M. Güneş, E. Tekin and A. Ulker (2023). [Socialized Healthcare and Women's Fertility Decisions](https://doi.org/10.3368/jhr.59.1.0419-10155R3). *Journal of Human Resources*, 58(3), 1028–1055. Table 2, Panel A and Figure 4, Panel B.
