---
title: "Week 11 exercise: Synthetic West Germany"
subtitle: "PP5001 · Martinmas 2026 · Ungraded practice"
format:
  html:
    embed-resources: true
    toc: true
execute:
  eval: true
---

Use the German reunification data to construct synthetic West Germany, interpret the estimated gap, and assess sensitivity. This exercise is **ungraded**. Work in Quarto and include your code, figures, tables, and explanations.

[Download the data](https://www.djaeger.org/teaching/mpp/pp5001/resources/week-11/germany.csv)

Save the notebook and `germany.csv` in the same folder. The supplied fitting code uses `pysyncon`, tested with version 1.7.0. Install it once in the Python environment used by Quarto:

```bash
python -m pip install pysyncon==1.7.0
```

The data and fitting specification come from the package's [German reunification example](https://github.com/sdfordham/pysyncon/blob/main/examples/germany.ipynb), based on Abadie, Diamond, and Hainmueller (2015). This implementation produces weights close to the published table. Numerical fitting choices can produce some differences in the weights and predictor matches.

## 1. Policy, data, and comparison countries

**(a)** State the policy question. Explain why the outcome is West German GDP per capita and what the counterfactual represents.

**(b)** Inspect the data. How many countries and years are included? Which countries supply the donor pool?

**(c)** Identify one reason why a donor might provide a useful comparison and one concern about spillovers from reunification.

### Variables

| Variable | Description |
|:--|:--|
| `country`, `code`, `year` | Country name, country identifier, and calendar year |
| `gdp` | GDP per capita in PPP **current** US dollars, following the erratum |
| `trade` | Trade openness |
| `infrate` | Inflation rate |
| `industry` | Industry share |
| `schooling` | Schooling measure, observed at five-year intervals |
| `invest70`, `invest80` | Prepared investment-rate averages used in the two fitting stages |

The data also contain `invest60`, which is not used below. Missing values in the predictor series reflect their coverage. The fitting code selects the relevant observation years and averages the available values.

## 2. Construct synthetic West Germany

The code has two steps. First, choose the **predictor weights** using predictor values from the 1970s and outcome predictions over 1981–1990. Then use those predictor weights to fit the **donor weights** using the later predictor values.

The published specification includes 1990 in some fitting windows. Reunification took place during 1990. For the summaries below, use 1960–1989 as the pre-treatment period and 1991–2003 as the post-treatment period, leaving the transition year out of these summaries. This follows the example's fitting choices while making the evaluation windows explicit.

### Step 1: Predictor weights

`Dataprep` records the variables, countries, and periods to use. `predictors_op="mean"` averages each predictor over the stated years. `special_predictors` gives selected variables their own observation windows. Python's `range(1971, 1981)` includes 1971 through 1980.

```{python}
import pandas as pd
import pysyncon as sc
import matplotlib.pyplot as plt
import copy
df = pd.read_csv("germany.csv")

dataprep_train = sc.Dataprep(
    foo=df,
    predictors=["gdp", "trade", "infrate"],
    predictors_op="mean",
    time_predictors_prior=range(1971, 1981),
    special_predictors=[
        ("industry", range(1971, 1981), "mean"),
        ("schooling", [1970, 1975], "mean"),
        ("invest70", [1980], "mean"),
    ],
    dependent="gdp",
    unit_variable="country",
    time_variable="year",
    treatment_identifier="West Germany",
    controls_identifier=[
        "USA",
        "UK",
        "Austria",
        "Belgium",
        "Denmark",
        "France",
        "Italy",
        "Netherlands",
        "Norway",
        "Switzerland",
        "Japan",
        "Greece",
        "Portugal",
        "Spain",
        "Australia",
        "New Zealand",
    ],
    time_optimize_ssr=range(1981, 1991),
)


synth_train = sc.Synth()
synth_train.fit(dataprep=dataprep_train)
```

### Step 2: Donor weights

`custom_V=synth_train.V` carries the predictor weights from the first fit into the second.

```{python}
dataprep = sc.Dataprep(
    foo=df,
    predictors=["gdp", "trade", "infrate"],
    predictors_op="mean",
    time_predictors_prior=range(1981, 1991),
    special_predictors=[
        ("industry", range(1981, 1991), "mean"),
        ("schooling", [1980, 1985], "mean"),
        ("invest80", [1980], "mean"),
    ],
    dependent="gdp",
    unit_variable="country",
    time_variable="year",
    treatment_identifier="West Germany",
    controls_identifier=[
        "USA",
        "UK",
        "Austria",
        "Belgium",
        "Denmark",
        "France",
        "Italy",
        "Netherlands",
        "Norway",
        "Switzerland",
        "Japan",
        "Greece",
        "Portugal",
        "Spain",
        "Australia",
        "New Zealand",
    ],
    time_optimize_ssr=range(1960, 1990),
)


synth = sc.Synth()
synth.fit(dataprep=dataprep, custom_V=synth_train.V)
```

**(a)** Report the positive donor weights in a labelled table and compare them with Table 1 of the paper. Check that the weights sum to one.

```{python}
# Retain precision for calculations. Round only the displayed table.
weights = synth.weights(round=12)
print(weights.round(3))
print(weights.sum())
```

**(b)** Compare the treated and synthetic predictor values. Where is the fit closest? Where are the largest discrepancies?

```{python}
balance = synth.summary()
print(balance)
```

The package's `sample mean` column is an **unweighted** donor average. The corrected paper table uses population weights for its OECD comparison column. Focus on the treated and synthetic columns when comparing the fit.

## 3. Outcome paths and estimated effects

The code below applies each donor weight to that country's GDP in every year.

```{python}
# One row per year, with one GDP column per country.
gdp = df.pivot(index="year", columns="country", values="gdp")

paths = pd.DataFrame(index=gdp.index)
paths["West Germany"] = gdp["West Germany"]
paths["Synthetic West Germany"] = (
    gdp[weights.index].mul(weights, axis="columns").sum(axis=1)
)
paths["gap"] = paths["West Germany"] - paths["Synthetic West Germany"]
```

**(a)** Plot the observed and synthetic GDP paths, with a vertical line at 1990. Use different line styles and label the units correctly. Compare your figure with Figure 2 of the paper.

**(b)** Plot the gap and compare it with Figure 3. Calculate the average gap over 1991–2003. What is the estimated effect in 2003?

**(c)** Express the 2003 gap as a percentage of synthetic GDP in 2003. Explain the denominator. Discuss how you interpret a dollar gap measured in current prices.

**(d)** Give a substantive assumption needed to interpret the gap as the effect of reunification. What other development could threaten that interpretation?

## 4. Remove an influential donor

Remove the country with the largest weight and refit, retaining the predictor weights from the original fit. This examines sensitivity to donor availability at the original predictor weights.

```{python}
largest_donor = weights.idxmax()
print(largest_donor)

# Copy the setup so the original model remains available.
dataprep_leaveout = copy.deepcopy(dataprep)
dataprep_leaveout.controls_identifier = [
    country for country in dataprep.controls_identifier
    if country != largest_donor
]

synth_leaveout = sc.Synth()
synth_leaveout.fit(dataprep=dataprep_leaveout, custom_V=synth_train.V)
leaveout_weights = synth_leaveout.weights(round=12)

paths["Synthetic without largest donor"] = (
    gdp[leaveout_weights.index].mul(leaveout_weights, axis="columns").sum(axis=1)
)
```

**(a)** Plot the two synthetic paths alongside West Germany. Compare both pre-treatment fit and post-treatment differences.

**(b)** Recalculate the average gap over 1991–2003. Does removing this donor change your substantive conclusion?

## 5. Placebo comparisons

The supplied loop repeats the two fitting steps for each donor country. In each repetition, that country becomes the placebo treated unit. West Germany stays out of the donor pool. Each placebo has its predictor weights selected again, using the same time windows as the original fit. The loop may take a few minutes.

```{python}
# Store West Germany's gap first, then add each placebo gap.
placebo_gaps = pd.DataFrame(index=gdp.index)
placebo_gaps["West Germany"] = paths["gap"]

for placebo_country in dataprep.controls_identifier:
    # This country is the placebo treated unit, so remove it from its donors.
    placebo_donors = [
        country for country in dataprep.controls_identifier
        if country != placebo_country
    ]

    # Re-estimate predictor weights for this placebo country.
    placebo_train = copy.deepcopy(dataprep_train)
    placebo_train.treatment_identifier = placebo_country
    placebo_train.controls_identifier = placebo_donors
    placebo_train_fit = sc.Synth()
    placebo_train_fit.fit(dataprep=placebo_train)

    # Estimate its donor weights using the selected predictor weights.
    placebo_setup = copy.deepcopy(dataprep)
    placebo_setup.treatment_identifier = placebo_country
    placebo_setup.controls_identifier = placebo_donors
    placebo_fit = sc.Synth()
    placebo_fit.fit(dataprep=placebo_setup, custom_V=placebo_train_fit.V)
    placebo_weights = placebo_fit.weights(round=12)

    # Apply those weights to GDP in every year, then calculate the gap.
    placebo_path = (
        gdp[placebo_weights.index].mul(placebo_weights, axis="columns").sum(axis=1)
    )
    placebo_gaps[placebo_country] = gdp[placebo_country] - placebo_path

# Mean squared gaps before and after reunification, for each country.
pre_mspe = placebo_gaps.loc[1960:1989].pow(2).mean()
post_mspe = placebo_gaps.loc[1991:2003].pow(2).mean()
mspe_ratio = post_mspe / pre_mspe

placebo_results = pd.DataFrame({
    "Pre-treatment MSPE": pre_mspe,
    "Post-treatment MSPE": post_mspe,
    "Post/pre ratio": mspe_ratio
})
print(placebo_results.sort_values("Post/pre ratio", ascending=False).round(3))
```

**(a)** Plot the placebo gaps in grey, with West Germany's gap as a thick black line. Add a vertical line at 1990 and a horizontal line at zero. Which placebo countries fit poorly before reunification?

**(b)** Calculate the fraction of the 17 countries whose post/pre-MSPE ratio is at least as large as West Germany's. Include West Germany in both counts. What is the smallest possible value of this fraction?

**(c)** How unusual is West Germany's result? Explain what the placebo comparison contributes to your assessment and what is required for an exact randomization interpretation.

## 6. Your conclusion

What do the estimated magnitudes, pre-treatment fit, leave-one-out result, and placebo evidence suggest about the effect of reunification on West German GDP per capita? What are the limitations of the counterfactual comparison?

## References

- Abadie, Diamond, and Hainmueller (2015), [Comparative Politics and the Synthetic Control Method](https://doi.org/10.1111/ajps.12116).
- [2025 erratum](https://economics.mit.edu/sites/default/files/2025-09/Comparative%20Politics%20and%20the%20Synthetic%20Control%20Method%20erratum.pdf).
- [pysyncon documentation](https://sdfordham.github.io/pysyncon/) and [German reunification example](https://github.com/sdfordham/pysyncon/blob/main/examples/germany.ipynb).
