---
title: "Week 1: Vaccination Experiment"
subtitle: "Ungraded after-class exercise"
---

[Week 1 notes](../sessions/01-experiments.qmd) · [Download this notebook](../resources/exercises/01-experiments.qmd) · [Teaching data](../data/vaccination-teaching.dta) · [Data dictionary](../data/README.md)

Replicate **Table S3 and Table S5** from Campos-Mercade et al. (2021), and **comment on Table S4**. Use the [supplement](../resources/readings/vaccination-supplement.pdf), pp. 28–31. You may work together. Include short interpretations with your code; matching the typography of the published tables is unnecessary.

## 1. Load and understand the data

The teaching file contains the authors' 8,286 main-analysis observations and the variables needed here. `vaccinated` and `intention1` are coded **0/1**. The assignment indicators are also 0/1. Each participant belongs to exactly one of six arms. The [dictionary](../data/README.md) explains the retained variables and transformations.

Download the notebook and data. You can put the data in the same folder as the notebook, or retain the site's `data/` folder. The following code checks both locations.

```{python}
from pathlib import Path
import pandas as pd
import statsmodels.formula.api as smf

candidates = [
    Path("data/vaccination-teaching.dta"),
    Path("vaccination-teaching.dta"),
    Path("../data/vaccination-teaching.dta"),
]
data_path = next((p for p in candidates if p.exists()), None)
if data_path is None:
    raise FileNotFoundError("Download vaccination-teaching.dta beside this notebook.")
df = pd.read_stata(data_path, convert_categoricals=False)
print(f"{len(df):,} observations")
```

1. Tabulate the number of participants in each arm. Which arm is the reference group in the published regressions? What does that group receive?
2. Construct a variable called `nudge`, equal to one for the social-impact, argument, or information conditions and zero otherwise. Verify that it has the intended values. Do not include the no-reminders arm in this group.
3. Check the units and possible values of the two outcomes.

```{python}
# Your code: inspect the assignment indicators and construct nudge.
```

## 2. Replicate Table S3: balance

Begin with group means for the eight characteristics: `age`, `female`, `single`, `haschildren`, `university`, `income`, `foreign`, and `unemployed`.

Table S3 presents **regressions of baseline characteristics on assignment indicators**, without additional controls. The coefficients describe differences from the control group; they are not causal effects of assignment on pre-existing characteristics.

For **each** characteristic, estimate:

- A pooled specification with `treat_pay`, your `nudge`, and `treat_min`.
- A separate-arm specification with `treat_pay`, `treat_soc`, `treat_arg`, `treat_info`, and `treat_min`.

Include an intercept and omit the control indicator. For the income regressions, create a separate variable measured in **10,000 SEK**, as in the table. Retain the original `income` for the later categorical control.

::: {.callout-note}
## Standard errors
Use `.fit(cov_type="HC1", use_t=True)` for the heteroskedasticity-consistent standard errors and t-based inference used in this exercise. Record coefficients, standard errors, and the number of observations. A loop can save repetition across baseline characteristics.
:::

```{python}
# Your code: pooled and separate-arm balance regressions.
```

Explain in two or three sentences: are the differences substantively large? How would you interpret the presence of some significant coefficients? Does this table establish balance on unobserved characteristics?

## 3. Comment on Table S4: external validity

You do **not** need to replicate this table or obtain population data.

- Identify two differences between the experimental sample and the comparison population. Why might they matter for transporting the treatment effect?
- Read the footnotes. Are income and immigration background defined identically in the two sources?
- Would a sample that matched these population means necessarily tell us the effect of a different incentive in another country or at a later stage of vaccination rollout?

Distinguish a concern about **external validity** from a failure of **internal validity**.

## 4. Replicate Table S5: the linear probability models

Estimate the two treatment specifications from Table S3 for each of `vaccinated` and `intention1`, now adding the authors' controls. This produces the four columns of Table S5. Keep all six arms and use the control arm as the omitted category.

::: {.callout-note}
## Categorical controls in Python
In a statsmodels formula, `C(region)` treats numeric region codes as categories and includes indicators, omitting a reference category. `region` alone would treat the codes as a quantitative variable.

The authors also allow baseline differences across **age-group–region combinations**. `C(age_band) * C(region)` includes age-group indicators, region indicators, and their interactions. We supply this part of the formula; you do not need to construct those indicators yourself.
:::

```{python}
controls = (
    "C(female) + C(age_band) * C(region) + C(covid1_2)"
    " + C(civilstatus) + C(haschildren) + C(education)"
    " + C(occupation) + C(mother) + C(father) + C(income)"
)
```

Build a formula by joining your outcome, treatment indicators, and `controls`. For example, the general pattern is:

```{python}
#| eval: false
formula = "outcome ~ assignment_indicators + " + controls
model = smf.ols(formula, data=df).fit(cov_type="HC1", use_t=True)
```

Replace `outcome` and `assignment_indicators` with the appropriate variable names/formula terms. They are placeholders, not variables in the data. Report the treatment coefficients and standard errors; there is no need to print every control coefficient.

```{python}
# Your code: the four models in Table S5.
```

Answer these questions alongside your results:

1. Your outcome coefficients are one-hundredth of the published coefficients. Have you replicated the findings? Explain the units, including those of the standard errors.
2. Interpret the payment coefficient in percentage points. Distinguish this from a percentage change relative to control uptake.
3. What do the pooled nudge estimates suggest about intentions and actual behaviour? Does an insignificant coefficient establish a zero effect?
4. Why would controlling for `intention1` in the vaccination equation be inappropriate for estimating the total effect of assignment?

### A standard-error check

Refit one model using `.fit()` with no covariance option. Compare its coefficients, standard errors, and confidence intervals with the HC1 version. What changes and what stays the same? Relate the difference to the LPM result $\operatorname{Var}(u_i\mid X_i)=p_i(1-p_i)$ when its conditional mean is correctly specified. Robust standard errors need not always be larger.

```{python}
# Your code and interpretation: default versus HC1 standard errors.
```

## Sources

Campos-Mercade et al. (2021), [paper](https://doi.org/10.1126/science.abm0475), [supplement](../resources/readings/vaccination-supplement.pdf), and [data and code](https://doi.org/10.5281/zenodo.5529626). See the [week 1 notes](../sessions/01-experiments.qmd) for the analysis-plan discussion and conceptual background.
