Replicate Table S3 and Table S5 from Campos-Mercade et al. (2021), and comment on Table S4. Use the supplement, pp. 28–31. You may work together. Include short interpretations with your code; matching the typography of the published tables is unnecessary.
1. Load and understand the data
The teaching file contains the authors’ 8,286 main-analysis observations and the variables needed here. vaccinated and intention1 are coded 0/1. The assignment indicators are also 0/1. Each participant belongs to exactly one of six arms. The dictionary explains the retained variables and transformations.
Download the notebook and data. You can put the data in the same folder as the notebook, or retain the site’s data/ folder. The following code checks both locations.
from pathlib import Pathimport pandas as pdimport statsmodels.formula.api as smfcandidates = [ Path("data/vaccination-teaching.dta"), Path("vaccination-teaching.dta"), Path("../data/vaccination-teaching.dta"),]data_path =next((p for p in candidates if p.exists()), None)if data_path isNone:raiseFileNotFoundError("Download vaccination-teaching.dta beside this notebook.")df = pd.read_stata(data_path, convert_categoricals=False)print(f"{len(df):,} observations")
8,286 observations
Tabulate the number of participants in each arm. Which arm is the reference group in the published regressions? What does that group receive?
Construct a variable called nudge, equal to one for the social-impact, argument, or information conditions and zero otherwise. Verify that it has the intended values. Do not include the no-reminders arm in this group.
Check the units and possible values of the two outcomes.
# Your code: inspect the assignment indicators and construct nudge.
2. Replicate Table S3: balance
Begin with group means for the eight characteristics: age, female, single, haschildren, university, income, foreign, and unemployed.
Table S3 presents regressions of baseline characteristics on assignment indicators, without additional controls. The coefficients describe differences from the control group; they are not causal effects of assignment on pre-existing characteristics.
For each characteristic, estimate:
A pooled specification with treat_pay, your nudge, and treat_min.
A separate-arm specification with treat_pay, treat_soc, treat_arg, treat_info, and treat_min.
Include an intercept and omit the control indicator. For the income regressions, create a separate variable measured in 10,000 SEK, as in the table. Retain the original income for the later categorical control.
NoteStandard errors
Use .fit(cov_type="HC1", use_t=True) for the heteroskedasticity-consistent standard errors and t-based inference used in this exercise. Record coefficients, standard errors, and the number of observations. A loop can save repetition across baseline characteristics.
# Your code: pooled and separate-arm balance regressions.
Explain in two or three sentences: are the differences substantively large? How would you interpret the presence of some significant coefficients? Does this table establish balance on unobserved characteristics?
3. Comment on Table S4: external validity
You do not need to replicate this table or obtain population data.
Identify two differences between the experimental sample and the comparison population. Why might they matter for transporting the treatment effect?
Read the footnotes. Are income and immigration background defined identically in the two sources?
Would a sample that matched these population means necessarily tell us the effect of a different incentive in another country or at a later stage of vaccination rollout?
Distinguish a concern about external validity from a failure of internal validity.
4. Replicate Table S5: the linear probability models
Estimate the two treatment specifications from Table S3 for each of vaccinated and intention1, now adding the authors’ controls. This produces the four columns of Table S5. Keep all six arms and use the control arm as the omitted category.
NoteCategorical controls in Python
In a statsmodels formula, C(region) treats numeric region codes as categories and includes indicators, omitting a reference category. region alone would treat the codes as a quantitative variable.
The authors also allow baseline differences across age-group–region combinations. C(age_band) * C(region) includes age-group indicators, region indicators, and their interactions. We supply this part of the formula; you do not need to construct those indicators yourself.
Replace outcome and assignment_indicators with the appropriate variable names/formula terms. They are placeholders, not variables in the data. Report the treatment coefficients and standard errors; there is no need to print every control coefficient.
# Your code: the four models in Table S5.
Answer these questions alongside your results:
Your outcome coefficients are one-hundredth of the published coefficients. Have you replicated the findings? Explain the units, including those of the standard errors.
Interpret the payment coefficient in percentage points. Distinguish this from a percentage change relative to control uptake.
What do the pooled nudge estimates suggest about intentions and actual behaviour? Does an insignificant coefficient establish a zero effect?
Why would controlling for intention1 in the vaccination equation be inappropriate for estimating the total effect of assignment?
A standard-error check
Refit one model using .fit() with no covariance option. Compare its coefficients, standard errors, and confidence intervals with the HC1 version. What changes and what stays the same? Relate the difference to the LPM result \(\operatorname{Var}(u_i\mid X_i)=p_i(1-p_i)\) when its conditional mean is correctly specified. Robust standard errors need not always be larger.
# Your code and interpretation: default versus HC1 standard errors.
---title: "Week 1: Vaccination Experiment"subtitle: "Ungraded after-class exercise"---[Week 1 notes](../sessions/01-experiments.qmd) · [Download this notebook](../resources/exercises/01-experiments.qmd) · [Teaching data](../data/vaccination-teaching.dta) · [Data dictionary](../data/README.md)Replicate **Table S3 and Table S5** from Campos-Mercade et al. (2021), and **comment on Table S4**. Use the [supplement](../resources/readings/vaccination-supplement.pdf), pp. 28–31. You may work together. Include short interpretations with your code; matching the typography of the published tables is unnecessary.## 1. Load and understand the dataThe teaching file contains the authors' 8,286 main-analysis observations and the variables needed here. `vaccinated` and `intention1` are coded **0/1**. The assignment indicators are also 0/1. Each participant belongs to exactly one of six arms. The [dictionary](../data/README.md) explains the retained variables and transformations.Download the notebook and data. You can put the data in the same folder as the notebook, or retain the site's `data/` folder. The following code checks both locations.```{python}from pathlib import Pathimport pandas as pdimport statsmodels.formula.api as smfcandidates = [ Path("data/vaccination-teaching.dta"), Path("vaccination-teaching.dta"), Path("../data/vaccination-teaching.dta"),]data_path =next((p for p in candidates if p.exists()), None)if data_path isNone:raiseFileNotFoundError("Download vaccination-teaching.dta beside this notebook.")df = pd.read_stata(data_path, convert_categoricals=False)print(f"{len(df):,} observations")```1. Tabulate the number of participants in each arm. Which arm is the reference group in the published regressions? What does that group receive?2. Construct a variable called `nudge`, equal to one for the social-impact, argument, or information conditions and zero otherwise. Verify that it has the intended values. Do not include the no-reminders arm in this group.3. Check the units and possible values of the two outcomes.```{python}# Your code: inspect the assignment indicators and construct nudge.```## 2. Replicate Table S3: balanceBegin with group means for the eight characteristics: `age`, `female`, `single`, `haschildren`, `university`, `income`, `foreign`, and `unemployed`.Table S3 presents **regressions of baseline characteristics on assignment indicators**, without additional controls. The coefficients describe differences from the control group; they are not causal effects of assignment on pre-existing characteristics.For **each** characteristic, estimate:- A pooled specification with `treat_pay`, your `nudge`, and `treat_min`.- A separate-arm specification with `treat_pay`, `treat_soc`, `treat_arg`, `treat_info`, and `treat_min`.Include an intercept and omit the control indicator. For the income regressions, create a separate variable measured in **10,000 SEK**, as in the table. Retain the original `income` for the later categorical control.::: {.callout-note}## Standard errorsUse `.fit(cov_type="HC1", use_t=True)` for the heteroskedasticity-consistent standard errors and t-based inference used in this exercise. Record coefficients, standard errors, and the number of observations. A loop can save repetition across baseline characteristics.:::```{python}# Your code: pooled and separate-arm balance regressions.```Explain in two or three sentences: are the differences substantively large? How would you interpret the presence of some significant coefficients? Does this table establish balance on unobserved characteristics?## 3. Comment on Table S4: external validityYou do **not** need to replicate this table or obtain population data.- Identify two differences between the experimental sample and the comparison population. Why might they matter for transporting the treatment effect?- Read the footnotes. Are income and immigration background defined identically in the two sources?- Would a sample that matched these population means necessarily tell us the effect of a different incentive in another country or at a later stage of vaccination rollout?Distinguish a concern about **external validity** from a failure of **internal validity**.## 4. Replicate Table S5: the linear probability modelsEstimate the two treatment specifications from Table S3 for each of `vaccinated` and `intention1`, now adding the authors' controls. This produces the four columns of Table S5. Keep all six arms and use the control arm as the omitted category.::: {.callout-note}## Categorical controls in PythonIn a statsmodels formula, `C(region)` treats numeric region codes as categories and includes indicators, omitting a reference category. `region` alone would treat the codes as a quantitative variable.The authors also allow baseline differences across **age-group–region combinations**. `C(age_band) * C(region)` includes age-group indicators, region indicators, and their interactions. We supply this part of the formula; you do not need to construct those indicators yourself.:::```{python}controls = ("C(female) + C(age_band) * C(region) + C(covid1_2)"" + C(civilstatus) + C(haschildren) + C(education)"" + C(occupation) + C(mother) + C(father) + C(income)")```Build a formula by joining your outcome, treatment indicators, and `controls`. For example, the general pattern is:```{python}#| eval: falseformula ="outcome ~ assignment_indicators + "+ controlsmodel = smf.ols(formula, data=df).fit(cov_type="HC1", use_t=True)```Replace `outcome` and `assignment_indicators` with the appropriate variable names/formula terms. They are placeholders, not variables in the data. Report the treatment coefficients and standard errors; there is no need to print every control coefficient.```{python}# Your code: the four models in Table S5.```Answer these questions alongside your results:1. Your outcome coefficients are one-hundredth of the published coefficients. Have you replicated the findings? Explain the units, including those of the standard errors.2. Interpret the payment coefficient in percentage points. Distinguish this from a percentage change relative to control uptake.3. What do the pooled nudge estimates suggest about intentions and actual behaviour? Does an insignificant coefficient establish a zero effect?4. Why would controlling for `intention1` in the vaccination equation be inappropriate for estimating the total effect of assignment?### A standard-error checkRefit one model using `.fit()` with no covariance option. Compare its coefficients, standard errors, and confidence intervals with the HC1 version. What changes and what stays the same? Relate the difference to the LPM result $\operatorname{Var}(u_i\mid X_i)=p_i(1-p_i)$ when its conditional mean is correctly specified. Robust standard errors need not always be larger.```{python}# Your code and interpretation: default versus HC1 standard errors.```## SourcesCampos-Mercade et al. (2021), [paper](https://doi.org/10.1126/science.abm0475), [supplement](../resources/readings/vaccination-supplement.pdf), and [data and code](https://doi.org/10.5281/zenodo.5529626). See the [week 1 notes](../sessions/01-experiments.qmd) for the analysis-plan discussion and conceptual background.
3. Comment on Table S4: external validity
You do not need to replicate this table or obtain population data.
Distinguish a concern about external validity from a failure of internal validity.