import statsmodels.formula.api as smf
first_stage = smf.wls(
"medicaid ~ selected + C(stratum)",
data=analysis,
weights=analysis["weight_12m"],
).fit(
cov_type="cluster",
cov_kwds={"groups": analysis["household_id"]},
)Week 3 exercise: Insurance, offers, and take-up
The Oregon Health Insurance Experiment · Ungraded practice
Week 3 notes · Download the Quarto setup file
Use the public-use data to estimate the effect of Medicaid on the probability of an outpatient visit. Work in Quarto, showing the code, a labelled results table, and your interpretation. This exercise is ungraded.
Data and setup
Download PP5001-week-03-oregon-data.zip from the Week 3 section of Moodle and unzip it. Save the Quarto setup file beside the extracted data and documentation folders. The data folder contains:
oregonhie_descriptive_vars.dtaoregonhie_stateprograms_vars.dtaoregonhie_survey12m_vars.dta
Read the supplied Public Use Data User Guide (documentation/ohie_userguide.pdf), especially the descriptions of randomisation, merging files, and the twelve-month survey. Each file has one row per person; person_id is the identifier used to merge them. The setup file supplies the import and merge code.
Source: Oregon Health Insurance Experiment public-use data, distributed by NBER. The Moodle download contains unchanged copies of the three original data files, the user guide, and their codebooks.
If needed, install the IV package from your terminal:
python -m pip install linearmodelsVariables used in the exercise
| Name in our analysis | Original variable | Meaning |
|---|---|---|
selected |
treatment |
1 if selected by the lottery, 0 otherwise |
medicaid |
ohp_all_ever_firstn_30sep2009 |
Any Medicaid enrolment between 10 March 2008 and 30 September 2009 |
outpatient |
doc_any_12m |
Any outpatient visit in the previous six months |
numhh_list |
Same | Number of household members on the lottery list |
wave_survey12m |
Same | Survey wave |
weight_12m |
Same | Supplied survey weight |
household_id |
Same | Household identifier for clustered standard errors |
We keep twelve-month respondents with positive survey weights and complete information on these variables. The setup file makes this selection explicitly, before estimating any model. Use this same sample throughout so differences between estimates reflect the method rather than changes in the observations used.
The setup file also creates stratum, identifying each observed combination of household-list size and survey wave. C(stratum) adds an indicator for each group, with one omitted as the reference group. This implements the household-size indicators, survey-wave indicators, and their interactions described in the paper.
Survey weights account for the survey’s sampling and follow-up design. Use the supplied weight_12m values in every regression below.
1. Understand the comparison
Explain what was randomised and what remained a matter of take-up and eligibility. Identify the instrument, treatment, and outcome. Why might directly comparing Medicaid recipients with nonrecipients be misleading?
Report the number of observations in your analysis sample. Tabulate lottery selection and Medicaid coverage. Explain why the two indicators differ.
2. First stage
Estimate the regression of Medicaid coverage on lottery selection and the group indicators. The following code shows how to use the survey weights and cluster standard errors by household:
Report the coefficient on selected and its standard error. Interpret the coefficient in percentage points. Compare it with the survey first stage in the first row of Table III.
Clustering permits different disturbance variances and correlation among people in the same household. The household is the appropriate unit here because members share lottery assignment.
3. Intention-to-treat effect
Use the same specification, replacing the dependent variable with outpatient. Report and interpret the coefficient on selected. Which policy question does this ITT answer?
This is a linear probability model. Keep the outcome coded zero or one; multiply the coefficient by 100 when interpreting it in percentage points.
4. Instrumental variables
Divide your ITT coefficient by your first-stage coefficient. Then estimate the IV model:
import linearmodels.iv as iv
iv_model = iv.IV2SLS.from_formula(
"outpatient ~ 1 + C(stratum) + [medicaid ~ selected]",
data=analysis,
weights=analysis["weight_12m"],
)
iv_result = iv_model.fit(
cov_type="clustered",
clusters=analysis["household_id"],
debiased=True,
)Explain the meaning of [medicaid ~ selected]. Verify that the IV coefficient equals the ratio you calculated. Report its household-clustered standard error and 95% confidence interval from iv_result.
Explain why dividing the ITT standard error by the first-stage coefficient would fail to account for all the uncertainty in the ratio.
5. Compare and interpret
Estimate the direct regression of outpatient on medicaid and C(stratum), using the same survey weights and household-clustered standard errors.
Create a table showing the direct regression estimate, the first stage, the ITT, and the IV estimate. Label the coefficient in each row, its standard error, and the sample size. You can use Claude to help format the table; check that the entries come from your fitted models.
Compare your IV result with the “Outpatient visits last six months” / “Extensive margin (any)” row of Table V. Your regression uses the public files and the design controls described above. Small differences can arise from outcome-specific missing observations and rounding; explain the sample you actually use.
Finally, explain:
- Who are the compliers for this instrument?
- What do independence, exclusion, and monotonicity require in this application?
- What can you conclude about the effect and its uncertainty?
- What further information would you want before applying the result to a different insurance expansion?
6. Critique Claude
Give Claude the following prompt:
Explain what the Oregon Health Insurance Experiment tells us about the effects of Medicaid and what policymakers should learn from it. Write approximately 400 words for a Master of Public Policy student.
Include Claude’s response in your notebook. Then answer:
Assess Claude’s explanation using the assigned paper and your estimates. Identify two claims that are well supported and two statements that require correction or qualification. Explain your reasoning and rewrite the latter statements.
Presenting your work
Create your own tables and do not rely only on regression output. Give effects in probability units or percentage points, and label the choice clearly. The setup file includes print styling to keep tables within the page width. To make a PDF without LaTeX, open the rendered HTML in a browser and use Print → Save as PDF.
If Python reports a warning or cannot estimate a model, check the data and regression code. You can ask Claude to explain the message and help diagnose the problem. Make sure you understand suggested changes before using them.