---
title: "Week 7: Regression Discontinuity"
subtitle: "PP5001 · Martinmas 2026 · Ungraded practice"
format:
  pdf:
    papersize: a4
    fontsize: 11pt
    geometry: margin=25mm
    toc: false
  html:
    embed-resources: true
    code-fold: false
    toc: false
    html-math-method: katex
    include-in-header:
      text: |
        <style>
        @media print {
          @page { size: A4; margin: 15mm; }
          html, body { width: auto !important; min-width: 0 !important; }
          #quarto-content, main.content { display: block !important; width: 100% !important; max-width: 100% !important; min-width: 0 !important; margin: 0 !important; padding: 0 !important; }
          .cell, .cell-output, .cell-output-display, .table-responsive, .quarto-table-container { min-width: 0 !important; max-width: 100% !important; overflow: visible !important; }
          table { display: table !important; width: 100% !important; max-width: 100% !important; table-layout: fixed !important; border-collapse: collapse; font-size: 8.5pt !important; }
          th, td { min-width: 0 !important; width: auto !important; white-space: normal !important; overflow-wrap: anywhere !important; padding: 3pt !important; }
          thead { display: table-header-group; }
          tr { break-inside: avoid; }
          pre, pre code, pre code span { white-space: pre-wrap !important; overflow-wrap: anywhere !important; }
          pre { max-width: 100% !important; overflow: visible !important; font-size: 8pt; }
          img, svg { max-width: 100% !important; height: auto; }
          .code-copy-button, .anchorjs-link { display: none !important; }
        }
        </style>
execute:
  echo: true
  warning: true
  error: false
jupyter: python3
---

Replicate selected results from David S. Lee (2008), *Randomized Experiments from Non-random Selection in U.S. House Elections*. Then use `rdrobust` to examine how the estimates change and assess the evidence supporting the design. This post-class exercise is **ungraded**.

Use **Table 2, columns (1), (5), and (8)**. Work in the downloadable Quarto notebook and render your completed analysis to HTML or PDF. Create a comparison table with clear labels, units, standard errors or confidence intervals, and sample sizes.

## Getting started

Download the Quarto notebook from this site and `table_two_final.dta` from **Moodle**. Save both files in the same folder. Open the Quarto file and add your code and written answers beneath each question. You will need `pandas`, `numpy`, `matplotlib`, `statsmodels`, `rdrobust`, and `rddensity`. Install any missing packages in the Python environment used by Quarto.

```{python}
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import statsmodels.formula.api as smf
import rdrobust as rd
import rddensity as density

raw = pd.read_stata("table_two_final.dta", convert_categoricals=False)

# Match the estimation sample in Lee's Table 2.
# One observation marked use == 1 has missing treatment and interaction variables.
data = raw.loc[(raw["use"] == 1) & raw["right"].notna()].copy()
print("Observations:", len(data))
```

The resulting sample contains **6,558 observations**. Vote shares and the margin of victory are recorded as proportions. An estimated effect of 0.01 therefore represents one percentage point.

| Variable | Meaning |
|:--|:--|
| `difdemshare` | Democratic margin of victory in election $t$. The cutoff is zero. |
| `right` | Indicator for Democratic victory in election $t$. |
| `demsharenext` | Democratic vote share in election $t+1$. |
| `demshareprev` | Democratic vote share in election $t-1$. |
| `demwinprev` | Democratic victory measure for election $t-1$. |
| `demofficeexp`, `othofficeexp` | Accumulated previous election victories of the Democratic and opposition candidates. |
| `demelectexp`, `othelectexp` | Accumulated previous election attempts of those candidates. |
| `statedisdec` | State–district–decade identifier used to cluster standard errors. |
| `difdemshare2`, `difdemshare3`, `difdemshare4` | Squared, cubed, and fourth-power terms in the margin. |
| `rdifdemshare`, `rdifdemshare2`, `rdifdemshare3`, `rdifdemshare4` | Each margin term multiplied by `right`. |

## 1. The design and the graph

**(a)** Explain the treatment, running variable, cutoff, and outcome. State the identifying assumption and explain whose incumbency advantage is being estimated.

**(b)** Plot average Democratic vote share in election $t+1$ against the margin in election $t$, using bins that keep observations on opposite sides of zero separate. Mark the cutoff. Describe the pattern near zero. The binning example in the notes provides a starting point.

```{python}
# Add your code for Question 1 here.
```

## 2. Replicate Lee's Table 2

**(a)** Replicate column (1). Regress `demsharenext` on `right`, a fourth-order polynomial in the margin, and interactions of every polynomial term with `right`. Use the full estimation sample. Explain how the interaction terms allow the fitted relationship to differ on the two sides of the cutoff.

**(b)** Replicate column (5) by adding `demshareprev`, `demwinprev`, `demofficeexp`, `othofficeexp`, `demelectexp`, and `othelectexp`.

**(c)** Replicate column (8). Use `demshareprev` as the dependent variable, retain the fourth-order polynomial and its interactions, and include `demwinprev`, `demofficeexp`, `othofficeexp`, `demelectexp`, and `othelectexp` as controls.

Use standard errors clustered by `statedisdec` in all three regressions. In `statsmodels`, append the following to the model specification:

```python
.fit(cov_type="cluster", cov_kwds={"groups": data["statedisdec"]})
```

**(d)** Report the coefficient on `right`, its standard error, and the sample size for each regression. Compare your results with the published table. Interpret the coefficients in columns (1) and (5). Why does column (8) provide a check on the design?

```{python}
# Estimate each regression separately and assemble your results here.
```

## 3. Re-estimate using rdrobust

**(a)** Re-estimate the discontinuity in `demsharenext` using a local linear regression with triangular weights and an automatically selected bandwidth. Retain clustering by `statedisdec`. The following code shows the required specification:

```{python}
#| eval: false
rd_column1 = rd.rdrobust(
    y=data["demsharenext"],
    x=data["difdemshare"],
    c=0,
    p=1,
    q=2,
    kernel="tri",
    bwselect="mserd",
    cluster=data["statedisdec"]
)
print(rd_column1)
```

Here `p=1` requests a local linear fit, `q=2` uses a local quadratic fit to estimate its bias, and `bwselect="mserd"` selects the estimation bandwidth. To run this example when rendering your notebook, remove `#| eval: false` from the code cell.

**(b)** Repeat the analysis with the column (5) controls. Supply these to `rdrobust` using its `covs` argument, `covs=data[["demshareprev", "demwinprev", "demofficeexp", "othofficeexp", "demelectexp", "othelectexp"]]`. Then repeat the column (8) placebo analysis with its dependent variable and controls.

**(c)** For each specification, report the conventional estimate, bias-corrected estimate, robust 95% confidence interval, bandwidth, and number of observations inside the bandwidth. The result's `.coef`, `.ci`, `.bws`, and `.N_h` attributes contain these quantities. The robust interval is centered on the bias-corrected estimate.

**(d)** Compare these results with your replication. Discuss the estimated magnitudes, uncertainty, and substantive conclusions. Explain how using a local linear fit, triangular weights, and a selected bandwidth changes the estimation approach.

```{python}
# Add your rdrobust estimates and comparison table here.
```

## 4. Sorting around the cutoff

**(a)** Plot a histogram of the margin of victory near zero. Choose bins with zero as an edge and use equal bin widths on both sides. Describe any apparent difference in the density immediately below and above zero.

**(b)** Apply the density-discontinuity test implemented by `rddensity` to the running variable in the estimation sample. Use the full estimation sample as the input, allowing the procedure to select its bandwidth. The following code reports the test statistic and p-value:

```{python}
#| eval: false
sorting_test = density.rddensity(data["difdemshare"], c=0)
print("Test statistic:", sorting_test.test["t_jk"])
print("p-value:", sorting_test.test["p_jk"])
```

Remove `#| eval: false` to run the code when rendering. This is the Cattaneo–Jansson–Ma density test. Review “How the Cattaneo–Jansson–Ma Test Works” in the Week 7 notes: the procedure estimates density from the slopes of local fits to the empirical cumulative distribution, then tests whether the two densities at the cutoff are equal.

**(c)** State the null hypothesis. Why would a discontinuity in the density raise concerns about the design? Interpret your result and explain how much reassurance it provides.

**(d)** Consider the density test alongside column (8). What does each check tell you? Taken together with the effect estimates, how persuasive do you find the evidence for an incumbency advantage?

```{python}
# Add your histogram and density test here.
```

## Reference

Lee, David S. (2008). [Randomized Experiments from Non-random Selection in U.S. House Elections](https://www.princeton.edu/~davidlee/wp/RDrand.pdf). *Journal of Econometrics*, 142(2), 675–697. Table 2.
