---
title: "Week 5: Weak Instruments"
subtitle: "Compulsory schooling · PP5001 · Martinmas 2026"
---

[Class slides](../slides/05-weak-instruments.qmd) · [Slides PDF](../resources/slides/05-weak-instruments.pdf) · [Notes PDF](../resources/notes/05-weak-instruments.pdf) · [Quarto source](../resources/sessions/05-weak-instruments.qmd) · [Problem Set 2](../exercises/05-weak-instruments.qmd)

## Learning Objectives

By the end of this topic, you should be able to:

- Distinguish unbiasedness from consistency using the sampling distributions of estimators.
- Explain how a weak first stage can magnify violations of instrument validity and produce substantial finite-sample bias towards OLS.
- Interpret the first-stage F statistic and partial $R^2$, and explain why the second-stage results alone cannot establish instrument strength.
- Assess the quarter-of-birth identification strategy and compare Wald and IV estimates under different sets of controls and instruments.

## Before class

Read these notes and **Angrist and Krueger (1991)** and **Bound, Jaeger, and Baker (1995)**. Both papers are on Moodle. We will examine the compulsory-schooling argument, the first-stage evidence, and the interpretation of the IV estimates. We also use **Figure 1 in Bound and Jaeger (2000)** to discuss the exclusion restriction.

### Textbook references

- **The Effect:** [Chapter 19, Instrumental Variables](https://theeffectbook.net/ch-InstrumentalVariables.html), particularly instrument relevance and weak instruments.
- **Causal Inference: The Remix:** [Chapter 7, Instrumental Variables](https://mixtape.scunning.com/07-instrumental_variables), particularly the discussion of weak instruments.
- **Mastering 'Metrics:** Chapter 3, *Instrumental Variables*. [Companion resources](https://www.masteringmetrics.com/online-metrics-resources/mastering-iv/). [Mastering Econometrics videos at MRU](https://learn.mru.org/courses/mastering-econometrics).
- **Mostly Harmless Econometrics — advanced reading:** Section 4.6 on IV with weak instruments and finite-sample behaviour.

## 1. Unbiasedness and consistency

### Repeated samples

An estimate changes when the sample changes. Imagine repeatedly drawing samples of the same size from the same population and estimating the same relationship. The distribution of those estimates is the **sampling distribution**. Its centre and spread tell us different things about an estimator.

An estimator is **unbiased** if its expectation equals the true parameter:

$$E\left[\widehat\beta\right]=\beta.$$

This is a statement about the average over repeated samples of a particular size. It does not mean that any individual estimate equals the truth. An unbiased estimator can have a wide sampling distribution.

An estimator is **consistent** if, as the sample grows, its sampling distribution concentrates around the true parameter. For any fixed margin of error, the probability of falling outside that margin eventually becomes arbitrarily small. We can understand this graphically without deriving an asymptotic distribution.

```{=html}
<style>
.notes-animation video {display:block; width:100%; height:auto; margin:1rem auto;}
@media print { .notes-animation video {display:none;} }
</style>
<div class="notes-animation">
<video controls playsinline preload="metadata" poster="../resources/images/week-05/unbiasedness-consistency-poster.png" aria-label="Narrated animation explaining unbiasedness and consistency">
<source src="../resources/videos/week-05/unbiasedness-consistency-narrated.mp4" type="video/mp4">
</video>
<p><a href="../resources/videos/week-05/unbiasedness-consistency-narrated.mp4">Watch the narrated animation on unbiasedness and consistency</a> (5½ minutes).</p>
</div>
```

These distinctions matter directly for instrumental variables. OLS with an omitted variable behaves like the sample mean plus 20: more data gives a more precise estimate of the wrong quantity. IV with a valid, strong instrument behaves more like the sample mean plus $20/n$: biased in finite samples, but consistent. With a weak instrument, that finite-sample bias can remain large even in very big samples, and it pulls IV towards OLS. That is the problem we turn to next.

![Schematic sampling distributions. The dashed line marks the true effect. All panels use the same scales.](../resources/images/week-05/unbiasedness-consistency.png)

**The probability limit** is the value that the estimate approaches as the sample size grows: the probability that the estimate differs from that value by more than any fixed positive amount goes to zero.

We write $\operatorname{plim}\widehat\beta=b$. An estimator is **consistent** if its probability limit equals the true parameter, $b=\beta$.

The curves above are schematic normal densities chosen to illustrate the properties. The left panel is centred on the truth at every sample size and becomes more concentrated as the sample grows. In the middle panel the estimator is biased at each of the displayed sample sizes, but the centre moves towards the truth while the spread shrinks. In the right panel the distribution becomes increasingly concentrated around the wrong value. More observations make that estimator more precise about the wrong target.

::: {.callout-note title="Your turn"}
Which panels illustrate unbiasedness? Which illustrate consistency? Why is a smaller standard error insufficient to establish that an estimator is estimating the causal effect?
:::

### OLS and IV

When an omitted variable is correlated with the explanatory variable, OLS can converge to a slope different from the causal effect. As the sample size increases, the OLS estimate becomes more precise, but omitted-variable bias remains. We can therefore obtain a very precise estimate of the wrong population parameter. Under the IV assumptions, a relevant instrument allows IV to converge to the causal effect. Finite-sample IV estimates can nevertheless be biased and highly variable. With weak instruments, this distinction matters even in datasets containing hundreds of thousands of observations.

The concentration around a point describes the distribution of the **unscaled estimator**. A centred and standardised estimator is a different object. Our graphs here concern the coefficient estimates themselves.

## 2. What goes wrong when relevance is weak?

We use the notation in Bound, Jaeger, and Baker. Let $x_i$ denote schooling, $y_i$ log earnings, $z_i$ an excluded instrument, and $\epsilon_i$ the outcome disturbance. Initially omit additional controls:

$$x_i=\pi_0+\pi_1z_i+\nu_i,$$
$$y_i=\beta_0+\beta x_i+\epsilon_i.$$

Schooling $x_i$ can be endogenous because it is related to unobserved determinants of earnings in $\epsilon_i$. An instrument must be related to schooling and orthogonal to the outcome disturbance. This week asks what happens when the relationship between the instrument and schooling is very small.

For a binary instrument the Wald estimator is

$$\widehat\beta^{IV}=\frac{\overline y_{z=1}-\overline y_{z=0}}{\overline x_{z=1}-\overline x_{z=0}}.$$

A small denominator makes the estimated effect sensitive to small changes in the estimated group means. Sampling variation can dominate the schooling difference. If the population schooling difference is zero, the instrument fails relevance entirely.

### Deriving the inconsistency of OLS and IV

Start with $y_i=\beta_0+\beta x_i+\epsilon_i$. Substituting into the OLS formula gives

$$\widehat\beta_{OLS}=\beta+
\frac{\frac{1}{n}\sum_i\left(x_i-\bar x\right)\epsilon_i}
{\frac{1}{n}\sum_i\left(x_i-\bar x\right)^2}.$$

The probability limit, written $\operatorname{plim}$, is the value around which an estimator concentrates as the sample grows. Under the usual sampling conditions, the sample covariance and variance converge to their population counterparts:

$$\operatorname{plim}\widehat\beta_{OLS}=\beta+
\frac{\sigma_{x,\epsilon}}{\sigma_x^2}.$$

Here $\sigma_{x,\epsilon}=\operatorname{Cov}\left(x_i,\epsilon_i\right)$ and $\sigma_x^2=\operatorname{Var}\left(x_i\right)$.

Let $\widehat x_i$ be the first-stage fitted value. With an intercept, $\overline{\widehat x}=\bar x$ and the first-stage residual is orthogonal to $\widehat x_i$.

$$\widehat\beta_{IV}
=\frac{\sum_i\left(\widehat x_i-\bar x\right)y_i}
{\sum_i\left(\widehat x_i-\bar x\right)x_i}
=\beta+\frac{\frac{1}{n}\sum_i\left(\widehat x_i-\bar x\right)\epsilon_i}
{\frac{1}{n}\sum_i\left(\widehat x_i-\bar x\right)^2}.$$

Therefore,

$$\operatorname{plim}\widehat\beta_{IV}=\beta+
\frac{\sigma_{\hat x,\epsilon}}{\sigma_{\hat x}^2}.$$

The population moments refer to the population first-stage fitted component.

### What Does $R^2$ Measure?

**$R^2$ is the proportion of the sample variation in the dependent variable explained by the regression.**

In the first stage, the dependent variable is $x_i$. For OLS with an intercept,

$$R^2=\frac{\text{explained variation}}{\text{total variation}}
=\frac{\sum_i\left(\widehat x_i-\bar x\right)^2}{\sum_i\left(x_i-\bar x\right)^2}
=1-\frac{\sum_i\widehat\nu_i^2}{\sum_i\left(x_i-\bar x\right)^2}.$$

Here $\widehat\nu_i=x_i-\widehat x_i$. An $R^2$ of **0.01** means the instruments explain **1%** of the variation in $x_i$.

The population counterpart is $R^2_{x,z}=\sigma_{\hat x}^2/\sigma_x^2$, the share of population variance explained by the instruments.

The total variation is measured by the sum of squared deviations of $x_i$ from its sample mean. The explained variation is the corresponding sum for the fitted values. The residual sum of squares measures the variation left unexplained. With an intercept, OLS splits total variation into these two parts:

$$\sum_i\left(x_i-\bar x\right)^2
=\sum_i\left(\widehat x_i-\bar x\right)^2+\sum_i\widehat\nu_i^2.$$

Thus $R^2=0$ means the regression explains none of the sample variation, while $R^2=1$ means it fits every observation exactly. In the population expression, $\sigma_{\hat x}^2$ refers to the variance of the population first-stage fitted component. This is the variance ratio that enters the next derivation. Later, when we include controls, partial $R^2$ will measure how much of the variation remaining after those controls the excluded instruments explain.

### Relative Inconsistency of IV and OLS

Divide the two discrepancies from the true effect:

$$\frac{\operatorname{plim}\widehat\beta_{IV}-\beta}
{\operatorname{plim}\widehat\beta_{OLS}-\beta}
=\frac{\sigma_{\hat x,\epsilon}/\sigma_{\hat x}^2}
{\sigma_{x,\epsilon}/\sigma_x^2}
=\frac{\sigma_{\hat x,\epsilon}/\sigma_{x,\epsilon}}
{\sigma_{\hat x}^2/\sigma_x^2}.$$

The denominator is the population first-stage $R^2$:

$$\boxed{\frac{\operatorname{plim}\widehat\beta_{IV}-\beta}
{\operatorname{plim}\widehat\beta_{OLS}-\beta}
=\frac{\sigma_{\hat x,\epsilon}/\sigma_{x,\epsilon}}{R^2_{x,z}}.}$$

A small $R^2$ magnifies the consequences of imperfect instrument validity.

### One Instrument: Expressing the Result in Correlations

With one instrument,

$$\operatorname{plim}\widehat\beta_{IV}-\beta
=\frac{\sigma_{z,\epsilon}}{\sigma_{z,x}}
=\frac{\rho_{z,\epsilon}}{\rho_{z,x}}\frac{\sigma_\epsilon}{\sigma_x},$$

while

$$\operatorname{plim}\widehat\beta_{OLS}-\beta
=\rho_{x,\epsilon}\frac{\sigma_\epsilon}{\sigma_x}.$$

Dividing cancels $\sigma_\epsilon/\sigma_x$:

$$\boxed{\frac{\operatorname{plim}\widehat\beta_{IV}-\beta}
{\operatorname{plim}\widehat\beta_{OLS}-\beta}
=\frac{\rho_{z,\epsilon}/\rho_{x,\epsilon}}{\rho_{x,z}}.}$$

### When Can IV Be Worse Than OLS?

IV has a larger absolute inconsistency when

$$\left|\frac{\rho_{z,\epsilon}}{\rho_{x,\epsilon}}\right|
>\left|\rho_{x,z}\right|.$$

An instrument can be much less correlated with the disturbance than $x_i$ is, yet still produce a larger inconsistency if it barely predicts $x_i$.

We can measure the instrument's predictive power. We must assess its validity using the institutional setting and supporting evidence.

These expressions assume a fixed, nonzero population first stage and finite population moments. They describe what happens with imperfect validity even as the sample becomes arbitrarily large. The relative-inconsistency ratios also assume that OLS inconsistency is nonzero. If the instrument is valid, $\sigma_{\hat x,\epsilon}=0$ and IV is consistent.

With several instruments, the population first-stage fitted component is the linear combination of instruments that predicts $x_i$. Its variance divided by the variance of $x_i$ gives $R^2_{x,z}$. In the sample, the identity in the IV derivation follows by writing $x_i=\widehat x_i+\widehat\nu_i$ and using $\sum_i\left(\widehat x_i-\bar x\right)\widehat\nu_i=0$. With additional controls, the same argument applies after removing their linear contribution from the variables. This motivates the partial $R^2$ below.

### Finite-sample bias with valid instruments

Now restore instrument validity. The population first-stage fitted component is uncorrelated with $\epsilon_i$. With a fixed, nonzero first stage, IV is consistent. To study finite-sample bias, however, we must consider the estimated first stage in repeated samples of a given size:

$$\widehat\beta_{IV}-\beta=
\frac{\sum_i\left(\widehat x_i-\bar x\right)\epsilon_i}
{\sum_i\left(\widehat x_i-\bar x\right)^2}.$$

Both parts of this ratio are random. Population orthogonality does not imply that its expectation is zero. In particular, the first-stage fitted values depend on the same observations used to estimate the outcome equation. The expectation of the ratio generally differs from the ratio of expectations.

### Why the bias is related to OLS

Write the first stage with $K$ excluded instruments as

$$x_i=\pi_0+\sum_{j=1}^{K}\pi_jz_{ji}+\nu_i.$$

The instruments are exogenous, so the correlation between $x_i$ and the outcome disturbance comes from $\nu_i$:

$$\operatorname{Cov}\left(x_i,\epsilon_i\right)
=\operatorname{Cov}\left(\nu_i,\epsilon_i\right)
=\sigma_{\epsilon,\nu}.$$

In a particular sample, the instruments will also predict some of the variation in $\nu_i$ by chance. The fitted first stage consequently combines the instruments' true predictive contribution with some fitted noise. Because $\nu_i$ is related to $\epsilon_i$, this fitted noise can introduce the same endogeneity that affects OLS into the second stage. When the instruments are weak, the fitted noise can be a large share of the variation in $\widehat x_i$.

This explains the pull towards OLS in overidentified 2SLS. Adding instruments gives the first-stage regression more opportunities to fit noise. If they add little true predictive power, the finite-sample bias can increase.

A large sample helps when the population first stage is fixed and nonzero: the true predictive contribution accumulates as observations are added. If the population first stage is zero, more observations cannot create identification.

## 3. First-stage evidence

### Partial $R^2$

Suppose we include $L$ controls $w_{ki}$ and three excluded instruments:

$$x_i=\pi_0+\pi_1z_{1i}+\pi_2z_{2i}+\pi_3z_{3i}+\sum_{k=1}^{L}\delta_kw_{ki}+\nu_i.$$

The overall first-stage $R^2$ includes the predictive contribution of the controls. Our question is how much the **excluded instruments** add after accounting for those controls. Estimate a restricted regression containing the controls alone, and an unrestricted regression that adds the excluded instruments. Then

$$\text{partial }R^2=\frac{R_U^2-R_R^2}{1-R_R^2}=\frac{SSR_R-SSR_U}{SSR_R},$$

where $SSR$ is the sum of squared residuals. This measures the proportion of the schooling variation left after the controls that is explained by adding the excluded instruments. A value of 0.00012 means 0.012%, not 1.2%.

### The first-stage F statistic

Test the joint hypothesis

$$H_0:\pi_1=\pi_2=\pi_3=0.$$

The test concerns the excluded instruments. The overall regression test also tests the coefficients on the controls and answers a different question. With one excluded instrument, the corresponding first-stage $F$ is the square of its $t$ statistic.

Report the first-stage coefficients, their standard errors, the excluded-instrument $F$, and partial $R^2$ alongside the IV estimates. Inspect these for each control set because adding controls changes the variation available to identify the effect.

The $F<10$ rule of thumb flags a possible weak-instrument problem. An $F$ above 10 does not guarantee reliable conventional IV inference. The interpretation of critical values depends on the model and the assumptions about the disturbances. A heteroskedasticity-robust joint test and the conventional homoskedastic test need not give identical statistics.

### Relative bias and the first-stage F

Deriving the finite-sample bias of IV is complicated. Interested readers can find the argument in Bound, Jaeger, and Baker (1995).

Under certain assumptions, the bias of IV relative to the bias of OLS is approximately

$$\frac{\operatorname{Bias}\left(\widehat\beta_{IV}\right)}
{\operatorname{Bias}\left(\widehat\beta_{OLS}\right)}\approx\frac{1}{F}.$$

For example, $F=10$ suggests an IV bias roughly one-tenth of the OLS bias.

**This is why we report the first-stage $F$ alongside the IV estimate: it helps us assess how serious finite-sample bias may be.**

The ratio describes the share of OLS bias remaining in IV. For example, if OLS has a bias of 2, a relative bias of 10% corresponds to an IV bias of about 0.2.

This approximation provides guidance for overidentified IV under the assumptions used in the bias calculation. Its accuracy depends on the model and instrument strength. The first-stage $F$ is a diagnostic, and the approximation does not guarantee a particular amount of bias in an application.

### Inference

When instruments are weak, the usual IV coefficient divided by its reported standard error need not have the distribution used to calculate conventional p-values. Heteroskedasticity-consistent standard errors address heteroskedasticity, but do not solve weak identification. A seemingly precise second-stage result can be misleading.

One approach is to test a proposed effect directly. If $\beta=b$ and the instrument is valid, then $y_i-bx_i$ should have no relationship with the excluded instruments after accounting for the controls. An **Anderson–Rubin test** checks this implication. Using a version appropriate to the disturbance structure allows inference that remains valid with weak instruments. A confidence set contains the values of $b$ that are not rejected. It may be very wide or unbounded when the data provide little information about the effect.

## 4. The simulation illustrations

### How the data are generated

We retain the population relationships from the original lecture illustration:

$$x_i=1+\pi_1z_i+2w_i+\nu_i,$$
$$y_i=2+5x_i+20w_i+e_i.$$

The true causal slope is **5**. The omitted variable $w_i$ raises both $x_i$ and $y_i$. Draw $w_i$ and $\nu_i$ independently from standard normal distributions and $e_i$ independently from a normal distribution with standard deviation 10. The binary instrument has probability 0.4 of equalling one and is independent of these draws.

Set $\pi_1=5$ for the strong-instrument example. For the other two examples set $\pi_1=0$, as in the original slides. These illustrate the limiting case of weak relevance: **no population first stage**. We use 1,000 observations in the first two examples and 329,509 in the third.

Each figure shows the true relationship, the sample OLS line and the sample Wald/IV line, with the two instrument-group means marked. Every observation is plotted. The three figures use the same axis scales. Dotted lines project each group mean onto both axes, showing how the first-stage difference collapses when relevance fails.

### Strong instrument: 1,000 observations

![Strong instrument.](../resources/images/week-05/wald-strong.png)

The instrument creates a substantial difference in mean $x_i$. The IV line has slope about 5.13, close to the true slope of 5. OLS has slope about 8.53 because ability is omitted. The first-stage coefficient is about 4.91 and its robust $F$ is about 1,262.

### No population first stage: 1,000 observations

![No population first stage, smaller sample.](../resources/images/week-05/wald-zero-small.png)

The instrument-group means are close horizontally. The estimated IV slope is about −2.18, with a reported standard error of about 25.78. The first-stage $F$ is about 0.39. The sample provides no convincing evidence of relevance.

### No population first stage: 329,509 observations

![No population first stage, large sample.](../resources/images/week-05/wald-zero-large.png)

In this draw, OLS is about 13.00 and IV about 13.80, even though the true effect is 5. The reported IV standard error is about 4.27. Looking only at the second stage, we might regard the positive IV estimate and its proximity to OLS as reassuring. The first-stage $F$ is only about 2.02. The instrument has no population relevance, and the conventional second-stage inference is unreliable.

These are **illustrative draws**, using a fixed seed chosen to show that IV can resemble OLS when relevance fails. Repeating the simulation with different seeds can yield very different IV estimates. Agreement in this figure is not a claim that exactly identified IV converges to OLS when the instrument is irrelevant.

::: {.callout-note title="Your turn"}
Would the second-stage coefficient and reported standard error alone alert you to the problem? What changes when you inspect the first-stage results? Why do 329,509 observations fail to resolve the problem in this example?
:::

### Python code

[Download the Python simulation code](../resources/week-05/weak-instrument-simulations.py). The code generates the three datasets, estimates OLS, first-stage and IV regressions, and saves the figures. Change the seed to examine other samples. The first graph uses schematic densities to illustrate unbiasedness and consistency.

::: {.callout-tip title="Simulation code" collapse="true"}

```python
"""PP5001 Week 5: run beside an empty folder called figures, or let this script create it."""
import pathlib as pl
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import scipy.stats as st
import statsmodels.formula.api as smf
import linearmodels.iv as iv

output = pl.Path('figures')
output.mkdir(exist_ok=True)

# These are schematic sampling distributions to distinguish two properties.
# They are normal curves chosen for illustration, not simulated IV densities.
grid = np.linspace(0, 10, 1500)
fig, axes = plt.subplots(1, 3, figsize=(12, 3.8), sharex=True, sharey=True)
for ax, title, means in zip(axes,
    ['Unbiased and consistent', 'Biased and consistent', 'Biased and inconsistent'],
    [[5, 5, 5], [6.5, 5.6, 5.15], [7, 7, 7]]):
    for mean, sd, label, color in zip(means, [1.2, .7, .35],
        ['Smaller n', 'Larger n', 'Much larger n'], ['#a9cce3', '#2980b9', '#154360']):
        ax.plot(grid, st.norm.pdf(grid, mean, sd), label=label, color=color)
    ax.axvline(5, color='#b03a2e', linestyle='--', label='True effect')
    ax.set_title(title, fontsize=12)
    ax.set_xlabel('Estimated effect')
    ax.set_yticks([])
axes[0].set_ylabel('Sampling density')
axes[1].legend(fontsize=9)
fig.tight_layout()
fig.savefig(output / 'unbiasedness-consistency.png', dpi=200)
plt.close(fig)

# The same population relationships as the original lecture illustration.
# w, e and v are independent, and Z is independent of all three.
# Seed 4 is an illustrative draw showing how IV can resemble OLS without relevance.
# Try other seeds: a zero first stage produces very unstable estimates.
def make_sample(n, first_stage_effect, seed=4):
    rng = np.random.default_rng(seed)
    z = (rng.random(n) < .4).astype(int)
    ability = rng.normal(0, 1, n)
    e = rng.normal(0, 10, n)
    v = rng.normal(0, 1, n)
    d = 1 + first_stage_effect * z + 2 * ability + v
    y = 2 + 5 * d + 20 * ability + e
    return pd.DataFrame({'y': y, 'x': d, 'z': z})

strong = make_sample(1000, 5)
weak_small = make_sample(1000, 0)
weak_large = make_sample(329509, 0)

# Fit each example explicitly. HC1 / robust request consistent standard errors.
# The usual IV standard errors can nevertheless be misleading without relevance.
strong_ols = smf.ols('y ~ x', data=strong).fit(cov_type='HC1')
strong_first = smf.ols('x ~ z', data=strong).fit(cov_type='HC1')
strong_iv = iv.IV2SLS.from_formula('y ~ 1 + [x ~ z]', data=strong).fit(cov_type='robust', debiased=True)
small_ols = smf.ols('y ~ x', data=weak_small).fit(cov_type='HC1')
small_first = smf.ols('x ~ z', data=weak_small).fit(cov_type='HC1')
small_iv = iv.IV2SLS.from_formula('y ~ 1 + [x ~ z]', data=weak_small).fit(cov_type='robust', debiased=True)
large_ols = smf.ols('y ~ x', data=weak_large).fit(cov_type='HC1')
large_first = smf.ols('x ~ z', data=weak_large).fit(cov_type='HC1')
large_iv = iv.IV2SLS.from_formula('y ~ 1 + [x ~ z]', data=weak_large).fit(cov_type='robust', debiased=True)

# Use the same axis limits in all three illustrations.
all_data = pd.concat([strong, weak_small, weak_large])
x_limits = (all_data['x'].min() - 1, all_data['x'].max() + 1)
y_limits = (all_data['y'].min() - 10, all_data['y'].max() + 10)

# Every observation is plotted, including all 329,509 in the large sample.
def draw_example(data, ols_model, iv_model, filename):
    means = data.groupby('z')[['x', 'y']].mean()
    fig, ax = plt.subplots(figsize=(10, 5.8))
    z0 = data.loc[data['z'] == 0]
    z1 = data.loc[data['z'] == 1]
    large = len(data) > 10000
    size, alpha = (2, .035) if large else (13, .4)
    ax.scatter(z0['x'], z0['y'], facecolors='none', edgecolors='#888888',
               s=size, alpha=alpha, linewidths=.5, label='z = 0', rasterized=True)
    ax.scatter(z1['x'], z1['y'], marker='^', color='#888888',
               s=size, alpha=alpha, label='z = 1', rasterized=True)
    grid = np.linspace(*x_limits, 200)
    ax.plot(grid, 2 + 5*grid, color='black', linewidth=2.8,
            label='True relationship: slope 5')
    ax.plot(grid, ols_model.params['Intercept'] + ols_model.params['x']*grid,
            color='#D55E00', linestyle='--', linewidth=2.8,
            label=f"OLS: {ols_model.params['x']:.2f}")
    ax.plot(grid, iv_model.params['Intercept'] + iv_model.params['x']*grid,
            color='#0072B2', linestyle='-.', linewidth=2.8,
            label=f"Wald / IV: {iv_model.params['x']:.2f}")
    ax.set_xlim(x_limits)
    ax.set_ylim(y_limits)
    # Project each group mean onto both axes. Offset labels so close means remain legible.
    for group in [0, 1]:
        mx, my = means.loc[group, ['x', 'y']]
        ax.plot([mx, mx], [y_limits[0], my], ':', color='black', linewidth=1.4)
        ax.plot([x_limits[0], mx], [my, my], ':', color='black', linewidth=1.4)
        offset = -22 if group == 0 else 22
        ax.annotate(r'$\bar{x}_{z=' + str(group) + '}$', xy=(mx, 0),
                    xycoords=('data', 'axes fraction'), xytext=(offset, -28),
                    textcoords='offset points', ha='center', fontsize=10,
                    arrowprops=dict(arrowstyle='-', color='black', lw=.6))
        ax.annotate(r'$\bar{y}_{z=' + str(group) + '}$', xy=(0, my),
                    xycoords=('axes fraction', 'data'), xytext=(-38, offset),
                    textcoords='offset points', ha='right', fontsize=10,
                    arrowprops=dict(arrowstyle='-', color='black', lw=.6))
    ax.scatter(means['x'], means['y'], marker='D', s=70, color='white',
               edgecolor='black', zorder=5, label='Instrument-group means')
    ax.set_xlabel('x', labelpad=32)
    ax.set_ylabel('y', labelpad=48)
    ax.set_title(f'{len(data):,} observations')
    ax.legend(fontsize=9, loc='upper left', framealpha=.95)
    fig.tight_layout()
    fig.savefig(output / filename, dpi=250)
    plt.close(fig)

draw_example(strong, strong_ols, strong_iv, 'wald-strong.png')
draw_example(weak_small, small_ols, small_iv, 'wald-zero-small.png')
draw_example(weak_large, large_ols, large_iv, 'wald-zero-large.png')

results = pd.DataFrame({
    'Example': ['Strong', 'No first stage, small', 'No first stage, large'],
    'N': [len(strong), len(weak_small), len(weak_large)],
    'OLS': [strong_ols.params['x'], small_ols.params['x'], large_ols.params['x']],
    'IV': [strong_iv.params['x'], small_iv.params['x'], large_iv.params['x']],
    'IV SE': [strong_iv.std_errors['x'], small_iv.std_errors['x'], large_iv.std_errors['x']],
    'First-stage coefficient': [strong_first.params['z'], small_first.params['z'], large_first.params['z']],
    'First-stage F': [float(strong_first.f_test('z=0').fvalue),
                      float(small_first.f_test('z=0').fvalue),
                      float(large_first.f_test('z=0').fvalue)],
    'First-stage R2': [strong_first.rsquared, small_first.rsquared, large_first.rsquared],
})
print(results.round(5).to_string(index=False))
results.to_csv(output / 'simulation-results.csv', index=False)
```

:::

## 5. Angrist and Krueger's natural experiment

School-entry rules and compulsory-schooling laws together can create a relationship between quarter of birth and completed schooling. Children entering school at different ages can reach the legal school-leaving age after completing different amounts of education. Angrist and Krueger use quarter of birth as an instrument to estimate the effect of schooling on log weekly earnings.

The 1980 Census sample used here contains **329,509 men born in 1930–1939**. In our notation, $x_i$ is years of completed schooling and $y_i$ is log weekly earnings. With a binary instrument, $z_i=1$ denotes a first-quarter birth and $z_i=0$ a birth in the other three quarters.

::: {.callout-note title="Your turn: The policy and the instrument"}
Explain how school-entry and school-leaving rules generate a first stage. Whose schooling might change because of these rules? What assumptions are needed to interpret the quarter-of-birth comparison as an effect of schooling on earnings?
:::

![Angrist and Krueger (1991), Figure I.](../resources/images/week-05/ak-fig1.png)

![Angrist and Krueger (1991), Figure II.](../resources/images/week-05/ak-fig2.png)

::: {.callout-note title="Your turn: Figures I and II"}
Identify the outcome in each figure. Explain the pattern across quarters within birth years. What do the figures tell us about the proposed mechanism and the comparison underlying IV?
:::

![Angrist and Krueger (1991), Table III, Panel B.](../resources/images/week-05/ak-tab3.png)

The first-quarter group has mean schooling of about 12.6881 years and mean log weekly earnings of about 5.8916. For the other quarters the corresponding means are about 12.7969 and 5.9027. Dividing the earnings difference by the schooling difference gives a Wald estimate of approximately 0.1020. Use the unrounded sample means in the exercise. In a log-earnings equation this is approximately a 10.2% earnings increase for another year of schooling. The exact percentage transformation is $100\left(\exp\left(\widehat\beta\right)-1\right)$.

### Controls and additional instruments

The empirical analysis also uses quarter-of-birth indicators and their interactions with year-of-birth indicators. Write the three quarter indicators as $Q_{1i}$, $Q_{2i}$ and $Q_{3i}$, leaving the fourth quarter as the reference category. With controls $w_{ki}$, the three-instrument specification is

$$x_i=\pi_0+\sum_{q=1}^{3}\pi_q Q_{qi}+\sum_{k=1}^{L}\delta_kw_{ki}+\nu_i,$$
$$y_i=\beta_0+\beta x_i+\sum_{k=1}^{L}\theta_kw_{ki}+\epsilon_i.$$

The same included controls belong in both equations. Additional interactions provide more excluded instruments in the paper's larger specifications.

![Angrist and Krueger (1991), Table V.](../resources/images/week-05/ak-tab5.png)

![Angrist and Krueger (1991), Table VII.](../resources/images/week-05/ak-tab7.png)

::: {.callout-note title="Your turn: Tables V and VII"}
Explain what changes across the specifications, including the controls and excluded instruments. Compare OLS and IV estimates and their reported uncertainty. Which first-stage information would you want before judging the IV results?
:::

## 6. Bound, Jaeger, and Baker

Bound, Jaeger, and Baker revisit the quarter-of-birth application to illustrate two concerns. Small violations of instrument validity can matter greatly when relevance is weak. Separately, even exogenous instruments can lead to poor finite-sample behaviour when they contain little information about the endogenous variable.

![Bound, Jaeger, and Baker (1995), Table 1.](../resources/images/week-05/bjb-tab1.png)

::: {.callout-note title="Your turn: Table 1"}
Compare the three IV specifications. Identify the excluded instruments and the controls. How does your assessment change when you read the first-stage F and partial $R^2$ alongside the second-stage estimates? Pay attention to the scaling of partial $R^2$ in the table.
:::

![Bound, Jaeger, and Baker (1995), Table 2.](../resources/images/week-05/bjb-tab2.png)

![Bound, Jaeger, and Baker (1995), Table 3.](../resources/images/week-05/bjb-tab3.png)

::: {.callout-note title="Your turn: Tables 2 and 3"}
For Table 2, explain what changes when state-of-birth controls and interactions are introduced. For Table 3, explain how the authors simulate quarter of birth and what they hold fixed. What would we expect if the instruments contained no information? Explain what the results reveal about second-stage estimates and the number of instruments.
:::

In the exercise we keep the three quarter-of-birth instruments and vary the controls. Adding age changes the comparison because earnings and schooling can vary with age, while quarter of birth also determines age within a birth year. In BJB's Table 1, columns 1–2 include age and age squared, together with the demographic and region controls. The first-stage F in this specification is about 13.486, and partial $R^2$ is about 0.000123. The table reports the latter multiplied by 100.

## 7. Returning to exclusion

Bound and Jaeger examine whether compulsory-schooling laws alone account for the relationship between earnings and quarter of birth. Their Figure 1 compares third-quarter minus first-quarter differences across birth decades.

![Bound and Jaeger (2000), Figure 1.](../resources/images/week-05/bj-fig1.png)

::: {.callout-note title="Your turn: Exclusion"}
Compare the quarter-of-birth patterns in schooling and earnings across the three birth decades. If schooling were the only channel and its effect were stable, how would you expect the two patterns to relate? What further explanations or evidence would you consider before accepting the exclusion restriction?
:::

This brings us back to the argument behind an instrument. The first-stage statistics assess relevance. Exclusion requires evidence and reasoning about the channels through which the instrument can affect the outcome. Weak relevance makes a small failure of that argument especially consequential for IV.

## After class

[Complete Problem Set 2](../exercises/05-weak-instruments.qmd), due **Friday 30 October 2026 at 12 noon (UK time), in Week 7**. Submit the completed Quarto source and rendered PDF through Moodle. Calculate the Wald estimate from sample means, reproduce it with IV, and estimate the three-instrument specifications with different age controls. Present first-stage results alongside the estimated returns to schooling.

## References

- Angrist, J. D., and A. B. Krueger (1991). “Does Compulsory School Attendance Affect Schooling and Earnings?” *Quarterly Journal of Economics* 106(4): 979–1014. [Paper](https://doi.org/10.2307/2937954).
- Bound, J., D. A. Jaeger, and R. M. Baker (1995). “Problems with Instrumental Variables Estimation When the Correlation Between the Instruments and the Endogenous Explanatory Variable Is Weak.” *Journal of the American Statistical Association* 90(430): 443–450. [Paper](https://www.djaeger.org/research/pubs/jasav90n430.pdf).
- Bound, J., and D. A. Jaeger (2000). “Do Compulsory School Attendance Laws Alone Explain the Association between Earnings and Quarter of Birth?” *Research in Labor Economics* 19: 83–108. [Paper](https://www.djaeger.org/research/pubs/rle2000.pdf).
- Staiger, D., and J. H. Stock (1997). “Instrumental Variables Regression with Weak Instruments.” *Econometrica* 65(3): 557–586. [Paper](https://doi.org/10.2307/2171753).
