Formulae
A reference for the main formulae in the completed weeks. Each section links to the notes for the derivations and examples. See Greek letters for the symbols and Python for estimation code.
Expectations refer to population quantities. Hats denote sample estimates, and bars denote sample means. The meaning of a symbol follows the model in its section.
Background: Population quantities and sample analogues
The analogy principle replaces population expectations with sample averages. For observations \(\left(X_i,Y_i\right)\), \(i=1,\ldots,n\), define \(\mu_X=E[X_i]\) and \(\mu_Y=E[Y_i]\). The sample column below uses \(1/n\) throughout.
| Quantity | Population | Sample analogue |
|---|---|---|
| Mean | \(\mu_X=E[X_i]\) | \(\overline X=\frac{1}{n}\sum_{i=1}^{n}X_i\) |
| Variance | \(\sigma_X^2=E\left[\left(X_i-\mu_X\right)^2\right]\) | \(\widehat\sigma_X^2=\frac{1}{n}\sum_{i=1}^{n}\left(X_i-\overline X\right)^2\) |
| Standard deviation | \(\sigma_X=\sqrt{\sigma_X^2}\) | \(\widehat\sigma_X=\sqrt{\widehat\sigma_X^2}\) |
| Covariance | \(\sigma_{X,Y}=E\left[\left(X_i-\mu_X\right)\left(Y_i-\mu_Y\right)\right]\) | \(\widehat\sigma_{X,Y}=\frac{1}{n}\sum_{i=1}^{n}\left(X_i-\overline X\right)\left(Y_i-\overline Y\right)\) |
| Correlation | \(\rho_{X,Y}=\frac{\sigma_{X,Y}}{\sigma_X\sigma_Y}\) | \(r_{X,Y}=\frac{\widehat\sigma_{X,Y}}{\widehat\sigma_X\widehat\sigma_Y}\) |
Variance and covariance require finite second moments. Correlation also requires positive variances for both variables.
The correction from \(n\) to \(n-1\)
For independent, identically distributed observations and \(n>1\), the usual unbiased estimators of population variance and covariance are
\[s_X^2=\frac{1}{n-1}\sum_{i=1}^{n}\left(X_i-\overline X\right)^2 =\frac{n}{n-1}\widehat\sigma_X^2,\]
\[s_{X,Y}=\frac{1}{n-1}\sum_{i=1}^{n}\left(X_i-\overline X\right)\left(Y_i-\overline Y\right) =\frac{n}{n-1}\widehat\sigma_{X,Y}.\]
Estimating the means from the same observations makes the \(1/n\) variance estimator downward biased and multiplies the expected covariance by \(\left(n-1\right)/n\). The correction above removes that bias. Taking a square root gives the usual sample standard deviation \(s_X\), although \(s_X\) itself is generally biased for \(\sigma_X\).
The common correction cancels in the covariance-to-variance ratio used for the OLS slope:
\[\frac{s_{X,Y}}{s_X^2}=\frac{\widehat\sigma_{X,Y}}{\widehat\sigma_X^2}.\]
It also cancels in the sample correlation when the same convention is used for covariance and both variances.
Week 1: Selection bias and experiments
Potential outcomes and selection bias
For binary treatment \(D_i\), the observed outcome is
\[Y_i=D_iY_i^1+\left(1-D_i\right)Y_i^0.\]
The average treatment effect and the average treatment effect on the treated are
\[\mathrm{ATE}=E\left[Y_i^1-Y_i^0\right],\qquad \mathrm{ATT}=E\left[Y_i^1-Y_i^0\mid D_i=1\right].\]
The observed difference in population means decomposes as
\[\begin{aligned} E\left[Y_i\mid D_i=1\right]-E\left[Y_i\mid D_i=0\right] &=\mathrm{ATT}\\ &\quad+\underbrace{E\left[Y_i^0\mid D_i=1\right]-E\left[Y_i^0\mid D_i=0\right]}_{\text{selection bias}}. \end{aligned}\]
Random assignment of treatment makes the selection-bias term zero and makes ATT equal ATE. This interpretation assumes well-defined treatments and no interference between individuals.
OLS and omitted-variable bias
With an intercept and no additional controls, the population short-regression slope and its sample analogue are:
| Method | Population | Sample estimator |
|---|---|---|
| OLS slope | \(\beta_1^S=\frac{\operatorname{Cov}\left(D_i,Y_i\right)}{\operatorname{Var}\left(D_i\right)}\) | \(\widehat\beta_1^S=\frac{\sum_i\left(D_i-\overline D\right)\left(Y_i-\overline Y\right)}{\sum_i\left(D_i-\overline D\right)^2}\) |
Each denominator requires variation in the explanatory variable.
With an intercept and one explanatory variable,
\[\widehat\beta_1= \frac{\sum_i\left(D_i-\overline D\right)\left(Y_i-\overline Y\right)} {\sum_i\left(D_i-\overline D\right)^2},\qquad \widehat\beta_0=\overline Y-\widehat\beta_1\overline D.\]
For binary \(D_i\), the slope is \(\overline Y_{D=1}-\overline Y_{D=0}\).
Suppose the true population relationship and the auxiliary regression are
\[Y_i=\beta_0+\beta_1D_i+\beta_2A_i+\varepsilon_i, \qquad E\left[\varepsilon_i\mid D_i,A_i\right]=0,\]
\[A_i=\gamma_0+\gamma_1D_i+v_i,\qquad \gamma_1=\frac{\operatorname{Cov}\left(D_i,A_i\right)}{\operatorname{Var}\left(D_i\right)}.\]
Omitting ability \(A_i\) gives the population short-regression slope
\[\boxed{\beta_1^S=\beta_1+\beta_2\gamma_1.}\]
The superscript \(S\) denotes the short regression. Its difference from the causal slope is \(\beta_2\gamma_1\).
Power and sample size
For a specified alternative, power is \(1-\beta\), where \(\beta\) is the probability of a Type II error. With independent observations, common outcome variance \(\sigma^2\), and a two-sided test at level \(\alpha\), the normal approximation gives
\[\mathrm{MDE}=\left(z_{1-\alpha/2}+z_{1-\beta}\right) \sqrt{\frac{\sigma^2}{n_T}+\frac{\sigma^2}{n_C}}.\]
Here \(z_p\) is the \(p\) quantile of the standard normal distribution. With equal treatment and control arms, the required size per arm to detect an effect \(\Delta\) is
\[n=2\left(z_{1-\alpha/2}+z_{1-\beta}\right)^2\frac{\sigma^2}{\Delta^2}.\]
Week 2: Matching and weighting
Propensity scores
The propensity score is the probability of treatment conditional on observed pre-treatment characteristics:
\[p\left(X_i\right)=P\left(D_i=1\mid X_i\right).\]
For a single explanatory variable, the logit model is
\[p\left(X_i\right)=\frac{1}{1+\exp\left(-\left(\gamma_0+\gamma_1X_i\right)\right)}.\]
To identify ATT using untreated comparisons, we assume that conditional on \(X_i\), untreated potential outcomes have the same mean for treated and untreated individuals. We also need controls with comparable characteristics throughout the support of the treated group.
Exact matching
Within a covariate group \(x\),
\[\widehat\tau\left(x\right)=\overline Y_{D=1,X=x}-\overline Y_{D=0,X=x}.\]
Weight these differences by the distribution of characteristics among treated observations:
\[\widehat\tau_{\mathrm{ATT}}=\sum_x\frac{N_{1x}}{N_1}\widehat\tau\left(x\right).\]
Here \(N_{1x}\) is the number of treated observations in group \(x\) and \(N_1\) is the total number treated.
Inverse-probability weighting for ATT
Using estimated propensity scores \(\widehat p_i\), assign weights
\[w_i=\begin{cases} 1 & D_i=1,\\ \widehat p_i/\left(1-\widehat p_i\right) & D_i=0. \end{cases}\]
The normalised weighted difference in means is
\[\widehat\tau_{\mathrm{ATT}} =\frac{1}{N_1}\sum_{i:D_i=1}Y_i -\frac{\sum_{i:D_i=0}w_iY_i}{\sum_{i:D_i=0}w_i}.\]
It is also the treatment coefficient from weighted least squares with an intercept and treatment indicator only:
\[\min_{b_0,b_1}\sum_iw_i\left(Y_i-b_0-b_1D_i\right)^2.\]
OLS uses the same objective with all weights equal to one. Very large ATT weights signal that a few controls may dominate the comparison.
Weeks 3–4: Instrumental variables and LATE
In the omitted-ability model, the population short-regression slope differs from the causal effect by
\[\beta_1^S-\beta_1=\beta_2\gamma_1,\qquad \gamma_1=\frac{\operatorname{Cov}\left(D_i,A_i\right)}{\operatorname{Var}\left(D_i\right)}.\]
The finite-sample bias of the OLS estimator relative to the causal effect is
\[E\left[\widehat\beta_1^S\right]-\beta_1.\]
To connect the two, Week 3 writes
\[Y_i=\beta_0+\beta_1D_i+\beta_2A_i+\nu_i\]
and conditions on all the observed schooling values, \(\mathbf D=\left(D_1,\ldots,D_n\right)\). Assume
\[E\left[\nu_i\mid\mathbf D\right]=0,\qquad E\left[A_i\mid\mathbf D\right]=\gamma_0+\gamma_1D_i.\]
For samples with variation in \(D_i\), these assumptions imply
\[E\left[\widehat\beta_1^S\mid\mathbf D\right]-\beta_1=\beta_2\gamma_1.\]
Averaging over possible schooling samples gives \(E\left[\widehat\beta_1^S\right]-\beta_1=\beta_2\gamma_1\), provided the expectation exists. The linear conditional-mean assumption is what allows this population OVB expression to describe the finite-sample bias as well.
Wald and the IV estimator
In \(Y_i=\beta_0+\beta_1D_i+u_i\), instrument validity requires orthogonality, \(\operatorname{Cov}\left(Z_i,u_i\right)=0\), and relevance, \(\operatorname{Cov}\left(Z_i,D_i\right)\ne0\).
For a binary instrument, Wald divides the difference in conditional outcome means by the difference in conditional treatment means. The sample estimator replaces population expectations with sample means. In the constant-effect model, orthogonality and relevance identify the causal slope through the following ratios.
| Method | Population | Sample estimator |
|---|---|---|
| Wald, binary instrument | \(\beta_1=\frac{E\left[Y_i\mid Z_i=1\right]-E\left[Y_i\mid Z_i=0\right]}{E\left[D_i\mid Z_i=1\right]-E\left[D_i\mid Z_i=0\right]}\) | \(\widehat\beta_{1,\mathrm{Wald}}=\frac{\overline Y_{Z=1}-\overline Y_{Z=0}}{\overline D_{Z=1}-\overline D_{Z=0}}\) |
| Exactly identified IV, one instrument | \(\beta_1=\frac{\operatorname{Cov}\left(Z_i,Y_i\right)}{\operatorname{Cov}\left(Z_i,D_i\right)}\) | \(\widehat\beta_{1,\mathrm{IV}}=\frac{\sum_i\left(Z_i-\overline Z\right)\left(Y_i-\overline Y\right)}{\sum_i\left(Z_i-\overline Z\right)\left(D_i-\overline D\right)}\) |
Each denominator must be nonzero. With a binary instrument, the Wald and IV sample formulas give the same estimate. With heterogeneous effects and binary treatment, the assumptions in the LATE section give the population ratio a complier-effect interpretation.
First stage, reduced form and 2SLS
With exogenous controls \(X_{ik}\),
\[\begin{aligned} D_i&=\pi_0+\pi_1Z_i+\sum_{k=1}^{K}\pi_{k+1}X_{ik}+v_i,\\ Y_i&=\delta_0+\delta_1Z_i+\sum_{k=1}^{K}\delta_{k+1}X_{ik}+e_i. \end{aligned}\]
With one excluded instrument and one endogenous explanatory variable, the model is exactly identified:
\[\widehat\beta_{1,\mathrm{IV}}=\frac{\widehat\delta_1}{\widehat\pi_1}.\]
Use the same observations, controls and weights in both regressions. Two-stage least squares first predicts \(D_i\) from all the exogenous variables, then regresses \(Y_i\) on the predicted \(D_i\) and the controls. The two-stage procedure below writes out both stages explicitly and also applies when there is just one excluded instrument. Use an IV routine for the standard errors.
Overidentified IV: the two-stage procedure
With one endogenous explanatory variable and more than one excluded instrument, the model is overidentified. Two-stage least squares proceeds as follows:
- Regress \(D_i\) on all the excluded instruments and all the exogenous controls, including an intercept. Obtain \(\widehat D_i\).
- Regress \(Y_i\) on \(\widehat D_i\) and the same exogenous controls, including an intercept.
The fitted first stage with \(J\) excluded instruments and \(K\) exogenous controls is
\[\widehat D_i=\widehat\pi_0+\sum_{j=1}^{J}\widehat\pi_jZ_{ij} +\sum_{k=1}^{K}\widehat\pi_{J+k}X_{ik}.\]
The estimated second-stage regression is
\[Y_i=\widehat\beta_{0,\mathrm{2SLS}}+\widehat\beta_{1,\mathrm{2SLS}}\widehat D_i +\sum_{k=1}^{K}\widehat\beta_{k+1,\mathrm{2SLS}}X_{ik}+\widehat r_i.\]
Here \(\widehat r_i\) is the second-stage regression residual. For IV standard errors, the outcome-equation residual is computed using observed \(D_i\) in place of \(\widehat D_i\).
Use an IV routine to estimate the model and obtain the appropriate standard errors. Report the excluded instruments and their joint first-stage test. The ratio of a single reduced-form coefficient to a single first-stage coefficient applies to the exactly identified case above.
LATE
For binary treatment and instrument, let \(D_i^1\) and \(D_i^0\) be treatment status under the two instrument values. Under independence, exclusion, relevance and monotonicity, with well-defined treatments and no interference,
\[\frac{E\left[Y_i\mid Z_i=1\right]-E\left[Y_i\mid Z_i=0\right]} {E\left[D_i\mid Z_i=1\right]-E\left[D_i\mid Z_i=0\right]} =E\left[Y_i^1-Y_i^0\mid D_i^1>D_i^0\right] =\tau_{\mathrm{LATE}}.\]
The expectation on the right is the average treatment effect for compliers. If assignment is random only conditional on specified characteristics, the analysis must account for those characteristics.
Leave-one-out leniency
For a simple illustration, let \(j\left(i\right)\) identify the official assigned to individual \(i\). A leave-one-out treatment rate is
\[Z_i=\frac{\sum_{\ell\ne i:\,j\left(\ell\right)=j\left(i\right)}D_\ell} {n_{j\left(i\right)}-1}.\]
Here \(n_j\) counts cases assigned to official \(j\). The papers refine this construction to account for the assignment process and repeated cases. Leaving out the individual’s own treatment decision avoids a mechanical relationship between their treatment and the instrument. Instrument validity also depends on the assignment process, exclusion and monotonicity.
Week 5: Weak instruments
This section follows BJB’s notation: \(x_i\) is schooling, \(y_i\) is log earnings, and \(\epsilon_i\) is the outcome disturbance.
Unbiasedness and consistency
Unbiasedness concerns the mean across repeated samples of a given size:
\[E\left[\widehat\beta\right]=\beta.\]
Consistency means that for every fixed \(a>0\),
\[P\left(\left|\widehat\beta-\beta\right|>a\right)\longrightarrow0 \quad\text{as }n\longrightarrow\infty.\]
We write \(\operatorname{plim}\widehat\beta=\beta\). A probability limit can also be a value other than the true parameter.
Consequences of endogeneity and imperfect instruments
For the simple model and under the moment and sampling conditions discussed in the notes,
\[\operatorname{plim}\widehat\beta_{OLS}-\beta =\frac{\sigma_{x,\epsilon}}{\sigma_x^2},\]
\[\operatorname{plim}\widehat\beta_{IV}-\beta =\frac{\sigma_{z,\epsilon}}{\sigma_{z,x}} =\frac{\rho_{z,\epsilon}}{\rho_{z,x}}\frac{\sigma_\epsilon}{\sigma_x}.\]
The IV expression assumes fixed, nonzero population relevance. Small correlation between the instrument and schooling can magnify a violation of orthogonality.
First-stage diagnostics
For a first-stage regression with an intercept,
\[R^2=1-\frac{\sum_i\widehat\nu_i^2}{\sum_i\left(x_i-\overline x\right)^2}.\]
To test \(q\) excluded instruments jointly, the conventional homoskedastic first-stage statistic is
\[F=\frac{\left(SSR_R-SSR_U\right)/q}{SSR_U/\left(n-k_U\right)}.\]
\(SSR_R\) is the residual sum of squares when the excluded instruments are omitted, \(SSR_U\) includes them, and \(k_U\) counts all coefficients in the unrestricted first stage, including the intercept. Both regressions contain the same controls. Robust or clustered tests require the corresponding variance calculation.
The approximation discussed in BJB connects relative finite-sample bias to first-stage strength:
\[\frac{\operatorname{Bias}\left(\widehat\beta_{IV}\right)} {\operatorname{Bias}\left(\widehat\beta_{OLS}\right)}\approx\frac{1}{F}.\]
This is an approximation under particular conditions. The reported first-stage \(F\) is a diagnostic of instrument strength, and does not establish the exclusion restriction.
Week 7: Regression discontinuity
Let \(X_i\) be the running variable, \(x_0\) the cutoff, and \(R_i=X_i-x_0\). This section uses \(\rho\) for the treatment effect at the cutoff.
Fuzzy RD
When eligibility changes treatment probability at the cutoff, the fuzzy RD estimand is
\[\rho_{FRD}=\frac{ \lim_{x\downarrow x_0}E\left[Y_i\mid X_i=x\right]-\lim_{x\uparrow x_0}E\left[Y_i\mid X_i=x\right] }{ \lim_{x\downarrow x_0}E\left[D_i\mid X_i=x\right]-\lim_{x\uparrow x_0}E\left[D_i\mid X_i=x\right] }.\]
Under continuity, exclusion, monotonicity and a nonzero first-stage jump, this identifies the treatment effect for compliers at the cutoff. Retain the well-defined-treatment and no-interference assumptions.
Estimating fuzzy RD with IV
Instrument actual treatment \(D_i\) with the cutoff indicator
\[Z_i=1\left(X_i\ge x_0\right)=1\left(R_i\ge0\right).\]
The instrument is a deterministic function of the running variable. It records eligibility, while \(D_i\) records actual treatment. Identification comes from continuity at the cutoff and the discontinuous change in treatment probability induced by eligibility.
Using observations within the chosen bandwidth, the local linear first stage is
\[D_i=\pi_0+\pi_1Z_i+\pi_2R_i+\pi_3Z_i\cdot R_i+v_i.\]
Obtain predicted treatment from this regression:
\[\widehat D_i=\widehat\pi_0+\widehat\pi_1Z_i+\widehat\pi_2R_i+\widehat\pi_3Z_i\cdot R_i.\]
The estimated second stage is
\[Y_i=\widehat\beta_0+\widehat\rho\widehat D_i+\widehat\beta_1R_i+\widehat\gamma_1Z_i\cdot R_i+\widehat r_i.\]
The coefficient \(\widehat\rho\) is the fuzzy RD estimate. Include \(R_i\) and \(Z_i\cdot R_i\) as controls in both stages to allow different slopes on the two sides. The cutoff indicator \(Z_i\) is the excluded instrument. Use the same observations and any kernel weights in both stages.
This specification is exactly identified. Its IV estimate equals the ratio of the estimated reduced-form jump to the estimated first-stage jump, using the same local specification. Use an IV routine for standard errors, based on the outcome-equation residual using observed \(D_i\). Report the first-stage jump alongside the IV estimate.
Density continuity
The manipulation test examines
\[H_0:\quad \lim_{x\uparrow x_0}f_X(x)=\lim_{x\downarrow x_0}f_X(x).\]
Here \(f_X\) is the density of the running variable. This is a diagnostic for sorting around the cutoff. The identifying continuity assumption concerns the conditional means of the potential outcomes.
Week 8: Basic difference-in-differences
Week 8 notes. Let \(T_i\) identify the treatment group and \(Post_t\) the post-policy period.
Parallel trends and the effect
\[E\left[Y_{i,post}^0-Y_{i,pre}^0\mid T_i=1\right] =E\left[Y_{i,post}^0-Y_{i,pre}^0\mid T_i=0\right].\]
With no anticipation, parallel trends identifies the post-policy ATT from the difference of changes. Its sample analogue is
\[\widehat\delta=(\bar Y_{T,post}-\bar Y_{T,pre})-(\bar Y_{C,post}-\bar Y_{C,pre}).\]
Levels and differences
\[Y_{it}=\alpha+\gamma T_i+\tau Post_t+\delta(T_i\cdot Post_t)+u_{it}.\]
For a two-period panel, differencing gives
\[\Delta Y_i=\tau+\delta T_i+\Delta u_i.\]
Baseline covariates
Allow baseline covariates \(X_{ik}\) to predict both levels and changes:
\[Y_{it}=\alpha+\gamma T_i+\tau Post_t+\delta(T_i\cdot Post_t) +\sum_{k=1}^{K}\beta_kX_{ik}+\sum_{k=1}^{K}\theta_k(X_{ik}\cdot Post_t)+u_{it}.\]
\[\Delta Y_i=\tau+\delta T_i+\sum_{k=1}^{K}\theta_kX_{ik}+\Delta u_i.\]
The \(\beta_k\) terms cancel because the baseline covariates are constant across the two observations. The \(\theta_k\) terms remain. Causal interpretation requires an appropriate conditional parallel-trends assumption and the regression specification to represent those conditional changes adequately.
Week 9: TWFE and event studies
Two-way fixed effects
\[Y_{it}=\alpha_i+\lambda_t+\delta D_{it}+u_{it}.\]
\(\alpha_i\) captures time-invariant unit differences, \(\lambda_t\) captures common calendar-year effects, and \(D_{it}\) indicates treatment in that unit-year.
Calendar time and event time
If unit \(i\) adopts in calendar year \(G_i\), its event time is \(k=t-G_i\). Event time zero is the adoption year. Treatment remains on in subsequent years, while the event-time indicator changes each year.
For the common-adoption-date presentation in the notes, \(G_i=G\) for treated units:
\[Y_{it}=\alpha_i+\lambda_t+ \sum_{\substack{k=-K\\k\ne-1}}^{L}\beta_k \left(T_i\cdot\mathbf{1}(t-G=k)\right)+u_{it}.\]
Year −1 is the reference period. The indicators cover the observed treated event times, with endpoint bins if specified. \(\lambda_t\) continues to refer to calendar time. With staggered adoption, replace \(G\) by \(G_i\) for treated units and set the indicators to zero for never-treated units. Interpreting the resulting TWFE coefficients requires care when effects differ across adoption cohorts or time since treatment.
Week 10: Modern difference-in-differences
Effects by adoption cohort and calendar time
Let \(G_i=g\) mean that unit \(i\) first receives treatment in year \(g\). Write \(G_i=\infty\) for never-treated units. For \(t\ge g\),
\[ATT(g,t)=E\left[Y_{it}^{g}-Y_{it}^{\infty}\mid G_i=g\right].\]
\(Y_{it}^{g}\) is the outcome if treatment starts in year \(g\), and \(Y_{it}^{\infty}\) is the outcome if the unit remains untreated. Thus \(ATT(2004,2007)\) is the effect in 2007 for the cohort first treated in 2004.
Identification and sample analogue
Using the last pre-treatment year \(g-1\) as the baseline, parallel trends with never-treated controls requires
\[E\left[Y_{it}^{\infty}-Y_{i,g-1}^{\infty}\mid G_i=g\right] =E\left[Y_{it}^{\infty}-Y_{i,g-1}^{\infty}\mid G_i=\infty\right].\]
With no anticipation, this gives
\[\begin{aligned} ATT(g,t)={}&E\left[Y_{it}-Y_{i,g-1}\mid G_i=g\right]\\ &-E\left[Y_{it}-Y_{i,g-1}\mid G_i=\infty\right]. \end{aligned}\]
For a panel containing the same units in both periods, the sample analogue is
\[\widehat{ATT}(g,t) =\left(\bar Y_{g,t}-\bar Y_{g,g-1}\right) -\left(\bar Y_{\infty,t}-\bar Y_{\infty,g-1}\right).\]
The first subscript on each mean identifies the adoption cohort, and the second identifies calendar time. Not-yet-treated units with \(G_i>t\) can also supply the comparison under the corresponding parallel-trends and no-anticipation assumptions.
Averaging at a common event time
At event time \(k=t-g\ge0\),
\[\theta_k=\sum_{g\in\mathcal G_k}w_{g,k}\,ATT(g,g+k), \qquad w_{g,k}\ge0,\qquad\sum_{g\in\mathcal G_k}w_{g,k}=1.\]
\(\mathcal G_k\) contains cohorts observed at event time \(k\) with a valid comparison group. With cohort-size weights, larger cohorts receive more weight. Changes across event times can reflect both changing effects and changes in which cohorts contribute.
Week 11: Synthetic controls
Week 11 notes · Python implementation.
Counterfactual and estimated effect
Unit 1 receives treatment. Units \(j=2,\ldots,J+1\) form the untreated donor pool. Treatment begins in period \(T_0+1\).
\[\widehat{Y}_{1t}^{0}=\sum_{j=2}^{J+1}w_jY_{jt}, \qquad w_j\ge0,\qquad\sum_{j=2}^{J+1}w_j=1.\]
The donor weights are chosen using pre-treatment information and then held fixed across time. For post-treatment periods,
\[\widehat{\tau}_{1t}=Y_{1t}-\widehat{Y}_{1t}^{0}.\]
Interpreting this gap as a causal effect requires the synthetic path to approximate the treated unit’s outcome without the intervention, including the absence of spillovers affecting the donor outcomes.
Pre- and post-treatment fit
Let \(g_{it}\) be unit \(i\)’s observed outcome minus its synthetic outcome. Apply the same construction to the treated unit and the placebo units. With \(T_1=T-T_0\),
\[\operatorname{MSPE}_{i}^{\mathrm{pre}} =\frac{1}{T_0}\sum_{t=1}^{T_0}g_{it}^{2}, \qquad \operatorname{MSPE}_{i}^{\mathrm{post}} =\frac{1}{T_1}\sum_{t=T_0+1}^{T}g_{it}^{2}.\]
\[R_i=\frac{\operatorname{MSPE}_{i}^{\mathrm{post}}} {\operatorname{MSPE}_{i}^{\mathrm{pre}}}.\]
A larger ratio indicates a greater deterioration in fit after treatment relative to before treatment. The ratio requires a positive pre-treatment MSPE. For the German exercise, the summary windows are 1960–1989 and 1991–2003, omitting 1990, with the denominators adjusted to the number of years in each window.
Placebo rank
\[p=\frac{\text{number of units with }R_i\ge R_1}{J+1}.\]
Include the treated unit in both counts. The smallest possible rank fraction is \(1/(J+1)\). An exact randomization interpretation requires an appropriate treatment-assignment mechanism. In the application, the rank summarizes how unusual the treated unit’s result is within the chosen comparison set.