Robustness? Range Tests for Equivalence and Equality Across Multiple Specifications
Working Paper
This paper develops bootstrap-based tests of equality and equivalence across multiple specifications while accounting for their joint sampling distribution. These tests are used to assess whether the results are robust.
Paper
Download the current version of the paper (29 July 2026)
Abstract
Applied economists routinely compare estimates across specifications, observe that they are “similar,” and conclude that their results are “robust.” This common practice makes an implicit inferential claim about the range of estimates, but the estimates are usually generated from the same data and have a joint sampling distribution that published tables do not show. I formalize this claim with two bootstrap statistics. The minimum equivalence bound, \(R^*_{1-\alpha}\), is the smallest tolerance within which the estimates can be judged equivalent. The range-based equality \(p\)-value, \(p_R\), tests whether the estimates are statistically distinguishable. Together they separate two ideas that informal robustness checks often conflate: failure to detect differences and affirmative evidence of agreement. Because the bootstrap resamples the same observations and re-estimates all specifications in each replication, it captures dependence across estimates without requiring inversion of a potentially near-singular variance-covariance matrix. Simulations show approximately correct size and demonstrate that common visual rules often provide false comfort. Applications to five prominent papers validate some conventional robustness claims while revealing cases in which apparent agreement reflects imprecision rather than stability. In a pre-registered survey of CEPR and NBER affiliates, expert judgments align with the framework in obvious cases but diverge in intermediate cases, where the verdict depends on the joint sampling distribution rather than on the visible spread of point estimates alone. I suggest that \(R^*_{.95}\) and \(p_R\) be reported whenever multiple specifications are presented as evidence of robustness.
Software
Replication Materials
Replication materials will be posted here.
Survey Pre-registration
Citation
If you use this paper, please cite it as:
Jaeger, David A. (2026). “Robustness? Range Tests for Equivalence and Equality Across Multiple Specifications.” Working paper.
@article{jaeger2026robustness,
author = {Jaeger, David A.},
title = {Robustness? Range Tests for Equivalence and Equality Across Multiple Specifications},
year = {2026},
note = {Working Paper}
}