To make high-quality research more accessible and easier to explore.

Fields:
30 results

The Interpretation of Instrumental Variables Estimators in Simultaneous Equations Models with an Application to the Demand for Fish

Review of Economic Studies 2000 67(3), 499-527
In markets where prices are determined by the intersection of supply and demand curves, standard identification results require the presence of instruments that shift one curve but not the other. These results are typically presented in the context of linear models with fixed coefficients and additive residuals. The first contribution of this paper is an investigation of the consequences of relaxing both the linearity and the additivity assumption for the interpretation of linear instrumental variables estimators. Without these assumptions, the standard linear instrumental variables estimator identifies a weighted average of the derivative of the behavioural relationship of interest. A second contribution is the formulation of critical identifying assumptions in terms of demand and supply at different prices and instruments, rather than in terms of functional-form specific residuals. Our approach to the simultaneous equations problem and the average-derivative interpretation of instrumental variables estimates is illustrated by estimating the demand for fresh whiting at the Fulton fish market. Strong and credible instruments for identification of this demand function are available in the form of weather conditions at sea.

Efficient Estimation of Average Treatment Effects Using the Estimated Propensity Score

Econometrica 2003 71(4), 1161-1189
We are interested in estimating the average effect of a binary treatment on a scalar outcome. If assignment to the treatment is exogenous or unconfounded, that is, independent of the potential outcomes given covariates, biases associated with simple treatment-control average comparisons can be removed by adjusting for differences in the covariates. Rosenbaum and Rubin (1983) show that adjusting solely for differences between treated and control units in the propensity score removes all biases associated with differences in covariates. Although adjusting for differences in the propensity score removes all the bias, this can come at the expense of efficiency, as shown by Hahn (1998), Heckman, Ichimura, and Todd (1998), and Robins, Mark, and Newey (1992). We show that weighting by the inverse of a nonparametric estimate of the propensity score, rather than the true propensity score, leads to an efficient estimate of the average treatment effect. We provide intuition for this result by showing that this estimator can be interpreted as an empirical likelihood estimator that efficiently incorporates the information about the propensity score.

When Should You Adjust Standard Errors for Clustering?

Quarterly Journal of Economics 2022 138(1), 1-35 open access
Clustered standard errors, with clusters defined by factors such as geography, are widespread in empirical research in economics and many other disciplines. Formally, clustered standard errors adjust for the correlations induced by sampling the outcome variable from a data-generating process with unobserved cluster-level components. However, the standard econometric framework for clustering leaves important questions unanswered: (i) Why do we adjust standard errors for clustering in some ways but not others, for example, by state but not by gender, and in observational studies but not in completely randomized experiments? (ii) Is the clustered variance estimator valid if we observe a large fraction of the clusters in the population? (iii) In what settings does the choice of whether and how to cluster make a difference? We address these and other questions using a novel framework for clustered inference on average treatment effects. In addition to the common sampling component, the new framework incorporates a design component that accounts for the variability induced on the estimator by the treatment assignment mechanism. We show that, when the number of clusters in the sample is a nonnegligible fraction of the number of clusters in the population, conventional clustered standard errors can be severely inflated, and propose new variance estimators that correct for this bias.

Sampling‐Based versus Design‐Based Uncertainty in Regression Analysis

Econometrica 2020 88(1), 265-296 open access
Consider a researcher estimating the parameters of a regression function based on data for all 50 states in the United States or on data for all visits to a website. What is the interpretation of the estimated parameters and the standard errors? In practice, researchers typically assume that the sample is randomly drawn from a large population of interest and report standard errors that are designed to capture sampling variation. This is common even in applications where it is difficult to articulate what that population of interest is, and how it differs from the sample. In this article, we explore an alternative approach to inference, which is partly design‐based. In a design‐based setting, the values of some of the regressors can be manipulated, perhaps through a policy intervention. Design‐based uncertainty emanates from lack of knowledge about the values that the regression outcome would have taken under alternative interventions. We derive standard errors that account for design‐based uncertainty instead of, or in addition to, sampling‐based uncertainty. We show that our standard errors in general are smaller than the usual infinite‐population sampling‐based standard errors and provide conditions under which they coincide.

Estimating Average Treatment Effects: Supplementary Analyses and Remaining Challenges

American Economic Review 2017
There is a large literature on semiparametric estimation of average treatment effects under unconfounded treatment assignment in settings with a fixed number of covariates. More recently attention has focused on settings with a large number of covariates. In this paper we extend lessons from the earlier literature to this new setting. We propose that in addition to reporting point estimates and standard errors, researchers report results from a number of supplementary analyses to assist in assessing the credibility of their estimates.

Estimating the Effect of Unearned Income on Labor Earnings, Savings, and Consumption: Evidence from a Survey of Lottery Players

American Economic Review 2001 91(4), 778-794
This paper provides empirical evidence about the effect of unearned income on earnings, consumption, and savings. Using an original survey of people playing the lottery in Massachusetts in the mid-1980's, we analyze the effects of the magnitude of lottery prizes on economic behavior. The critical assumption is that among lottery winners the magnitude of the prize is randomly assigned. We find that unearned income reduces labor earnings, with a marginal propensity to consume leisure of approximately 11 percent, with larger effects for individuals between 55 and 65 years old. After receiving about half their prize, individuals saved about 16 percent.

The Surrogate Index: Combining Short-Term Proxies to Estimate Long-Term Treatment Effects More Rapidly and Precisely

Review of Economic Studies 2026 93(4), 2284-2312 open access
A common challenge in estimating the impact of interventions (e.g. job training programmes, educational programmes) is that many outcomes of interest (e.g. lifetime earnings or other labour market outcomes) are observed with a long delay. In biomedical settings, this is often addressed by using short-term outcomes as so-called “surrogates” for the outcome of interest, e.g. tumour size as a surrogate for mortality in cancer studies. We build on this literature by combining multiple, possibly qualitatively distinct, short-term outcomes (e.g. short-run earnings and employment indicators) systematically into a “surrogate index”. Under the Prentice surrogacy assumption, which requires that the primary outcome is independent of the treatment conditional on the surrogates, we show that the average treatment effect on the surrogate index equals the treatment effect on the long-term outcome. We also relate the surrogacy assumption to a set of structural, causal assumptions. We then characterize the bias that arises from violations of each of the key assumptions, and we provide simple methods to validate these assumptions using additional observed outcomes. We apply our method to analyse the long-term impacts of a multi-site job training experiment in California. Rather than waiting a full 9 years to directly observe the long-term impact, we show that it is possible to use short-term (the first six quarters) outcomes as surrogates. Given the surrogacy assumption one could have estimated the programme’s long-term impacts on mean employment rates using the employment rates observed in the first six quarters, with a 35% reduction in standard errors relative to a simple difference in means estimator based on all 9 years of data.

Nonparametric Tests for Treatment Effect Heterogeneity

The Review of Economics and Statistics 2008 90(3), 389-405
In this paper we develop two nonparametric tests of treatment effect heterogeneity. The first test is for the null hypothesis that the treatment has a zero average effect for all subpopulations defined by covariates. The second test is for the null hypothesis that the average effect conditional on the covariates is identical for all subpopulations, that is, that there is no heterogeneity in average treatment effects by covariates. We derive tests that are straightforward to implement and illustrate the use of these tests on data from two sets of experimental evaluations of the effects of welfare-to-work programs.

Synthetic Difference-in-Differences

American Economic Review 2021 111(12), 4088-4118 open access
We present a new estimator for causal effects with panel data that builds on insights behind the widely used difference-in-differences and synthetic control methods. Relative to these methods we find, both theoretically and empirically, that this “synthetic difference-in-differences” estimator has desirable robustness properties, and that it performs well in settings where the conventional estimators are commonly used in practice. We study the asymptotic behavior of the estimator when the systematic part of the outcome model includes latent unit factors interacted with latent time factors, and we present conditions for consistency and asymptotic normality.