To make high-quality research more accessible and easier to explore.

Fields:

More Data or Better Data? A Statistical Decision Problem

Review of Economic Studies 2017 84(4), 1583-1605
When designing data collection, crucial questions arise regarding how much data to collect and how much effort to expend to enhance the quality of the collected data. To make choice of sample design a coherent subject of study, it is desirable to specify an explicit decision problem. We use the Wald framework of statistical decision theory to study allocation of a budget between two or more sampling processes. These processes all draw random samples from a population of interest and aim to collect data that are informative about the sample realizations of an outcome. They differ in the cost of data collection and the quality of the data obtained. One may incur lower cost per sample member but yield lower data quality than another. Increasing the allocation of budget to a low-cost process yields more data, while increasing the allocation to a high-cost process yields better data. We initially view the concept of “better data” abstractly and then fix attention on two important cases. In both cases, a high-cost sampling process accurately measures the outcome of each sample member. The cases differ in the data yielded by a low-cost process. In one, the low-cost process has non-response and in the other it provides a low-resolution interval measure of each sample member’s outcome. In these settings, we study minimax-regret sample design for prediction of a real-valued outcome under square loss; that is, design which minimizes maximum mean square error. The analysis imposes no assumptions that restrict the unobserved outcomes. Hence, the decision maker must cope with both the statistical imprecision of finite samples and the partial identification of the true state of nature.

Inference with Imputed Data: The Allure of Making Stuff Up

Journal of Labor Economics 2025 43(S1), S333-S350
Incomplete observability of data generates an identification problem. What one can learn about a population parameter depends on the assumptions one finds credible. Rubin has promoted random multiple imputation (RMI) as a general way to deal with missing values. The recommendation has been influential to researchers who seek a simple fix to the nuisance of missing data. This paper provides a transparent assessment of the mix of Bayesian and frequentist thinking used by Rubin to argue for RMI. It evaluates random imputation to replace missing outcome or covariate data when the objective is to learn a conditional expectation.

Communicating Uncertainty in Official Economic Statistics: An Appraisal Fifty Years after Morgenstern

Journal of Economic Literature 2015 53(3), 631-653
Federal statistical agencies in the United States and analogous agencies elsewhere commonly report official economic statistics as point estimates, without accompanying measures of error. Users of the statistics may incorrectly view them as error free or may incorrectly conjecture error magnitudes. This paper discusses strategies to mitigate misinterpretation of official statistics by communicating uncertainty to the public. Sampling error can be measured using established statistical principles. The challenge is to satisfactorily measure the various forms of nonsampling error. I find it useful to distinguish transitory statistical uncertainty, permanent statistical uncertainty, and conceptual uncertainty. I illustrate how each arises as the Bureau of Economic Analysis periodically revises GDP estimates, the Census Bureau generates household income statistics from surveys with nonresponse, and the Bureau of Labor Statistics seasonally adjusts employment statistics. I anchor my discussion of communication of uncertainty in the contribution of Oskar Morgenstern (1963a), who argued forcefully for agency publication of error estimates for official economic statistics.

Social Learning from Private Experiences: The Dynamics of the Selection Problem

Review of Economic Studies 2004 71(2), 443-458
I analyse social interactions that stem from the successive endeavours of new cohorts of heterogeneous decision makers to learn from the experiences of past cohorts. A dynamic process of information accumulation and decision making occurs as the members of each cohort observe the experiences of earlier ones, and then make choices that yield experiences observable by future cohorts. Decision makers face the "selection problem" as they seek to learn from observation of past actions and outcomes, while not observing the counterfactual outcomes that would have occurred had other actions been chosen. Assuming that all cohorts face the same outcome distributions, I show that social learning is a process of sequential reduction in ambiguity. The specific nature of this process, and its terminal state, depend critically on how decision makers make choices under ambiguity. I use the problem of learning about innovations to illustrate. Copyright The Review of Economic Studies Limited, 2004.

Identification of Endogenous Social Effects: The Reflection Problem

Review of Economic Studies 1993 60(3), 531
This paper examines the reflection problem that arises when a researcher observing the distribution of behaviour in a population tries to infer whether the average behaviour in some group influences the behaviour of the individuals that comprise the group. It is found that inference is not possible unless the researcher has prior information specifying the compisition of reference groups. If this information is available, the prospects for inference depend critically on the population relationship between the variables defining reference groups and those directly affecting outcomes. Inference is difficult to implossible if these variables are functionally dependent or are statistically independent. The prospects are better if the variables defining reference groups and those directly affecting outcomes are moderately related in the population.

Measuring Expectations

Econometrica 2004 72(5), 1329-1376
This article discusses the history underlying the new literature, describes some of what has been learned thus far, and looks ahead towards making further progress

Statistical Treatment Rules for Heterogeneous Populations

Econometrica 2004 72(4), 1221-1246
An important objective of empirical research on treatment response is to provide decision makers with information useful in choosing treatments. This paper studies minimax-regret treatment choice using the sample data generated by a classical randomized experiment. Consider a utilitarian social planner who must choose among the feasible statistical treatment rules, these being functions that map the sample data and observed covariates of population members into a treatment allocation. If the planner knew the population distribution of treatment response, the optimal treatment rule would maximize mean welfare conditional on all observed covariates. The appropriate use of covariate information is a more subtle matter when only sample data on treatment response are available. I consider the class of conditional empirical success rules; that is, rules assigning persons to treatments that yield the best experimental outcomes conditional on alternative subsets of the observed covariates. I derive a closed-form bound on the maximum regret of any such rule. Comparison of the bounds for rules that conditional on smaller and larger subsets of the covariates yields sufficient sample sizes for productive use of covariate information. When the available sample size exceeds the sufficiency boundary, a planner can be certain that conditioning treatment choice on more covariates is preferable (in terms of minimax regret) to conditioning on fewer covariates.

Monotone Treatment Response

Econometrica 1997 65(6), 1311
The standard formalization of the econometric analysis of treatment response assumes that each member of a population of interest receives one of a set of mutually exclusive and exhaustive treatments, and that the outcome under the realized treatment is observable. Outcomes under the nonrealized treatments are necessarily unobservable; hence these outcomes are censored. This paper investigates what may be learned about treatment response when it is assumed that response functions are monotone, semi-monotone, or concave-monotone. The analysis assumes nothing about the process of treatment selection and imposes no cross-individual restrictions on response. The basic idea is to determine, for every member of the population, the set of response functions that pass through that person's realized (treatment, outcome) pair and that are consistent with the functional-form assumption imposed. These person-specific findings are then explicitly aggregated across the population to determine what can be learned about the distribution of response. The findings have application to the econometric analysis of market demand and of production.

Semiparametric Analysis of Random Effects Linear Models from Binary Panel Data

Econometrica 1987 55(2), 357
[Andersen (1970) considered the problem of inference on random effects linear models from binary response panel data. He showed that inference is possible if the disturbances for each panel member are known to be white noise with logistic distribution and if the observed explanatory variables vary over time. A conditional maximum likelihood estimator consistently estimates the model parameters up to scale. The present paper shows that inference remains possible if the disturbances for each panel member are known only to be time-stationary with unbounded support and if the explanatory variables vary enough over time. A conditional version of the maximum score estimator (Manski, 1975, 1985) consistently estimates the model parameters up to scale.]