To make high-quality research more accessible and easier to explore.

Fields:
24 results

The Effects of 401(K) Participation on the Wealth Distribution: An Instrumental Quantile Regression Analysis

The Review of Economics and Statistics 2004 86(3), 735-751
We use instrumental quantile regression approach to examine the effects of 401(k) plans on wealth using data from the Survey of Income and Program Participation. Using 401(k) eligibility as an instrument for 401(k) participation, we estimate the quantile treatment effects of participation in a 401(k) plan on several measures of wealth. The results show the effects of 401(k) participation on net financial assets are positive and significant over the entire range of the asset distribution, and that the increase in the low tail of the assets distribution appears to translate completely into an increase in wealth. However, there is significant evidence of substitution from other forms of wealth in the upper tail of the distribution.

Distribution Regression with Sample Selection and UK Wage Decomposition

Journal of Political Economy 2025 133(12), 3952-3992
We develop a distribution regression model under endogenous sample selection. This model is a semi-parametric generalization of the Heckman selection model. It accommodates much richer effects of the covariates on outcome distribution and patterns of heterogeneity in the selection process, and allows for drastic departures from the Gaussian error structure, while maintaining the same level tractability as the classical model. The model applies to continuous, discrete and mixed outcomes. We provide identification, estimation, and inference methods, and apply them to obtain wage decomposition for the UK. Here we decompose the difference between the male and female wage distributions into composition, wage structure, selection structure, and selection sorting effects. After controlling for endogenous employment selection, we still find substantial gender wage gap – ranging from 21% to 40% throughout the (latent) offered wage distribution that is not explained by composition. We also uncover positive sorting for single men and negative sorting for married women that accounts for a substantive fraction of the gender wage gap at the top of the distribution.

Inference on Causal and Structural Parameters using Many Moment Inequalities

Review of Economic Studies 2019 86(5), 1867-1900 open access
This article considers the problem of testing many moment inequalities where the number of moment inequalities, denoted by p, is possibly much larger than the sample size n. There is a variety of economic applications where solving this problem allows to carry out inference on causal and structural parameters; a notable example is the market structure model of Ciliberto and Tamer (2009) where p=2^m+1 with m being the number of firms that could possibly enter the market. We consider the test statistic given by the maximum of p Studentized (or t-type) inequality-specific statistics, and analyse various ways to compute critical values for the test statistic. Specifically, we consider critical values based upon (1) the union bound combined with a moderate deviation inequality for self-normalized sums, (2) the multiplier and empirical bootstraps, and (3) two-step and three-step variants of (1) and (2) by incorporating the selection of uninformative inequalities that are far from being binding and a novel selection of weakly informative inequalities that are potentially binding but do not provide first-order information. We prove validity of these methods, showing that under mild conditions, they lead to tests with the error in size decreasing polynomially in n while allowing for p being much larger than n; indeed p can be of order $\exp (n^c)$ for some $c > 0$. Importantly, all these results hold without any restriction on the correlation structure between p Studentized statistics, and also hold uniformly with respect to suitably large classes of underlying distributions. Moreover, in the online supplement, we show validity of a test based on the block multiplier bootstrap in the case of dependent data under some general mixing conditions.

Fisher–Schultz Lecture: Generic Machine Learning Inference on Heterogeneous Treatment Effects in Randomized Experiments, With an Application to Immunization in India

Econometrica 2025 93(4), 1121-1164 open access
We propose strategies to estimate and make inference on key features of heterogeneous effects in randomized experiments. These key features include best linear predictors of the effects using machine learning proxies, average effects sorted by impact groups , and average characteristics of most and least impacted units . The approach is valid in high‐dimensional settings, where the effects are proxied (but not necessarily consistently estimated) by predictive and causal machine learning methods. We post‐process these proxies into estimates of the key features. Our approach is generic; it can be used in conjunction with penalized methods, neural networks, random forests, boosted trees, and ensemble methods, both predictive and causal. Estimation and inference are based on repeated data splitting to avoid overfitting and achieve validity. We use quantile aggregation of the results across many potential splits, in particular taking medians of p ‐values and medians and other quantiles of confidence intervals. We show that quantile aggregation lowers estimation risks over a single split procedure, and establish its principal inferential properties. Finally, our analysis reveals ways to build provably better machine learning proxies through causal learning: we can use the objective functions that we develop to construct the best linear predictors of the effects, to obtain better machine learning proxies in the initial step. We illustrate the use of both inferential tools and causal learners with a randomized field experiment that evaluates a combination of nudges to stimulate demand for immunization in India.

Intersection Bounds: Estimation and Inference

Econometrica 2013 81(2), 667-737 open access
We develop a practical and novel method for inference on intersection bounds, namely bounds defined by either the infimum or supremum of a parametric or nonparametric function, or, equivalently, the value of a linear programming problem with a potentially infinite constraint set. We show that many bounds characterizations in econometrics, for instance bounds on parameters under conditional moment inequalities, can be formulated as intersection bounds. Our approach is especially convenient for models comprised of a continuum of inequalities that are separable in parameters, and also applies to models with inequalities that are nonseparable in parameters. Since analog estimators for intersection bounds can be severely biased in finite samples, routinely underestimating the size of the identified set, we also offer a median-bias-corrected estimator of such bounds as a by-product of our inferential procedures. We develop theory for large sample inference based on the strong approximation of a sequence of series or kernel-based empirical processes by a sequence of “penultimate” Gaussian processes. These penultimate processes are generally not weakly convergent, and thus are non-Donsker. Our theoretical results establish that we can nonetheless perform asymptotically valid inference based on these processes. Our construction also provides new adaptive inequality/moment selection methods. We provide conditions for the use of nonparametric kernel and series estimators, including a novel result that establishes strong approximation for any general series estimator admitting linearization, which may be of independent interest.

Automatic Debiased Machine Learning of Causal and Structural Effects

Econometrica 2022 90(3), 967-1027 open access
Many causal and structural effects depend on regressions. Examples include policy effects, average derivatives, regression decompositions, average treatment effects, causal mediation, and parameters of economic structural models. The regressions may be high‐dimensional, making machine learning useful. Plugging machine learners into identifying equations can lead to poor inference due to bias from regularization and/or model selection. This paper gives automatic debiasing for linear and nonlinear functions of regressions. The debiasing is automatic in using Lasso and the function of interest without the full form of the bias correction. The debiasing can be applied to any regression learner, including neural nets, random forests, Lasso, boosting, and other high‐dimensional methods. In addition to providing the bias correction, we give standard errors that are robust to misspecification, convergence rates for the bias correction, and primitive conditions for asymptotic inference for estimators of a variety of estimators of structural and causal effects. The automatic debiased machine learning is used to estimate the average treatment effect on the treated for the NSW job training data and to estimate demand elasticities from Nielsen scanner data while allowing preferences to be correlated with prices and income.

The Sorted Effects Method: Discovering Heterogeneous Effects Beyond Their Averages

Econometrica 2018 86(6), 1911-1938 open access
This zip file contains the replication files for the manuscript. It also contains an online appendix. The supplementary material contains 7 appendices with additional results and some omitted proofs. Appendix C introduces some notation. Appendix D includes a brief review of differential geometry. Appendix E gathers the proofs of the key mathematical results in Appendix A. Appendix F provides sufficient conditions for the u-Donsker properties in Section 4. Appendix G extends the theoretical analysis to include discrete covariates. Appendices H and I report the results of 3 numerical simulations and an empirical application to the effect of race on mortgage denials, respectively.

Post-Selection and Post-Regularization Inference in Linear Models with Many Controls and Instruments

American Economic Review 2015 105(5), 486-490 open access
We consider estimation of and inference about coefficients on endogenous variables in a linear instrumental variables model where the number of instruments and exogenous control variables are each allowed to be larger than the sample size. We work within an approximately sparse framework that maintains that the signal available in the instruments and control variables may be effectively captured by a small number of the available variables. We provide a LASSO-based method for this setting which provides uniformly valid inference about the coefficients on endogenous variables. We illustrate the method through an application to demand estimation.