Journal of Political Economy2018126(6), 2179-2223open access
This paper models decisions to apply to and attend charter schools in Boston using a generalized Roy selection framework linking preferences to the achievement gains generated by charter attendance. The model is estimated with instruments based on randomized admission lotteries and distance to charter schools. Charter schools generate larger gains for disadvantaged students, but demand for charters is stronger among more advantaged students. Similarly, gains are inversely related to unobserved preferences for charters. As a result, counterfactual simulations indicate that charter expansion is likely to be most effective when accompanied by efforts to target students who are unlikely to apply.
This paper develops methods for detecting discrimination by individual employers using correspondence experiments that send fictitious resumes to real job openings. We establish identification of higher moments of the distribution of job‐level callback rates as a function of the number of resumes sent to each job and propose shape‐constrained estimators of these moments. Applying our methods to three experimental data sets, we find striking job‐level heterogeneity in the extent to which callback probabilities differ by race or sex. Estimates of higher moments reveal that while most jobs barely discriminate, a few discriminate heavily. These moment estimates are then used to bound the share of jobs that discriminate and the posterior probability that each individual job is engaged in discrimination. In a recent experiment manipulating racially distinctive names, we find that at least 85% of jobs that contact both of two white applications and neither of two black applications are engaged in discrimination. To assess the potential value of our methods for regulators, we consider the accuracy of decision rules for investigating suspicious callback behavior in various experimental designs under a simple two‐type model that rationalizes the experimental data. Though we estimate that only 17% of employers discriminate on the basis of race, we find that an experiment sending 10 applications to each job would enable detection of 7–10% of discriminatory jobs while yielding Type I error rates below 0.2%. A minimax decision rule acknowledging partial identification of the distribution of callback rates yields only slightly fewer investigations than a Bayes decision rule based on the two‐type model. These findings suggest illegal labor market discrimination can be reliably monitored with relatively small modifications to existing correspondence designs.
Quarterly Journal of Economics2016131(4), 1795-1848open access
We use data from the Head Start Impact Study (HSIS) to evaluate the cost-effectiveness of Head Start, the largest early childhood education program in the United States. Head Start draws roughly a third of its participants from competing preschool programs, many of which receive public funds. We show that accounting for the fiscal impacts of such program substitution pushes estimates of Head Start’s benefit-cost ratio well above one under a wide range of assumptions on the structure of the market for preschool services and the dollar value of test score gains. To parse the program’s test score impacts relative to home care and competing preschools, we selection-correct test scores in each care environment using excluded interactions between experimental assignments and household characteristics. We find that Head Start generates larger test score gains for children who would not otherwise attend preschool and for children who are less likely to participate in the program.
Structural econometric methods are often criticized for being sensitive to functional form assumptions. We study parametric estimators of the local average treatment effect (LATE) derived from a widely used class of latent threshold crossing models and show they yield LATE estimates algebraically equivalent to the instrumental variables (IV) estimator. Our leading example is Heckman's (1979) two‐step (“Heckit”) control function estimator which, with two‐sided non‐compliance, can be used to compute estimates of a variety of causal parameters. Equivalence with IV is established for a semiparametric family of control function estimators and shown to hold at interior solutions for a class of maximum likelihood estimators. Our results suggest differences between structural and IV estimates often stem from disagreements about the target parameter rather than from functional form assumptions per se. In cases where equivalence fails, reporting structural estimates of LATE alongside IV provides a simple means of assessing the credibility of structural extrapolation exercises.
Quarterly Journal of Economics2022137(4), 1963-2036open access
We study the results of a massive nationwide correspondence experiment sending more than 83,000 fictitious applications with randomized characteristics to geographically dispersed jobs posted by 108 of the largest U.S. employers. Distinctively Black names reduce the probability of employer contact by 2.1 percentage points relative to distinctively white names. The magnitude of this racial gap in contact rates differs substantially across firms, exhibiting a between-company standard deviation of 1.9 percentage points. Despite an insignificant average gap in contact rates between male and female applicants, we find a between-company standard deviation in gender contact gaps of 2.7 percentage points, revealing that some firms favor male applicants and others favor women. Company-specific racial contact gaps are temporally and spatially persistent, and negatively correlated with firm profitability, federal contractor status, and a measure of recruiting centralization. Discrimination exhibits little geographical dispersion, but two-digit industry explains roughly half of the cross-firm variation in both racial and gender contact gaps. Contact gaps are highly concentrated in particular companies, with firms in the top quintile of racial discrimination responsible for nearly half of lost contacts to Black applicants in the experiment. Controlling false discovery rates to the 5% level, 23 companies are found to discriminate against Black applicants. Our findings establish that discrimination against distinctively Black names is concentrated among a select set of large employers, many of which can be identified with high confidence using large-scale inference methods.
We develop an empirical Bayes ranking procedure that assigns ordinal grades to noisy measurements, balancing the information content of the assigned grades against the expected frequency of ranking errors. Applying the method to a massive correspondence experiment, we grade the race and gender contact gaps of 97 US employers, the identities of which we disclose for the first time. The grades are presented alongside measures of uncertainty about each firm’s contact gap in an accessible report card that is easily adaptable to other settings where ranks and levels are of simultaneous interest.
We use admissions lotteries to estimate the effects of large-scale public preschool in Boston on college-going, college preparation, standardized test scores, and behavioral outcomes. Preschool enrollment boosts college attendance as well as SAT test taking and high school graduation. Preschool also decreases high school disciplinary measures including juvenile incarceration, but has no detectable effect on state achievement test scores. An analysis of subgroups shows that effects on college enrollment, SAT-taking, and disciplinary outcomes are larger for boys than for girls. Our findings illustrate possibilities for large-scale modern, public preschool and highlight the importance of measuring long-term and non–test score outcomes in evaluating the effectiveness of education programs.
The Review of Economics and Statistics2024106(1), 1-19open access
We introduce two empirical strategies harnessing the randomness in school assignment mechanisms to measure school value-added. The first estimator controls for the probability of school assignment, treating take-up as ignorable. We test this assumption using randomness in assignments. The second approach uses assignments as instrumental variables (IVs) for low-dimensional models of value-added and forms empirical Bayes posteriors from these IV estimates. Both strategies solve the underidentification challenge arising from school undersubscription. Models controlling for assignment risk and lagged achievement in Denver and New York City yield reliable value-added estimates. Estimates from models with lower-quality achievement controls are improved by IV.
American Economic Review2016106(5), 388-392open access
We develop over-identification tests that use admissions lotteries to assess the predictive value of regression-based value-added models (VAMs). These tests have degrees of freedom equal to the number of quasi-experiments available to estimate school effects. By contrast, previously implemented VAM validation strategies look at a single restriction only, sometimes said to measure forecast bias. Tests of forecast bias may be misleading when the test statistic is constructed from many lotteries or quasi-experiments, some of which have weak first stage effects on school attendance. The theory developed here is applied to data from the Charlotte-Mecklenberg School district analyzed by Deming (2014).
We use a large-scale randomized experiment to study the impact of augmenting staffing in the world’s largest public early-childhood program: India’s Integrated Child Development Services. Adding a worker doubled net preschool instructional time and led to increases of 0.28σ and 0.46σ in math and language test scores after 18 months for children who remained enrolled in the program and 0.13σ and 0.10σ for all children enrolled at baseline. Rates of stunting and severe malnutrition were also lower in the treatment group for children who remained enrolled. A cost-benefit analysis suggests that the benefits of augmenting staffing significantly exceed its costs. These effects are likely to replicate even at larger scales of program implementation.