To make high-quality research more accessible and easier to explore.

Fields:
29 results

Machine Learning as a Tool for Hypothesis Generation

Quarterly Journal of Economics 2024 139(2), 751-827
While hypothesis testing is a highly formalized activity, hypothesis generation remains largely informal. We propose a systematic procedure to generate novel hypotheses about human behavior, which uses the capacity of machine learning algorithms to notice patterns people might not. We illustrate the procedure with a concrete application: judge decisions about whom to jail. We begin with a striking fact: the defendant’s face alone matters greatly for the judge’s jailing decision. In fact, an algorithm given only the pixels in the defendant’s mug shot accounts for up to half of the predictable variation. We develop a procedure that allows human subjects to interact with this black-box algorithm to produce hypotheses about what in the face influences judge decisions. The procedure generates hypotheses that are both interpretable and novel: they are not explained by demographics (e.g., race) or existing psychology research, nor are they already known (even if tacitly) to people or experts. Though these results are specific, our procedure is general. It provides a way to produce novel, interpretable hypotheses from any high-dimensional data set (e.g., cell phones, satellites, online behavior, news headlines, corporate filings, and high-frequency time series). A central tenet of our article is that hypothesis generation is a valuable activity, and we hope this encourages future work in this largely “prescientific” stage of science.

Diagnosing Physician Error: A Machine Learning Approach to Low-Value Health Care

Quarterly Journal of Economics 2022 137(2), 679-727 open access
How effective are physicians at diagnosing heart attacks? To answer this question, we contrast physician testing decisions with a machine learning model of risk. When the two deviate, we use actual health outcome data to judge whether the algorithm or the physician was right. We find physicians over-test: tests that are predictably useless are still performed. At the same time, physicians also under-test: many predicted high-risk patients are untested and then suffer adverse health events (including death) at high rates. A natural experiment using shift-to-shift testing variation confirms these findings: increasing testing improves health and reduces mortality, but only for patients flagged as high-risk by the algorithm. The simultaneous existence of over- and under-testing cannot easily be explained by incentives alone, and instead suggests errors. We provide suggestive evidence on the psychology behind these errors:(i) physicians use too simple a model of risk, suggesting bounded rationality; (ii) they over-weight salient information; and (iii) they over-weight symptoms that are representative or stereotypical of heart attack. Together, these results suggest the need for health care models and policies to incorporate not just physician incentives, but also physician mistakes.

Network Effects and Welfare Cultures*

Quarterly Journal of Economics 2000 115(3), 1019-1055
We empirically examine the role of social networks in welfare participation using data on language spoken at home to better infer networks within an area. Our empirical strategy asks whether being surrounded by others who speak the same language increases welfare use more for those from high welfare-using language groups. This methodology allows us to include local area and language group fixed effects and to control for the direct effect of being surrounded by one's language group; these controls eliminate many ofthe problems in previous studies. The results strongly confirm the importance of networks in welfare participation.

Does Machine Learning Automate Moral Hazard and Error?

American Economic Review 2017 107(5), 476-480 open access
Machine learning tools are beginning to be deployed en masse in health care. While the statistical underpinnings of these techniques have been questioned with regard to causality and stability, we highlight a different concern here, relating to measurement issues. A characteristic feature of health data, unlike other applications of machine learning, is that neither y nor x is measured perfectly. Far from a minor nuance, this can undermine the power of machine learning algorithms to drive change in the health care system--and indeed, can cause them to reproduce and even magnify existing errors in human judgment.

Behavioral Hazard in Health Insurance *

Quarterly Journal of Economics 2015 130(4), 1623-1667 open access
A fundamental implication of standard moral hazard models is overuse of low-value medical care because copays are lower than costs. In these models, the demand curve alone can be used to make welfare statements, a fact relied on by much empirical work. There is ample evidence, though, that people misuse care for a different reason: mistakes, or "behavioral hazard." Much high-value care is underused even when patient costs are low, and some useless care is bought even when patients face the full cost. In the presence of behavioral hazard, welfare calculations using only the demand curve can be off by orders of magnitude or even be the wrong sign. We derive optimal copay formulas that incorporate both moral and behavioral hazard, providing a theoretical foundation for value-based insurance design and a way to interpret behavioral "nudges." Once behavioral hazard is taken into account, health insurance can do more than just provide financial protection - it can also improve health care efficiency.

Learning Through Noticing: Theory and Evidence from a Field Experiment *

Quarterly Journal of Economics 2014 129(3), 1311-1353
We consider a model of technological learning under which people “learn through noticing”: they choose which input dimensions to attend to and subsequently learn about from available data. Using this model, we show how people with a great deal of experience may persistently be off the production frontier because they fail to notice important features of the data they possess. We also develop predictions on when these learning failures are likely to occur, as well as on the types of interventions that can help people learn. We test the model’s predictions in a field experiment with seaweed farmers. The survey data reveal that these farmers do not attend to pod size, a particular input dimension. Experimental trials suggest that farmers are particularly far from optimizing this dimension. Furthermore, consistent with the model, we find that simply having access to the experimental data does not induce learning. Instead, behavioral changes occur only after the farmers are presented with summaries that highlight previously unattended-to relationships in the data.

The Psychological Lives of the Poor

American Economic Review 2016 106(5), 435-440 open access
All individuals rely on a fundamental set of mental capacities and functions, or bandwidth, in their economic and non-economic lives. Yet, many factors associated with poverty, such as malnutrition, alcohol consumption, or sleep deprivation, may tax this capacity. Previous research has demonstrated that such taxes often significantly alter judgments, preferences, and decision-making. A more suggestive but growing body of evidence points toward potential effects on productivity and utility. Considering the lives of the poor through the lens of bandwidth may improve our understanding of potential causes and consequences of poverty.

Do Financial Concerns Make Workers Less Productive?

Quarterly Journal of Economics 2025 140(1), 635-689
Workers who are worried about their personal finances may find it hard to focus at work. If so, reducing financial concerns could increase productivity. We test this hypothesis in a sample of low-income Indian piece-rate manufacturing workers. We stagger when wages are paid out: some workers are paid earlier and receive a cash infusion while others remain liquidity constrained. The cash infusion leads workers to reduce their financial concerns by immediately paying off debts and buying household essentials. Subsequently, they become more productive at work: their output increases by 7% (0.11 std. dev.), and they make fewer costly, unintentional mistakes. Workers with more cash on hand thus not only work faster but also more attentively, suggesting improved cognition. These effects are concentrated among more financially constrained workers. We argue that mechanisms such as gift exchange or nutrition cannot account for our results. Instead, our findings suggest that financial strain, at least partly through psychological channels, has the potential to reduce earnings exactly when money is most needed.

Helping Consumers Know Themselves

American Economic Review 2011 101(3), 417-422
Firms sometimes know more about a consumer's expected usage than the consumer herself. We explore the consequences of this reversal in the information asymmetry. We analyze the consequences of making consumers more informed about themselves. While making consumers more informed decreases their expenditure conditional on a given set of prices, equilibrium prices may increase, offsetting the direct benefit of information. We discuss theoretical and practical issues surrounding so-called RECAP regulation that would require firms to provide each consumer with information about her own usage of the firm's product.