To make high-quality research more accessible and easier to explore.

Fields:
5 results ✕ Clear filters

Human Decisions and Machine Predictions

Quarterly Journal of Economics 2018 133(1), 237-293 open access
Can machine learning improve human decision making? Bail decisions provide a good test case. Millions of times each year, judges make jail-or-release decisions that hinge on a prediction of what a defendant would do if released. The concreteness of the prediction task combined with the volume of data available makes this a promising machine-learning application. Yet comparing the algorithm to judges proves complicated. First, the available data are generated by prior judge decisions. We only observe crime outcomes for released defendants, not for those judges detained. This makes it hard to evaluate counterfactual decision rules based on algorithmic predictions. Second, judges may have a broader set of preferences than the variable the algorithm predicts; for instance, judges may care specifically about violent crimes or about racial inequities. We deal with these problems using different econometric strategies, such as quasi-random assignment of cases to judges. Even accounting for these concerns, our results suggest potentially large welfare gains: one policy simulation shows crime reductions up to 24.7% with no change in jailing rates, or jailing rate reductions up to 41.9% with no increase in crime rates. Moreover, all categories of crime, including violent crimes, show reductions; these gains can be achieved while simultaneously reducing racial disparities. These results suggest that while machine learning can be valuable, realizing this value requires integrating these tools into an economic framework: being clear about the link between predictions and decisions; specifying the scope of payoff functions; and constructing unbiased decision counterfactuals

Diagnosing Physician Error: A Machine Learning Approach to Low-Value Health Care

Quarterly Journal of Economics 2022 137(2), 679-727 open access
How effective are physicians at diagnosing heart attacks? To answer this question, we contrast physician testing decisions with a machine learning model of risk. When the two deviate, we use actual health outcome data to judge whether the algorithm or the physician was right. We find physicians over-test: tests that are predictably useless are still performed. At the same time, physicians also under-test: many predicted high-risk patients are untested and then suffer adverse health events (including death) at high rates. A natural experiment using shift-to-shift testing variation confirms these findings: increasing testing improves health and reduces mortality, but only for patients flagged as high-risk by the algorithm. The simultaneous existence of over- and under-testing cannot easily be explained by incentives alone, and instead suggests errors. We provide suggestive evidence on the psychology behind these errors:(i) physicians use too simple a model of risk, suggesting bounded rationality; (ii) they over-weight salient information; and (iii) they over-weight symptoms that are representative or stereotypical of heart attack. Together, these results suggest the need for health care models and policies to incorporate not just physician incentives, but also physician mistakes

Folklore

Quarterly Journal of Economics 2021 136(4), 1993-2046 open access
Folklore is the collection of traditional beliefs, customs, and stories of a community passed through the generations by word of mouth. We introduce to economics a unique catalog of oral traditions spanning approximately 1,000 societies. After validating the catalog's content by showing that the groups' motifs reflect known geographic and social attributes, we present two sets of applications. First, we illustrate how to fill in the gaps and expand upon a group's ethnographic record, focusing on political complexity, high gods, and trade. Second, we discuss how machine learning and human classification methods can help shed light on cultural traits, using gender roles, attitudes toward risk, and trust as examples. Societies with tales portraying men as dominant and women as submissive tend to relegate their women to subordinate positions in their communities, both historically and today. More risk-averse and less entrepreneurial people grew up listening to stories wherein competitions and challenges are more likely to be harmful than beneficial. Communities with low tolerance toward antisocial behavior, captured by the prevalence of tricksters being punished, are more trusting and prosperous today. These patterns hold across groups, countries, and second-generation immigrants. Overall, the results highlight the significance of folklore in cultural economics, calling for additional applications

Machine Learning as a Tool for Hypothesis Generation

Quarterly Journal of Economics 2024 139(2), 751-827
While hypothesis testing is a highly formalized activity, hypothesis generation remains largely informal. We propose a systematic procedure to generate novel hypotheses about human behavior, which uses the capacity of machine learning algorithms to notice patterns people might not. We illustrate the procedure with a concrete application: judge decisions about whom to jail. We begin with a striking fact: the defendant’s face alone matters greatly for the judge’s jailing decision. In fact, an algorithm given only the pixels in the defendant’s mug shot accounts for up to half of the predictable variation. We develop a procedure that allows human subjects to interact with this black-box algorithm to produce hypotheses about what in the face influences judge decisions. The procedure generates hypotheses that are both interpretable and novel: they are not explained by demographics (e.g., race) or existing psychology research, nor are they already known (even if tacitly) to people or experts. Though these results are specific, our procedure is general. It provides a way to produce novel, interpretable hypotheses from any high-dimensional data set (e.g., cell phones, satellites, online behavior, news headlines, corporate filings, and high-frequency time series). A central tenet of our article is that hypothesis generation is a valuable activity, and we hope this encourages future work in this largely “prescientific” stage of science

The Health Costs of Cost Sharing

Quarterly Journal of Economics 2024 139(4), 2037-2082 open access
What happens when patients suddenly stop their medications? We study the health consequences of drug interruptions caused by large, abrupt, and arbitrary changes in price. Medicare's prescription drug benefit as-if-randomly assigns 65-year-olds a drug budget as a function of their birth month, beyond which out-of-pocket costs suddenly increase. Those facing smaller budgets consume fewer drugs and die more: mortality increases 0.0164 percentage points per month (13.9%) for each $100 per month budget decrease (24.4%). This estimate is robust to a range of falsification checks and lies in the 97.8th percentile of 544 placebo estimates from similar populations that lack the same idiosyncratic budget policy. Several facts help make sense of this large effect. First, patients stop taking drugs that are both high value and suspected to cause life-threatening withdrawal syndromes when stopped. Second, using machine learning, we identify patients at the highest risk of drug-preventable adverse events. Contrary to the predictions of standard economic models, high-risk patients (e.g., those most likely to have a heart attack) cut back more than low-risk patients on exactly those drugs that would benefit them the most (e.g., statins). Finally, patients appear unaware of these risks. In a survey of 65-year-olds, only one-third believe that stopping their drugs for up to a month could have any serious consequences. We conclude that far from curbing waste, cost sharing is itself highly inefficient, resulting in missed opportunities to buy health at very low cost ($11,321 per life-year