To make high-quality research more accessible and easier to explore.

Fields:

Interpretable machine learning for creditor recovery rates

Journal of Banking & Finance 2024 164, 107187
Machine learning methods have achieved great success in modeling complex patterns in finance such as asset pricing and credit risk that enable them to outperform statistical models. In addition to the predictive accuracy of machine learning methods, the ability to interpret what a model has learned is crucial in the finance industry. We address this challenge by adapting interpretable machine learning to the context of corporate bond recovery rate modeling. In addition to the best performance, we show the value of interpretable machine learning by finding drivers of recovery rates and their relationship that cannot be discovered by the use of traditional machine learning methods. Our findings are financially meaningful and consistent with the findings in the existing credit risk literature

Charting by machines

Journal of Financial Economics 2024 153, 103791
We test the efficient market hypothesis by using machine learning to forecast stock returns from historical performance. These forecasts strongly predict the cross-section of future stock returns. The predictive power holds in most subperiods and is strong among the largest 500 stocks. The forecasting function has important nonlinearities and interactions, is remarkably stable through time, and captures effects distinct from momentum, reversal, and extant technical signals. These findings question the efficient market hypothesis and indicate that technical analysis and charting have merit. We also demonstrate that machine learning models that perform well in optimization continue to perform well out-of-sample

Missing values handling for machine learning portfolios

Journal of Financial Economics 2024 155, 103815
We characterize the structure and origins of missingness for 159 cross-sectional return predictors and study missing value handling for portfolios constructed using machine learning. Simply imputing with cross-sectional means performs well compared to rigorous expectation-maximization methods. This stems from three facts about predictor data: (1) missingness occurs in large blocks organized by time, (2) cross-sectional correlations are small, and (3) missingness tends to occur in blocks organized by the underlying data source. As a result, observed data provide little information about missing data. Sophisticated imputations introduce estimation noise that can lead to underperformance if machine learning is not carefully applied

Pre-publication revisions of bank financial statements: A novel way to monitor banks?

Journal of Financial Intermediation 2024 58, 101073
We investigate whether pre-publication revisions of bank financial statements contain forward-looking information about bank risk. Using 7.4 million observations of monthly financial reports from all banks in Brazil during 2007–2019, we show that 78 % of all revisions occur before the publication of these statements. The frequency, missing of reporting deadlines, and severity of revisions are positively related to future bank risk. Using machine learning techniques, we provide evidence on mechanisms through which revisions affect bank risk. Our findings suggest that private information about pre-publication revisions is useful for supervisors to monitor banks

Data and the Aggregate Economy

Journal of Economic Literature 2024 62(2), 458-484
Recent data technology innovations, such as artificial intelligence and machine learning, have transformed the production of knowledge and increased the importance of data. This review explores how data—digitized information—has been modeled within classic macroeconomic frameworks. It compares the economics of data to other concepts such as ideas, patents, and learning-by-doing. This paper also shows potential ways to model applications for data, including innovation, process optimization, and matching. Because this research area is nascent, much of the article is devoted to open questions and directions for future data economy research

Double Machine Learning: Explaining the Post-Earnings Announcement Drift

Journal of Financial and Quantitative Analysis 2024 59(3), 1003-1030
We demonstrate the benefits of merging traditional hypothesis-driven research with new methods from machine learning that enable high-dimensional inference. Because the literature on post-earnings announcement drift (PEAD) is characterized by a “zoo” of explanations, limited academic consensus on model design, and reliance on massive data, it will serve as a leading example to demonstrate the challenges of high-dimensional analysis. We identify a small set of variables associated with momentum, liquidity, and limited arbitrage that explain PEAD directly and consistently, and the framework can be applied broadly in finance

Cross-sectional expected returns: new Fama–MacBeth regressions in the era of machine learning

Review of Finance 2024 28(6), 1807-1831
We extend the Fama–MacBeth regression framework for cross-sectional return prediction to incorporate big data and machine learning. Our extension involves a three-step procedure for generating return forecasts based on Fama–MacBeth regressions with regularization and predictor selection as well as forecast combination and encompassing. As a by-product, it provides estimates of characteristic payoffs. We also develop three performance measures for assessing cross-sectional return forecasts, including a generalization of the popular time-series out-of-sample R2 statistic to the cross section. Applying our extension to over 200 firm characteristics, our cross-sectional return forecasts significantly improve out-of-sample predictive accuracy and provide substantial economic value to investors. Overall, our results suggest that a relatively large number of characteristics matter for determining cross-sectional expected returns. Our new method is straightforward to implement and interpret, and it performs well in our application

Why Naive $ 1/N $ Diversification Is Not So Naive, and How to Beat It?

Journal of Financial and Quantitative Analysis 2024 59(8), 3601-3632
We show theoretically that the usual estimated investment strategies will not achieve the optimal Sharpe ratio when the dimensionality is high relative to sample size, and the $ 1/N $ rule is optimal in a 1-factor model with diversifiable risks as dimensionality increases, which explains why it is difficult to beat the $ 1/N $ rule in practice. We also explore conditions under which it can be beaten, and find that we can outperform it by combining it with the estimated rules when $ N $ is small, and by combining it with anomalies or machine learning portfolios, conditional on the profitability of the latter, when $ N $ is large

Estimating Nursing Home Quality with Selection

The Review of Economics and Statistics 2024
We use variational inference (VI), a technique from the machine learning literature, to estimate a mortality-based Bayesian model of nursing home quality accounting for selection. We demonstrate how one can use VI to quickly and flexibly estimate a high-dimensional economic model with large datasets. Using our facility quality estimates, we examine the correlates of quality and find that public report cards have near-zero correlation. We then show that in contrast to prior literature, higher quality nursing homes fared better during the pandemic: a one standard deviation increase in quality corresponds to 2.5% fewer Covid-19 cases

Bank Supervision and Organizational Capital: The Case of Minority Lending

Journal of Accounting Research 2024 62(2), 505-549
We investigate whether improvements in banks' organizational capital and control systems facilitate increased loan origination to minority borrowers. We focus on bank supervisors' enforcement decisions and orders (EDOs) against banks and hypothesize that EDO‐imposed improvements in loan policies, internal governance, and employee training mitigate deficiencies in credit assessments and lending decisions that previously disadvantaged minority borrowers. We find that mortgage origination to minority borrowers increases following the resolution of EDOs, and more so for banks with stricter supervisors or more severe EDOs. Using a semisupervised machine learning method to analyze the text of EDOs, we find that such increases are higher for EDOs specifying revisions of loan policies, implementation of formal internal governance procedures, or more employee training. Overall, we find that EDO‐driven improvements in organizational capital generate unintended, positive social externalities that enhance access to credit for minority borrowers