To make high-quality research more accessible and easier to explore.

Fields:
28 results ✕ Clear filters

Interpretable machine learning for creditor recovery rates

Journal of Banking & Finance 2024 164, 107187
Machine learning methods have achieved great success in modeling complex patterns in finance such as asset pricing and credit risk that enable them to outperform statistical models. In addition to the predictive accuracy of machine learning methods, the ability to interpret what a model has learned is crucial in the finance industry. We address this challenge by adapting interpretable machine learning to the context of corporate bond recovery rate modeling. In addition to the best performance, we show the value of interpretable machine learning by finding drivers of recovery rates and their relationship that cannot be discovered by the use of traditional machine learning methods. Our findings are financially meaningful and consistent with the findings in the existing credit risk literature

How do machine learning and non-traditional data affect credit scoring? New evidence from a Chinese fintech firm

Journal of Financial Stability 2024 73, 101284 open access
This paper compares the predictive power of credit scoring models based on machine learning techniques with that of traditional loss and default models. Using proprietary transaction-level data from a leading fintech company in China, we test the performance of different models to predict losses and defaults both in normal times and when the economy is subject to a shock. In particular, we analyse the case of an (exogenous) change in regulation policy on shadow banking in China that caused credit conditions to deteriorate. We find that the model based on machine learning and non-traditional data is better able to predict losses and defaults than traditional models in the presence of a negative shock to the aggregate credit supply. This result reflects a higher capacity of non-traditional data to capture relevant borrower characteristics and of machine learning techniques to better mine the non-linear relationship between variables in a period of stress

Charting by machines

Journal of Financial Economics 2024 153, 103791
We test the efficient market hypothesis by using machine learning to forecast stock returns from historical performance. These forecasts strongly predict the cross-section of future stock returns. The predictive power holds in most subperiods and is strong among the largest 500 stocks. The forecasting function has important nonlinearities and interactions, is remarkably stable through time, and captures effects distinct from momentum, reversal, and extant technical signals. These findings question the efficient market hypothesis and indicate that technical analysis and charting have merit. We also demonstrate that machine learning models that perform well in optimization continue to perform well out-of-sample

Missing values handling for machine learning portfolios

Journal of Financial Economics 2024 155, 103815
We characterize the structure and origins of missingness for 159 cross-sectional return predictors and study missing value handling for portfolios constructed using machine learning. Simply imputing with cross-sectional means performs well compared to rigorous expectation-maximization methods. This stems from three facts about predictor data: (1) missingness occurs in large blocks organized by time, (2) cross-sectional correlations are small, and (3) missingness tends to occur in blocks organized by the underlying data source. As a result, observed data provide little information about missing data. Sophisticated imputations introduce estimation noise that can lead to underperformance if machine learning is not carefully applied

Discretionary dissemination on Twitter

Contemporary Accounting Research 2024 41(4), 2454-2487 open access
The study provides large‐scale descriptive evidence on the timing and nature of corporate financial tweeting. Using an unsupervised machine learning approach to analyze 24 million tweets posted by S&P 1500 firms from 2012 to 2020, we find that firms are more likely to tweet financial information around significantly negative or positive news events, such as earnings announcements and the filing of financial statements. This convex U‐shaped relation between the likelihood of posting financial tweets and the materiality of accounting events becomes stronger over time. Whereas research based on early samples concludes that firms are less likely to disseminate financial information on Twitter when the news is bad and material, the symmetric dissemination behavior we find suggests that these conclusions should be revised. We also show that a machine learning algorithm (Twitter‐Latent Dirichlet Allocation) is superior to a dictionary approach in classifying short messages like tweets

Pre-publication revisions of bank financial statements: A novel way to monitor banks?

Journal of Financial Intermediation 2024 58, 101073
We investigate whether pre-publication revisions of bank financial statements contain forward-looking information about bank risk. Using 7.4 million observations of monthly financial reports from all banks in Brazil during 2007–2019, we show that 78 % of all revisions occur before the publication of these statements. The frequency, missing of reporting deadlines, and severity of revisions are positively related to future bank risk. Using machine learning techniques, we provide evidence on mechanisms through which revisions affect bank risk. Our findings suggest that private information about pre-publication revisions is useful for supervisors to monitor banks

Data and the Aggregate Economy

Journal of Economic Literature 2024 62(2), 458-484
Recent data technology innovations, such as artificial intelligence and machine learning, have transformed the production of knowledge and increased the importance of data. This review explores how data—digitized information—has been modeled within classic macroeconomic frameworks. It compares the economics of data to other concepts such as ideas, patents, and learning-by-doing. This paper also shows potential ways to model applications for data, including innovation, process optimization, and matching. Because this research area is nascent, much of the article is devoted to open questions and directions for future data economy research

Do Anomalies Really Predict Market Returns? New Data and New Evidence

Review of Finance 2024 28(1), 1-44 open access
Using new data from US and global markets, we revisit market risk premium predictability by equity anomalies. We apply a repertoire of machine-learning methods to forty-two countries to reach a simple conclusion: anomalies, as such, cannot predict aggregate market returns. Any ostensible evidence from the USA lacks external validity in two ways: it cannot be extended internationally and does not hold for alternative anomaly sets—regardless of the selection and design of factor strategies. The predictability—if any—originates from a handful of specific anomalies and depends heavily on seemingly minor methodological choices. Overall, our results challenge the view that anomalies as a group contain helpful information for forecasting market risk premia

Double Machine Learning: Explaining the Post-Earnings Announcement Drift

Journal of Financial and Quantitative Analysis 2024 59(3), 1003-1030
We demonstrate the benefits of merging traditional hypothesis-driven research with new methods from machine learning that enable high-dimensional inference. Because the literature on post-earnings announcement drift (PEAD) is characterized by a “zoo” of explanations, limited academic consensus on model design, and reliance on massive data, it will serve as a leading example to demonstrate the challenges of high-dimensional analysis. We identify a small set of variables associated with momentum, liquidity, and limited arbitrage that explain PEAD directly and consistently, and the framework can be applied broadly in finance

Measuring Firm Complexity

Journal of Financial and Quantitative Analysis 2024 59(6), 2487-2514 open access
In business research, firm size is both ubiquitous and readily measured. Complexity, another firm-related construct, is also relevant, but difficult to measure and not well-defined. As a result, complexity is less frequently incorporated in empirical designs. We argue that most extant measures of complexity are one-dimensional, have limited availability, and/or are frequently misspecified. Using both machine learning and an application-specific lexicon, we develop a text solution that uses widely available data and provides an omnibus measure of complexity. Our proposed measure, used in tandem with 10-K file size, provides a useful proxy that dominates traditional measures