To make high-quality research more accessible and easier to explore.

Fields:
187 results

Relative Valuation with Machine Learning

Journal of Accounting Research 2023 61(1), 329-376 open access
We use machine learning for relative valuation and peer firm selection. In out‐of‐sample tests, our machine learning models substantially outperform traditional models in valuation accuracy. This outperformance persists over time and holds across different types of firms. The valuations produced by machine learning models behave like fundamental values. Overvalued stocks decrease in price and undervalued stocks increase in price in the following month. Determinants of valuation multiples identified by machine learning models are consistent with theoretical predictions derived from a discounted cash flow approach. Profitability ratios, growth measures, and efficiency ratios are the most important value drivers throughout our sample period. We derive a novel method to express valuation multiples predicted by our machine learning models as weighted averages of peer firm multiples. These weights are a measure of peer–firm comparability and can be used for selecting peer‐groups

How Can Innovation Screening Be Improved? A Machine Learning Analysis with Economic Consequences for Firm Performance

Journal of Financial and Quantitative Analysis 2025 60(8), 3722-3752 open access
This study utilizes U.S. Patent Office data to explore potential improvements in the patent examination process through machine learning. It shows that integrating machine learning with human expertise can increase patent citations by up to 26%. Using machine learning predictions as benchmarks, I find that the early expiration rate of granted patents positively correlates with examiners’ false acceptance rates. These errors negatively impact public companies’ operational performance and reduce successful IPO or M&A exits for private firms. Overall, this study highlights significant social and economic benefits of incorporating machine learning as a robo-advisor in patent screening

Interpretable machine learning for creditor recovery rates

Journal of Banking & Finance 2024 164, 107187
Machine learning methods have achieved great success in modeling complex patterns in finance such as asset pricing and credit risk that enable them to outperform statistical models. In addition to the predictive accuracy of machine learning methods, the ability to interpret what a model has learned is crucial in the finance industry. We address this challenge by adapting interpretable machine learning to the context of corporate bond recovery rate modeling. In addition to the best performance, we show the value of interpretable machine learning by finding drivers of recovery rates and their relationship that cannot be discovered by the use of traditional machine learning methods. Our findings are financially meaningful and consistent with the findings in the existing credit risk literature

Machine learning and the prediction of changes in profitability

Contemporary Accounting Research 2023 40(4), 2643-2672 open access
This study uses machinelearning methods to predict next‐period change in profitability based on a model proposed by Penman and Zhang (2004, Working paper, Columbia University and University of California, Berkeley; “PZ”). We find that new machinelearning methods predict out of sample substantially better than traditional regression methods and provide richer interpretations about the role and impact of different predictor variables through their nonlinear relationships and interaction effects. For example, our results contrast with previous research by showing that both components of the DuPont decomposition (change in profit margin and change in asset turnover) are informative of next‐period changes in profitability. Our results are robust across different performance metrics, alternative machinelearning models, and software. Furthermore, an unconstrained machinelearning model using a larger feature space could not significantly improve the performance of the PZ model. PZ variables alone accounted for most of the explanatory power of the unconstrained model, suggesting the PZ model is both well specified (in terms of feature selection) and robust in higher dimensional settings. With respect to the economic significance of this information, we find mixed results. The market appears to adjust its expectations more in line with the machinelearning predictions relative to the PZ model but the portfolio returns are not significantly different

Technical indicators and the cross-section of corporate bond returns in a machine learning era

Journal of Financial Markets 2026 79, 101029 open access
We explore the use of technical indicators to forecast corporate bond returns with various machine learning models. We show that technical indicators yield statistically significant and economically meaningful results, consistently outperforming bond characteristics. Although bond characteristics possess predictive power for bond returns, they do not provide incremental value beyond technical indicators across all bonds. Additionally, machine learning models do not offer substantial improvements over the benchmark linear model. These results underscore the significance of technical indicators in the corporate bond market

How do machine learning and non-traditional data affect credit scoring? New evidence from a Chinese fintech firm

Journal of Financial Stability 2024 73, 101284 open access
This paper compares the predictive power of credit scoring models based on machine learning techniques with that of traditional loss and default models. Using proprietary transaction-level data from a leading fintech company in China, we test the performance of different models to predict losses and defaults both in normal times and when the economy is subject to a shock. In particular, we analyse the case of an (exogenous) change in regulation policy on shadow banking in China that caused credit conditions to deteriorate. We find that the model based on machine learning and non-traditional data is better able to predict losses and defaults than traditional models in the presence of a negative shock to the aggregate credit supply. This result reflects a higher capacity of non-traditional data to capture relevant borrower characteristics and of machine learning techniques to better mine the non-linear relationship between variables in a period of stress

Bottom up vs. top down: What does firm 10-K tell us?

Journal of Financial Markets 2026 79, 101070 open access
While financial textual analysis increasingly relies on complex machine learning, we propose a simpler, data-driven alternative. Using elastic net regressions on a massive panel of 10-K n-grams, we construct a specialized dictionary that weights phrases by their marginal predictive power. This bottom-up methodology effectively forecasts expected stock returns, with a spread portfolio generating significant average returns. Our approach outperforms prominent financial dictionaries, off-the-shelf large language models, and machine learning algorithms. These results demonstrate the value of identifying financial meaning from the bottom up, highlighting the need for domain- specific models trained on relevant financial contexts

Machine-Learning-enhanced systemic risk measure: A Two-Step supervised learning approach

Journal of Banking & Finance 2022 136, 106416
This paper explores ways to improve the existing systemic risk measures by incorporating machine learning algorithms into the measurement. We aim to overcome the shortcomings of existing methods that rely on restricted modeling and are difficult to tap into various data resources. To this end, this paper unifies a dynamic quantification framework for systemic risk and links it to a two-step supervised learning problem, which allows for hierarchical structure of the systemic event and the return dependence. We leverage the generalization and predictive powers of machine learning to statistically model the tail events and the co-movements of the equity returns during the shocks to the macro-economy. Our results show that most machine learning algorithms enhance the systemic risk measure’s predictive power. Numerous comparative and sensitivity backtesting studies for United States and Hong Kong markets are conducted, from which we recommend the best machine learning algorithm for systemic risk measurement

International corporate bond returns: Uncovering predictability using machine learning

Journal of Financial Markets 2026 79, 101008 open access
We examine the cross-sectional predictability of corporate bond returns using a novel international dataset and a set of machine learning techniques. We find strong predictability in both U.S. and non-U.S. markets, with differing predictive factors. Bonds in developed markets show greater integration with the U.S. market and stronger ties to equity markets. Predictive performance of machine learning models varies over time and is greater before the onset of the COVID-19 pandemic and during periods of deteriorating business conditions, reduced market liquidity, elevated investor sentiment, and heightened risk aversion. The results offer insights into bond pricing and global diversification opportunities

Does Machine Learning Automate Moral Hazard and Error?

American Economic Review 2017 107(5), 476-480 open access
Machine learning tools are beginning to be deployed en masse in health care. While the statistical underpinnings of these techniques have been questioned with regard to causality and stability, we highlight a different concern here, relating to measurement issues. A characteristic feature of health data, unlike other applications of machine learning, is that neither y nor x is measured perfectly. Far from a minor nuance, this can undermine the power of machine learning algorithms to drive change in the health care system--and indeed, can cause them to reproduce and even magnify existing errors in human judgment