To make high-quality research more accessible and easier to explore.

Fields:

Interpretable machine learning for creditor recovery rates

Journal of Banking & Finance 2024 164, 107187
Machine learning methods have achieved great success in modeling complex patterns in finance such as asset pricing and credit risk that enable them to outperform statistical models. In addition to the predictive accuracy of machine learning methods, the ability to interpret what a model has learned is crucial in the finance industry. We address this challenge by adapting interpretable machine learning to the context of corporate bond recovery rate modeling. In addition to the best performance, we show the value of interpretable machine learning by finding drivers of recovery rates and their relationship that cannot be discovered by the use of traditional machine learning methods. Our findings are financially meaningful and consistent with the findings in the existing credit risk literature

Machine-Learning-enhanced systemic risk measure: A Two-Step supervised learning approach

Journal of Banking & Finance 2022 136, 106416
This paper explores ways to improve the existing systemic risk measures by incorporating machine learning algorithms into the measurement. We aim to overcome the shortcomings of existing methods that rely on restricted modeling and are difficult to tap into various data resources. To this end, this paper unifies a dynamic quantification framework for systemic risk and links it to a two-step supervised learning problem, which allows for hierarchical structure of the systemic event and the return dependence. We leverage the generalization and predictive powers of machine learning to statistically model the tail events and the co-movements of the equity returns during the shocks to the macro-economy. Our results show that most machine learning algorithms enhance the systemic risk measure’s predictive power. Numerous comparative and sensitivity backtesting studies for United States and Hong Kong markets are conducted, from which we recommend the best machine learning algorithm for systemic risk measurement

Consumer credit-risk models via machine-learning algorithms

Journal of Banking & Finance 2010 34(11), 2767-2787 open access
We apply machine-learning techniques to construct nonlinear nonparametric forecasting models of consumer credit risk. By combining customer transactions and credit bureau data from January 2005 to April 2009 for a sample of a major commercial bank’s customers, we are able to construct out-of-sample forecasts that significantly improve the classification rates of credit-card-holder delinquencies and defaults, with linear regression R2’s of forecasted/realized delinquencies of 85%. Using conservative assumptions for the costs and benefits of cutting credit lines based on machine-learning forecasts, we estimate the cost savings to range from 6% to 25% of total losses. Moreover, the time-series patterns of estimated delinquency rates from this model over the course of the recent financial crisis suggest that aggregated consumer credit-risk analytics may have important applications in forecasting systemic risk

The dynamics of non-performing loans during banking crises: A new database with post-COVID-19 implications

Journal of Banking & Finance 2021 133, 106140
We present a new dataset on the dynamics of non-performing loans (NPLs) during 92 banking crises since 1990. The data show similarities across crises in NPL buildup but much heterogeneity in the pace of NPL resolution. We document how high and unresolved NPLs deepen post-crisis recessions and use a machine learning approach to establish pre-crisis predictors of NPL problems. These predictors—a set of weak macroeconomic, institutional, corporate, and banking sector conditions—help shed light on post-COVID-19 NPL vulnerabilities

Predicting individual corporate bond returns

Journal of Banking & Finance 2025 171, 107372
Using machine learning and many predictors, we find strong bond return predictability, with an out-of-sample R-squared of 4.48% and an annualized Sharpe ratio of 3.27. ML models identify important predictors for aggregate predictors (bond market returns, TERM and HML factors, GDP growth) and bond characteristics (downside risk, short-term reversal, return skewness, and credit spreads). Predictability varies over time, being stronger during periods of high investor risk aversion, slow economic growth, and strong cross-sectional factor explanatory power. Our results highlight the benefits of leveraging both cross-sectional and time-series predictors to forecast corporate bond returns while considering public and private bonds

Machine learning in corporate bonds: Evidence from China

Journal of Banking & Finance 2026 184, 107636 open access
This study employs a broad set of machine learning (ML) methods to examine cross-sectional variation in corporate bond returns in China. Using macroeconomic indicators together with bond- and issuer-specific characteristics, we find that ML techniques outperform traditional linear models in both statistical and economic terms. These models are particularly effective at capturing distinctive features of the Chinese market, including the dominance of state-owned enterprises, implicit government guarantees, and rapid market evolution. We compare long-short and long-only portfolio strategies to account for practical constraints on short selling. The results indicate that ML methods are effective in markets where institutional features and information asymmetries play a central role in asset pricing

Predicting IPO first-day returns: Evidence from machine learning analyses*

Journal of Banking & Finance 2025 178, 107500 open access
Predicting IPO first-day returns is inherently challenging due to the wide range of contributing factors, each with distinct statistical properties. We assess the performance of several machine learning (ML) techniques and identify XGBoost as the most statistically effective model for forecasting first-day returns. Using a comprehensive set of 863 pre-IPO variables, our high-performing predictive model accurately estimates both the direction and magnitude of IPO first-day returns. The most influential predictors include underwriter agency measures, price revision, and the free-float fraction. Using a rolling-window predictive approach, the model demonstrates substantial practical value, generating approximately $300 billion in gains from IPOs with positive first-day returns and avoiding more than $22 billion in losses from those with negative returns over the 2000–2016 period

Subjectivity in sovereign credit ratings

Journal of Banking & Finance 2018 88, 366-392 open access
A sovereign creditrating is a function of hard and soft information that should reflect the creditworthiness and the probability of default of a country. We propose an alternative characterisation for the subjective component of a sovereign credit rating – the parts related to the ratee’s lobbying effort or its familiarity from a United States point of view – and apply it to S&P, Moody’s and Fitch ratings, using both traditional ordered-logit panel models and machine learning techniques. This subjective component turns out to be large, especially for the low-rated countries. Countries that are rated as investment grade tend to be positively influenced by it, and vice versa. Subjective judgment in credit ratings does have predictive value: it helps in identifying chances of sovereign defaults in the short-term. Still, the impact of subjectivity in sovereign ratings on borrowing costs is very limited on average

Citations and the readers’ information-extracting costs of finance articles

Journal of Banking & Finance 2021 131, 106188 open access
This paper focuses on the relationship between the reader's information-extracting costs of finance articles and the article's number of citations. The reader's information-extracting costs are measured using three metrics: (i) the Flesch-Kincaid readability score, (ii) the article's length, and (iii) the number of complex words. Based on a sample of more than 14,000 full text articles published between 2000 and 2016 in 16 finance journals, we show that the information-extracting costs of finance journals have significantly increased over time, while the topics of these articles, determined by machine-learning topic modeling, remained relatively constant. We find a positive correlation between the reader's information-extracting costs and the number of citations achieved by a paper for articles that are published in the top-three finance journals (JF, JFE, RFS), but do not observe this pattern for articles published in other major finance journals

Risk and risk management in the credit card industry

Journal of Banking & Finance 2016 72, 218-239 open access
Using account level credit-card data from six major commercial banks from January 2009 to December 2013, we apply machine-learning techniques to combined consumer-tradeline, credit-bureau, and macroeconomic variables to predict delinquency. In addition to providing accurate measures of loss probabilities and credit risk, our models can also be used to analyze and compare risk management practices and the drivers of delinquency across the banks. We find substantial heterogeneity in risk factors, sensitivities, and predictability of delinquency across banks, implying that no single model applies to all six institutions. We measure the efficacy of a bank's risk-management process by the percentage of delinquent accounts that a bank manages effectively, and find that efficacy also varies widely across institutions. These results suggest the need for a more customized approached to the supervision and regulation of financial institutions, in which capital ratios, loss reserves, and other parameters are specified individually for each institution according to its credit-risk model exposures and forecasts