To make high-quality research more accessible and easier to explore.

Fields:

How Can Innovation Screening Be Improved? A Machine Learning Analysis with Economic Consequences for Firm Performance

Journal of Financial and Quantitative Analysis 2025 60(8), 3722-3752 open access
This study utilizes U.S. Patent Office data to explore potential improvements in the patent examination process through machine learning. It shows that integrating machine learning with human expertise can increase patent citations by up to 26%. Using machine learning predictions as benchmarks, I find that the early expiration rate of granted patents positively correlates with examiners’ false acceptance rates. These errors negatively impact public companies’ operational performance and reduce successful IPO or M&A exits for private firms. Overall, this study highlights significant social and economic benefits of incorporating machine learning as a robo-advisor in patent screening

Estimating Stock Market Betas via Machine Learning

Journal of Financial and Quantitative Analysis 2025 60(3), 1074-1110 open access
Machine learning-based stock market beta estimators outperform established benchmark models both statistically and economically. Analyzing the predictability of time-varying market betas of U.S. stocks, we document that machine learning-based estimators produce the lowest forecast and hedging errors. They also help to create better market-neutral anomaly strategies and minimum variance portfolios. Among the various techniques, random forests perform the best overall. Model complexity is highly time-varying. Historical stock market betas, turnover, and size are the most important predictors. Compared to linear regressions, allowing for nonlinearity and interactions significantly improves predictive performance

Machine Learning and the Stock Market

Journal of Financial and Quantitative Analysis 2023 58(4), 1431-1472
Practitioners allocate substantial resources to technical analysis whereas academic theories of market efficiency rule out technical trading profitability. We study this long-standing puzzle by applying a diverse set of machine learning algorithms. The results show that an investor can find profitable technical trading rules using past prices, and that this out-of-sample profitability decreases through time, showing that markets have become more efficient over time. In addition, we find that the evolutionary genetic algorithm’s attitude in not shying away from erroneous predictions gives it an edge in building profitable strategies compared to the strict loss-minimization-focused machine learning algorithms

Double Machine Learning: Explaining the Post-Earnings Announcement Drift

Journal of Financial and Quantitative Analysis 2024 59(3), 1003-1030
We demonstrate the benefits of merging traditional hypothesis-driven research with new methods from machine learning that enable high-dimensional inference. Because the literature on post-earnings announcement drift (PEAD) is characterized by a “zoo” of explanations, limited academic consensus on model design, and reliance on massive data, it will serve as a leading example to demonstrate the challenges of high-dimensional analysis. We identify a small set of variables associated with momentum, liquidity, and limited arbitrage that explain PEAD directly and consistently, and the framework can be applied broadly in finance

A Trend Factor for the Cross Section of Cryptocurrency Returns

Journal of Financial and Quantitative Analysis 2025 60(7), 3116-3153 open access
We propose CTREND, a new trend factor for cryptocurrency returns, which aggregates price and volume information across different time horizons. Using data on more than 3,000 coins, we employ machine learning methods to exploit information from various technical indicators. The resulting signal reliably predicts cryptocurrency returns. The effect cannot be subsumed by known factors and remains robust across different subperiods, market states, and alternative research designs. Moreover, it survives the impact of transaction costs and persists in big and liquid coins. Finally, an asset pricing model that incorporates CTREND outperforms competing factor models, providing a superior explanation of cryptocurrency returns

Measuring Firm Complexity

Journal of Financial and Quantitative Analysis 2024 59(6), 2487-2514 open access
In business research, firm size is both ubiquitous and readily measured. Complexity, another firm-related construct, is also relevant, but difficult to measure and not well-defined. As a result, complexity is less frequently incorporated in empirical designs. We argue that most extant measures of complexity are one-dimensional, have limited availability, and/or are frequently misspecified. Using both machine learning and an application-specific lexicon, we develop a text solution that uses widely available data and provides an omnibus measure of complexity. Our proposed measure, used in tandem with 10-K file size, provides a useful proxy that dominates traditional measures

Information in Financial Contracts: Evidence from Securitization Agreements

Journal of Financial and Quantitative Analysis 2024 59(4), 1692-1725 open access
We introduce a novel application of machine learning to compare pooling and servicing agreements (PSAs) that govern commercial mortgage-backed securities. In contrast to the view that the PSA is largely boilerplate text, we document substantial variation across PSAs, both within- and across-underwriters and over time. A part of this variation is driven by differences in loan collateral across deals. Additionally, we find that differences in PSAs are correlated with ex post loan and bond performance. Collectively, our analysis suggests the importance of examining the entire governing document, rather than specific components, when analyzing complex financial securities

Uncovering Sparsity and Heterogeneity in Firm-Level Return Predictability Using Machine Learning

Journal of Financial and Quantitative Analysis 2023 58(8), 3384-3419
We develop an approach that combines the estimation of monthly firm-level expected returns with an assignment of firms to (possibly) latent groups, both based on observable characteristics, using machine learning principles with linear models. The best-performing methods are flexible two-stage sparse models that capture group-membership predictive relationships. Portfolios formed to exploit such group-varying predictions based on a parsimonious set of characteristics deliver economically meaningful returns with low turnover. We propose statistical tests based on nonparametric bootstrapping for our results, and detail how different characteristics may matter for different groups of firms, making comparisons to the existing literature

Venture Capital Communities

Journal of Financial and Quantitative Analysis 2020 55(2), 621-651
Although venture capitalists (VCs) can choose from thousands of potential syndicate partners, many co-syndicate with small groups of preferred partners. We term these groups “VC communities.” We apply computational methods from the physical sciences to 3 decades of syndication data to identify these communities. We find that communities comprise VCs that are similar in age, connectedness, and functional style but undifferentiated in spatial location. Machine-learning tools classify communities into 3 groups roughly ordered by their age and reach. Community VC financing is associated with faster maturation and greater innovation, especially for early-stage firms without an innovation history

Why Naive $ 1/N $ Diversification Is Not So Naive, and How to Beat It?

Journal of Financial and Quantitative Analysis 2024 59(8), 3601-3632
We show theoretically that the usual estimated investment strategies will not achieve the optimal Sharpe ratio when the dimensionality is high relative to sample size, and the $ 1/N $ rule is optimal in a 1-factor model with diversifiable risks as dimensionality increases, which explains why it is difficult to beat the $ 1/N $ rule in practice. We also explore conditions under which it can be beaten, and find that we can outperform it by combining it with the estimated rules when $ N $ is small, and by combining it with anomalies or machine learning portfolios, conditional on the profitability of the latter, when $ N $ is large