Bottom up vs. top down: What does firm 10-K tell us?
While financial textual analysis increasingly relies on complex machine learning, we propose a simpler, data-driven alternative. Using elastic net regressions on a massive panel of 10-K n-grams, we construct a specialized dictionary that weights phrases by their marginal predictive power. This bottom-up methodology effectively forecasts expected stock returns, with a spread portfolio generating significant average returns. Our approach outperforms prominent financial dictionaries, off-the-shelf large language models, and machine learning algorithms. These results demonstrate the value of identifying financial meaning from the bottom up, highlighting the need for domain- specific models trained on relevant financial contexts.