Empirical Asset Pricing via Machine Learning
Read the paperopens doi.org in a new tab
What they found
The authors compared machine-learning methods (regularized linear models, tree-based methods, and neural networks) for predicting monthly stock returns using 94 firm characteristics and macro variables on all U.S. stocks from 1957 to 2016, with strict out-of-sample evaluation. Neural networks and boosted trees performed best, roughly doubling the out-of-sample predictive power of the best linear methods, and a long-short portfolio built from neural-network forecasts earned a Sharpe ratio above 2 before costs. The most important predictors were price trends (momentum and reversal), liquidity, and volatility measures.
What you can use
- Machine learning genuinely improved return forecasts out of sample, mainly by capturing interactions among known signals, not by discovering new ones.
- Even the best models had an out-of-sample R-squared below 1% per month; predictability is real but tiny, and the value comes from portfolio aggregation.
- The signals machine learning found most useful were momentum, reversal, and liquidity, the same ones the older literature identified.
Caveats
Long-short portfolios heavy in small stocks, gross of costs; much of the paper profit is in illiquid names. Requires substantial data and computation. Free SSRN version exists.
Tags: backtesting, machine-learning, return-prediction, cross-sectional
Summaries are our own reading of the paper, not the authors' words. Educational only, not advice. Discuss it in Book Club.