Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance
Read the paperopens www.ams.org in a new tab
What they found
The authors show with simple math how easy it is to produce an impressive backtest by overfitting. If you try enough strategy variations on random data, the expected maximum Sharpe ratio grows with the number of trials; with only a few years of data and a modest number of trials, you will routinely find a 'strategy' with a Sharpe ratio above 1 that has zero true edge. They give a formula for the minimum backtest length needed to avoid this given the number of trials, and argue that presenting a backtest without disclosing the number of trials is scientifically meaningless. Overfit strategies do not just fail out of sample; they tend to lose money because they were fit to noise that reverses.
What you can use
- A great backtest is easy to produce and proves little unless you know how many variations were tried to find it.
- With five years of data, you only need to try about 45 random strategies to find one with an in-sample Sharpe of 1 by luck alone.
- Strategies overfit to noise often lose money out of sample, not just fail to make it, because the noise mean-reverts.
- Record and report your number of trials; it is the most important number in any backtest.
Caveats
Written as a polemic for a mathematics audience; the formulas assume independent trials, which understates the problem when variations are correlated and overstates it when they are near-duplicates.
Tags: backtesting, overfitting, sharpe-ratio, multiple-testing
Summaries are our own reading of the paper, not the authors' words. Educational only, not advice. Discuss it in Book Club.