Skip to content
GetProfitable
Search
Dictionary

Bonferroni correction

Divide your significance threshold by the number of tests you ran. Crude, conservative, and better than pretending you only ran one.

If you test 40 parameter sets and want an overall 5% chance of any false positive, judge each at 0.05/40 = 0.00125. Very few backtests survive this, which is informative rather than unfair.

The correction assumes independent tests, so it is too harsh when your 40 variants are near-copies of each other, such as moving average lengths of 48 to 52. In that case the effective number of independent tests is much smaller, and methods like reality-check or false-discovery-rate control are less punishing while still honest.

Whatever method you pick, the prerequisite is counting. Researchers who cannot say how many variants they tried cannot apply any correction, and their reported p-values have no interpretation.

Related: multiple-testing, false-discovery-rate, p-value, reality-check

Educational only, not advice. Spotted an error? Post in Site Feedback.