VALYNOXSTRATEGY RESEARCH TECHNOLOGY
AlphaPrimeAlphaPrime
LEARN / ROBUSTNESS

Overfitting in strategy generation: how not to fool yourself

Generate enough strategies and some will look excellent by chance alone. The more you search, the more that matters. Here is how to tell the difference.

AlphaPrime team ·

Why searching creates false winners

Imagine 10,000 strategies with no real edge at all. Their backtest results still vary, because markets are noisy. A few of them will, by chance, show smooth equity curves and strong statistics. If you sort by performance and keep the top, you will mostly keep the luckiest ones.

This is data-mining bias, also called selection bias or the multiple-testing problem. It is not specific to automated generation: a trader who manually tries fifty ideas on the same chart faces the same problem, only at a smaller scale. Automated generation makes it larger, because the search is larger.

A bell-shaped spread of backtest results from strategies with no edge; the best few in the right tail look excellent by chanceNo edge: average resultWorse backtestBetter backtestThe best few look excellentby luck aloneBacktest results of many strategies with no real edge
Illustration: strategies with no edge still spread out because markets are noisy. Keep only the top of a large search and you mostly keep the luckiest.

Signs a strategy is fitted to noise

  • Performance collapses out of sample. The clearest sign. Strong in the data used to build it, ordinary or negative on data it never saw.
  • Sharp parameter peaks. The strategy works at a lookback of 17 but not at 15 or 19. Real effects usually degrade gradually as parameters move.
  • Too many conditions for too few trades. Each extra filter can remove a few losing trades from the past. With enough filters, any history can be made to look good.
  • Results depend on a few trades. Remove the best five trades and the edge disappears.
  • Works on one market and one bar size only, with no logic explaining why.

Checks that help

1. Keep data the search never touches

Split the history before you start. Build and select on one part, then test the finalists once on a held-back part. If you go back and change the strategy after seeing the held-back results, that data is no longer unseen. Walk-forward analysis extends this idea by testing a re-tuning process on many consecutive unseen windows.

2. Raise the bar with the size of the search

The more candidates you test, the better the best one will look by luck. A result that would be impressive from 10 attempts may be expected from 100,000. Keep a record of how many strategies were generated and tested, and demand more from the finalists when the search was large.

3. Test stability, not just the best point

Run the finalist with nearby parameter values, slightly different start dates, and small amounts of noise added to prices. A robust strategy degrades gradually. A fitted one falls apart.

4. Monte Carlo the trade sequence

Reshuffle or resample the strategy's trades to see the range of drawdowns the same trades could have produced in a different order. The historical drawdown is one path. The distribution tells you what to prepare for.

5. Compare with random

Generate strategies with random entries under the same settings and see how your finalists compare with that distribution. If the best random strategies look as good as yours, the search has not found anything beyond noise.

6. Include realistic costs from the start

Costs remove many marginal strategies that look good on paper. Including them during the search, not only at the end, stops the search from favoring high-frequency noise. See trading costs in backtests.

A practical workflow

  1. Decide the data split, cost model and selection criteria before generating anything.
  2. Generate and filter on in-sample data with conservative thresholds.
  3. Run stability and Monte Carlo checks on the survivors.
  4. Run walk-forward analysis on the strategies that remain.
  5. Test the final few once on held-back data.
  6. Write down how many candidates the whole process tested.

None of these steps proves a strategy will work in the future. Together they make it much less likely that what you found is just the luckiest result of a large search.

What AlphaPrime does here

AlphaPrime's strategy validation lets you set your own thresholds and run basic, standard or extensive checks, including walk-forward analysis, stopping later checks when a strategy fails an earlier one. Faster GPU backtesting leaves more time for these checks.

Explore early access