Guides / How Much Backtest History Is Enough

GUIDE

How much backtest is enough? Count trades, not months.

"A year of backtest data" means very different things for a scalper and a swing strategy. Here's a more honest way to ask the question, and the warning signs that your backtest is telling you less than you think.

Cephic TeamUpdated Sep 29, 20267 min read

Quick answer: there's no single magic number, but under roughly 50–100 closed trades, statistical noise dominates over whatever real edge a strategy has — you genuinely can't tell a modest edge apart from a lucky run. The honest answer scales with your strategy's trade frequency, not a round number of months.

WRONG QUESTION

Why "X months of backtest" is the wrong question

A scalper making 15–20 trades a day has more statistically meaningful data after one month than a swing strategy making two trades a month has after a full year. Calendar time tells you almost nothing on its own — trade count is what actually determines how much you can trust the result.

THE SAMPLE-SIZE PROBLEM

Why small samples lie confidently

Statistical uncertainty around a win rate shrinks slowly as trades accumulate — cutting the noise in half roughly requires quadrupling the sample. With only 20–30 trades, a strategy that's actually a coin flip can easily produce a win rate that looks like a real edge, and a strategy with a genuine modest edge can look broken. This is exactly why a tool that reports a number should also say when it doesn't have enough data to mean anything, instead of returning a precise-looking figure regardless of sample size.

BEYOND TRADE COUNT

What matters besides raw trade count

COMMON MISTAKES

Where traders overtrust a backtest

WHAT TO DO WITH THE NUMBER

What to actually do about it

Check trade count before calendar span. Deliberately look for at least one rough stretch in the backtest window, not just a smooth one. Treat the backtest's max drawdown as a floor for what live can produce, not a ceiling. And re-grade the strategy once enough live trades accumulate, rather than trusting the original backtest indefinitely — live performance is the only dataset that can't be curve-fit after the fact.

Allocate grades your backtest's sample size and quality directly, instead of leaving you to eyeball it.

Open Allocate