Why is a backtest not the future?
A backtest measures how a rule would have behaved on one specific historical sample. That is a genuinely useful measurement and a terrible promise, and the distance between those two is where most trading systems die.
Overfitting: the loud failure
Every parameter you tune is an opportunity to fit noise. Consider a rule with four parameters — a lookback length, an entry threshold, an exit threshold, a filter — tested over five plausible values each. That is 5⁴ = 625 variants.
Now report "the best one." What you have reported is the maximum of a distribution of 625 results, not a discovery. If all 625 variants were pure noise, the best of them would still show an attractive equity curve, because that's what maxima of random samples do. The curve is real; the inference isn't.
The tell is sensitivity: if the "optimal" 14-day lookback makes money and the 13- and 15-day versions don't, you haven't found a market regularity — you've found a specific arrangement of past bars.
Data snooping: the same failure, one level up
The problem repeats at the level of ideas. Test 200 different strategies against the same decade and a handful will look excellent. The decade doesn't know you were searching, and nothing in the winning strategy's statistics records the 199 that failed. This is why the most important number about a backtest is usually one nobody reports: how many things were tried before this one.
Sample size
Thirty trades cannot separate skill from luck. Individual trade outcomes vary by roughly 1R while the edge you're trying to detect might be 0.2R, and the uncertainty of an average shrinks only as 1 ÷ √n. A result computed on a handful of trades is compatible with almost any true edge, including zero and including negative.
In-sample, out-of-sample, and the honesty problem
Splitting history into a design window and an untouched holdout is the standard defence, and it works — exactly once. Look at the holdout, adjust the rule, look again, and the holdout has quietly become training data. Walk-forward testing formalises the discipline by rolling the windows forward; nothing enforces it except your own record-keeping.
Worked illustration. The same rule shows +180% over a 2010–2019 design window and +6% over a 2020–2024 holdout. Reaching for bad luck to explain that gap is the wrong instinct — but so is treating it as a clean decomposition. Overfitting, a genuinely changed regime and ordinary variance all produce exactly this picture, and one holdout cannot separate them. What the gap does settle is that the 180% was a property of the design window rather than a number to carry forward.
What a backtest can honestly do
It can falsify. "This rule lost money even in the period I designed it on" is real, hard information, and it is cheap to obtain. It can never verify — a curve that rose in the past is a necessary condition for a strategy, never a sufficient one.
So the useful question is not "did it work?" It is: "how many things did I try before this one worked, and would I have chosen it in advance?"
In the data
A backtest's inputs are not as fixed as its dates. Here is one session of Coca-Cola, traded and adjusted:
The adjusted close is recomputed backwards at every split and dividend after the date, so the same backtest re-run over the same dates a year later reads slightly different inputs for any dividend payer; the traded close never moves. A result reproduced on one series and one reproduced on the other age differently, and so does every indicator computed from them, since most published indicators use the adjusted series. Record which one you used.
Try it now
- Take any rule you find interesting and count its free parameters — lookback lengths, thresholds, filters, entry times, the instrument list itself. Multiply the number of plausible values for each. That product is the size of the search you'd actually be running.
- Before testing anything, write down the parameter values you'd choose from reasoning alone. Compare them later to the "optimal" ones. A large gap is a warning about the optimum, not about your reasoning.
- Now split the history before you study either half. The chart below is deliberately the whole thing at once: pick a cut date on it, write the holdout dates down somewhere you cannot quietly edit, and only then look at the design half. The commitment is the entire mechanism, and where you cut is a decision you make rather than one you drift into.