‹ Backtest Strategies Lesson 14 of 17
Contents Lesson 14 of 17

3 min read · professional

In-sample, out-of-sample, and the honest way to tune

Unit 2 warned that tuning parameters until the curve looks good is fitting the answer to the sample. This lesson is the discipline that makes tuning legitimate.

Split before you tune, not after

Divide the history by date — the first 70% is in-sample, the rest is out-of-sample — and then do not look at the second part.

Tune in-sample as much as you like. Try fifty parameter pairs; that is what the in-sample period is for. Then run the winner once on the out-of-sample period, and whatever it produces is your result.

The discipline is entirely in the word once. Look at the out-of-sample result, go back, adjust, and look again, and you have converted it into a second in-sample period — just a smaller and noisier one. Nothing in the code stops you; only the log from unit 2 records that you did.

What you expect to see

Out-of-sample performance is worse. Always, more or less. A rule that returns 18% in-sample and 6% out is normal and possibly useful. A rule that returns 18% and 17% is either a rare genuine effect or a leak you have not found yet — and after unit 3, the prior should be leak.

A rule that returns 18% in-sample and −4% out has told you something valuable: the 18% was noise, and you learned that for the price of one run.

Walking forward

One split uses one out-of-sample period, which is one sample. Walk-forward gives you several: tune on years 1–3, test on year 4; tune on 2–4, test on 5; and so on. String the test periods together and that concatenated curve is a much more honest picture, because every point in it was produced by parameters chosen without seeing it.

It is also the first thing in this course that needs real history. Which is the next lesson.

Two cheap checks that catch most overfitting

The neighbour test, from unit 2's closing exercise: vary each parameter slightly and plot the results. You want a plateau, not a spike. A rule that works at 20/50 and fails at 21/50 has found a coincidence, and no amount of out-of-sample testing rescues that.

Count your attempts. If you tried forty parameter sets, the best of forty is expected to look good on noise alone. Report the number of attempts next to the result. It costs nothing, it is in your run log already, and it is the single most informative line in an honest report.

The finance behind it

How much evidence a run of good results actually is: How many years before a track record means anything?

Try it now

Split your data 70/30 before touching anything, tune only on the first part, and write your prediction for the out-of-sample return down before you run it. Comparing your prediction to the outcome calibrates your judgement, which is worth more than the strategy.