Contents Lesson 14 of 16

4 min read · professional

Which biases quietly inflate a backtest?

Overfitting is the loud failure and it gets all the attention. These are the quiet ones. Each makes a backtest better than reality without anyone lying, which is precisely what makes them dangerous — the researcher believes the number too.

Look-ahead bias

Using information that wasn't knowable at the decision timestamp. It hides in ordinary-looking code:

  • Deciding on today's close and assuming a fill at that same close.
  • Using restated financials at their restated values, months before the restatement existed.
  • Screening on an index membership list as of today across a ten-year history.

One test catches nearly all of it: "at that timestamp, on that clock, could this have been known?"

Survivorship bias

Run the test on today's index members and you have excluded everyone who was delisted, acquired in distress, or went bankrupt. Foundations covered this for return statistics; in a trading backtest it is worse, because failing names are exactly the ones a momentum or mean-reversion rule would have traded most enthusiastically. The universe must be reconstructed as it was, delisted names included.

Wrong price series

A backtest on unadjusted prices sees a phantom crash on every split date and a phantom gap on every ex-dividend date. Some rules "brilliantly" trade those non-events. Conversely, a test that assumes fully adjusted prices for an intraday strategy is modelling prices nobody could have transacted at.

Ignored costs

The most common and most quantifiable inflation. A system averaging +0.15R per trade across 500 trades a year looks like +75R annually:

Modelled cost per trade Net edge per trade Annual (500 trades)
0.00R (gross) +0.15R +75R
0.05R +0.10R +50R
0.08R +0.07R +35R
0.15R 0.00R 0R

Realistic costs of 0.08R per trade remove more than half the edge. Double the cost per trade — 0.15R — and it's gone entirely, which is the table's last row. Turnover doesn't erase the edge; it scales whatever survives the cost line, in whichever direction that points: 1,000 trades at +0.07R net is +70R, and 1,000 trades at a cost above the gross edge bleeds twice as fast as 500. That is why high-frequency edges are the most cost-sensitive — a small gross edge per trade, measured against a cost charged in full every time — and the most frequently reported gross.

Capacity and liquidity

Fills assumed at the mid, in full size, on an instrument trading 20,000 shares a day while the model buys 50,000. On paper the strategy works. In the order book, the strategy is the price move it was trying to capture.

Regime dependence

A rule tuned in one volatility regime, one interest-rate environment or one dominant trend is a precise description of that regime. Nothing in the backtest tells you the regime continues — and nothing in the equity curve distinguishes "this rule works" from "this decade was like that."

The checklist

Before believing a backtest number: point-in-time data, delisted names included, correct price adjustment, costs modelled explicitly, size checked against typical volume, and a written statement of which regime the sample actually covers. Six lines. They eliminate more bad strategies than any amount of cleverness.

In the data

Two items on that checklist are choices a data source makes for you unless you say otherwise. Most ticker lists show only what trades today, so a universe built from one has already dropped every name the survivorship section is about; the delisted names have to be asked for. And the price series is invisible once you have a result. Here is the same fund twice over its whole history:

Interactive candles chart: SPY.US (MAX)
Interactive line chart: SPY.US (MAX)

The candles are the prices that traded; the line is closes adjusted for every dividend since, and at the left edge the two sit far apart. A backtest number does not say which one it ran on, so write it down beside the result rather than hoping to recover it later.

Try it now

  1. Using the cost table above, find the per-trade cost at which a +0.20R gross edge over 300 trades a year would be halved, and the cost at which it disappears.
  2. Check the capacity line yourself. Read a typical session's participation off the volume pane below, or take the 50-session average from the table under the chart, and compute the largest position you could take without exceeding, say, 1% of it. Compare that with the position size Unit 1's formula gave you.
Interactive volume chart: IWM.US (1Y)
Live API response: mf2 iwm avgvol50 latest
  1. Say the look-ahead test out loud against one rule you like: "at that timestamp, could this have been known?" Most rules survive it. The ones that don't were never strategies.