‹ Measuring Performance Lesson 15 of 16
Contents Lesson 15 of 16

4 min read · professional

How many years before a track record means anything?

Everything measured so far has been an estimate drawn from a sample. Estimates carry error, and performance data carries an unusually large amount of it. This lesson puts a number on how much history is required before a record can be distinguished from chance — and the number is uncomfortable.

The core relationship

For an annualised information ratio measured over T years, the t-statistic of the excess return is approximately:

t ≈ IR × √T

The conventional bar for calling a result statistically distinguishable from zero is roughly t = 2. Rearranged, that gives the years required:

T ≈ (2 ÷ IR)²

Information ratio Years needed for t ≈ 2
1.00 4
0.50 16
0.40 25
0.33 37
0.25 64

Sustained information ratios of 1.0 are rare. Records in the 0.2–0.4 range are far more typical — which means the evidence requirement runs to decades, frequently longer than the manager's career, the fund's life, or the strategy's relevance. A five-year record is, statistically, close to silent.

The coin-flip arithmetic

Suppose 1,000 managers each have exactly a 50% chance of beating their benchmark each year, and no skill whatsoever. After five years:

1,000 × (1/2)⁵ = 1,000 ÷ 32 ≈ 31 managers with five consecutive winning years

Thirty-one flawless five-year records, produced by a room containing zero skill. They will have explanations. Some will be articulate. None of it will be evidence, because the process generating the records had none to give.

Run it with 500 managers over seven years: 500 ÷ 128 ≈ 4 with a perfect record. There is almost always someone.

Three ways the sample gets worse

Survivorship. Funds that close disappear from the tables. The average of what remains is therefore computed on the survivors, which lifts it. Compare a "peer group average" today with the full list of funds that existed at the start of the period and the two differ systematically.

Multiple testing. Screen 5,000 funds and rank them, and the top ten are selected on the same quantity you want to measure. With enough candidates, extreme results are guaranteed by arithmetic. Selecting the best out of thousands is not the same as testing one in advance.

Backtest overfitting. A strategy tuned on the history it is evaluated on will fit that history. The in-sample record is a description of the tuning, not evidence about anything else. Only out-of-sample, forward-dated performance carries information — and it starts a fresh, short clock.

What this does and doesn't say

It does not say skill is absent. It says the evidentiary bar is far higher than intuition suggests, that most published records sit well below it, and that confident conclusions drawn from a few years of data are usually conclusions about noise.

Held properly, this is liberating rather than cynical: it tells you exactly which questions the data can answer, and stops you spending attention on the ones it cannot. And it is measurement, not guidance — nothing here identifies a manager or an investment worth choosing.

Try it now

  1. Compute the years required for t ≈ 2 at an information ratio of 0.40, then at 0.60. Note how sharply the requirement falls as the ratio rises.
  2. Compute how many of 2,000 zero-skill managers would post six consecutive winning years by chance (2,000 ÷ 2⁶).
  3. A fund and a defensible benchmark are below. Measure each of the last five calendar years separately on both charts and subtract, giving five annual active returns. Count how many were positive — and then ask what that count, on its own, clears. Against step 2, the answer is nothing at all.
Interactive line chart: QQQ.US (5Y)
Interactive line chart: SPY.US (5Y)