Contents Lesson 6 of 16

4 min read · professional

What does a t-statistic of 3 mean, and why did finance raise the bar?

The previous lesson produced a mean and a standard error. Divide one by the other and you have a t-statistic: how many standard errors the mean is from zero. It is the number behind every "significant" in a strategy backtest or a factor paper, and this lesson is what it does and does not license.

The arithmetic

SPY.US's nineteen calendar years, measured on 2026-09-04: mean 12.25%, standard error 4.07 points, t-statistic 3.0. Under the conventional rule — a t-statistic above 2 is significant at roughly the 5% level — the equity premium on this window is significant. A strategy showing 6% a year over ten years at 20% volatility has a standard error of 6.3 points and a t-statistic below 1: not distinguishable from zero, however good the chart looked.

The p-value is the same number read the other way: the probability of seeing a t-statistic this large if the true mean were zero. A t of 2 is a p of about 5%; a t of 3, about 0.3%.

The bar was too low

The 5% rule means one test in twenty passes by chance. Nobody runs one test. A researcher who tries forty definitions of value, or a trader who backtests forty parameter settings, will find two that pass at 5% on pure noise, and the two are the ones that get published or traded. Harvey, Liu and Zhu made the argument formally in 2016 — the factor course's zoo lesson has the three hundred factors — and proposed that a new claim about returns should show a t-statistic of at least 3 before anyone believes it. Three is not magic; it is roughly what the 5% rule becomes once the number of things tried is counted.

What the number cannot say

Three limits. A large t on a short window is a large mean, not a proven one: the standard error already contains the window, but the window may not contain a crisis, and the tails lesson said what a sample without one is worth. Significance is not size: a t of 3 on a strategy earning 0.5% a year is a real but useless effect. The test assumes the observations are what they seem: overlapping windows, autocorrelated returns and volatility clustering all shrink the effective sample below the count of rows, and a t-statistic computed on daily returns as if each day were independent overstates its own confidence.

The whole claim

A claim about a return comes with four numbers — the mean, the standard error, the t-statistic and the number of things tried before this one — and a reader who gets three of the four has been told a story. The backtesting course's parameters lesson is the same idea for a strategy: the parameter you chose is one of many you looked at, and the t-statistic that matters is the one adjusted for the looking.

In the data

The inputs are the previous lesson's nineteen year-end closes. For a daily-return t-statistic on any symbol, /eod/{ticker}?from=…&to=…&fmt=json and one pass: mean over standard deviation times the square root of the count — and then a correction for the clustering, which is the part nobody's spreadsheet includes.

Try it now

  1. The daily version, computed on 28 September 2026 from /eod/SPY.US?from=2016-09-02&to=2026-09-03&fmt=json treating the 2,513 daily returns as independent: mean 0.063%, standard deviation 1.132%, so t = 0.063 ÷ 1.132 × √2,513 = 2.80. Here are the ten one-year returns in the same window, each from the first session on or after 2 September, on adjusted closes:
Year from Return
2016-09-02 +14.97%
2017-09-05 +19.98%
2018-09-04 +2.31%
2019-09-03 +25.49%
2020-09-02 +28.58%
2021-09-02 −12.22%
2022-09-02 +16.40%
2023-09-05 +24.61%
2024-09-03 +17.42%
2025-09-02 +22.09%

Compute the t-statistic from the ten: their mean over their standard deviation, times √10. Write both numbers. They differ, and not because either one knows more about SPY. The daily figure assumes independent days, and this decade's daily volatility annualised by √252, 18.0%, is wider than the spread of its actual years. Say which assumption produces the gap, and which way it would have gone in a decade whose bad days came in runs. 2. The window, for scale:

Interactive line chart: SPY.US (5Y)
  1. A backtest shows a t-statistic of 2.4 after trying twelve parameter settings. Roughly what is the chance that the best of twelve pure-noise tests reaches 2.0? (Hint: one minus 0.95 to the twelfth.)