Contents Lesson 3 of 16

4 min read · practitioner

How often does a five-sigma day happen, really?

A five-sigma event, in a normal world, happens about once in 3.5 million days — roughly once in fourteen thousand years of trading. SPY.US has had nineteen of them since 2006. Here is the count, the name for the property, and what it does to every risk number built on the bell curve.

The count, measured 2026-09-04

Daily returns on adjusted closes, 1 September 2006 to 3 September 2026, standard deviation 1.224%:

Threshold Normal distribution predicts Observed
beyond 2σ (about ±2.4%) 228.9 days 237
beyond 3σ (about ±3.7%) 13.6 days 83
beyond 4σ (about ±4.9%) 0.3 days 37
beyond 5σ (about ±6.1%) 0.0 days 19

At two sigma the bell curve is nearly right. At three it is wrong by six times, at four by a hundred, at five by more than the window can express. That is what fat tails means: not that markets are more volatile than a normal distribution — the standard deviation is the same by construction — but that the volatility is distributed differently, with more of it in a few extreme days and less in the ordinary ones.

Over the ten years to 3 September 2026 the pattern repeats at smaller scale: 32 days beyond three sigma against 6.8 predicted, 17 beyond four against 0.16, 10 beyond five against none. Bitcoin, on the same test from 2016: 61 days beyond three sigma against 9.9 predicted, with a daily standard deviation of 3.49% — fatter tails on a wider distribution.

1987, the twenty-two sigma day

On 19 October 1987 the S&P 500 fell 20.5% in a session. Measured against the daily standard deviation of the five years before it — 0.904%, from GSPC.INDX — that is a 22.6-sigma move. A normal distribution assigns a 22-sigma event a probability so small that the age of the universe, counted in days, does not contain one. It happened, and it happened on a Monday, and the next lesson is why it did not happen alone.

What fat tails break

Three things, each in wide use. Value at risk at 99%, computed from a normal distribution, says the worst one-in-a-hundred day is about 2.3 standard deviations; the data say the hundredth-worst day in 5,031 was well beyond that, and the days past it were where the money was lost. Option prices from a normal model underprice the far-out-of-the-money puts that pay on the 83 days — which is why, since 1987, the market has priced them higher than the model does, the smile the derivatives domain covers. A leveraged fund's decay, from the ETF course's daily reset lesson, scales with the variance, and fat tails put more of the variance into fewer days.

What to do about it

Not much can be modelled and one thing can be measured: count the tails on the data itself rather than reading them off a curve. A risk figure that says "worst day in a hundred" should be the hundredth-worst day observed, and a sample that contains no crisis has no tails to count — which is a statement about the sample, not the market.

In the data

The 1987 figure is two calls: /eod/GSPC.INDX?from=1987-10-19&to=1987-10-19&fmt=json for the day and /eod/GSPC.INDX?from=1982-10-01&to=1987-10-16&fmt=json for the standard deviation of what came before. The counts above are the twenty-year SPY.US pull from the previous lesson, one comparison per day against the thresholds.

Try it now

  1. The twenty-year series is 5,031 daily returns, too many for a page, so here is the count done on it: /eod/SPY.US?from=2006-09-01&to=2026-09-03&fmt=json, returns on adjusted closes, days beyond three standard deviations (±3.67%), grouped by calendar year. Computed on 28 September 2026:
Year Days beyond 3σ
2007 1
2008 29
2009 12
2010 2
2011 8
2015 2
2018 3
2020 17
2022 5
2025 4, all in April
All twenty years 83

Write down what fraction fall in 2008, 2020 and April 2025, and how many calendar years of the window, 2006 to 2026, have none at all. The tail is a calendar, not a lottery. 2. The index the 1987 figure comes from, at full range, where the day is a notch and 2008 is a valley:

Interactive line chart: GSPC.INDX (MAX)

Measure 16 October to 19 October 1987. Then measure the drawdown from 9 October 2007 to 9 March 2009 and write which of the two a normal model would have found less surprising, and why the answer is neither.