How likely is a −20% year?
The question has a textbook answer from a normal distribution and a better answer from the data itself, and this lesson computes both on SPY.US so the gap between them is a number rather than a warning.
Three answers, measured 2026-09-04
Inputs: the nineteen calendar-year returns of SPY.US from 2007 to 2025 — mean 12.25%, standard deviation 17.73% — and the 5,031 daily returns behind them.
The normal answer. Draw a hundred thousand years from a normal distribution with that mean and standard deviation. The share below −20%: 3.4%. The share below zero: 24.3%.
The bootstrap answer. Draw a hundred thousand years by resampling the nineteen observed years with replacement — each draw is one of the actual years, chosen at random. The share below −20%: 5.3%. Below zero: 15.8%, three in nineteen — a bootstrap of years assigns every outcome exactly its frequency in the sample, and cannot produce a year it has not seen.
The daily bootstrap. Build a hundred thousand synthetic years by drawing 252 days at random from the 5,031 observed daily returns and compounding them. The share below −20%: 4.6%.
The normal model says a −20% year comes once in thirty; the data, resampled, say once in twenty. On the tail that matters, the bell curve is optimistic by half, and it is optimistic for the reason the fat-tails lesson measured: the 2008 that produced −36.8% is the kind of year a normal distribution with a 17.7% standard deviation almost never draws.
What each method assumes
The normal draw assumes the shape. It is fast, smooth, and wrong in the tails by a factor that depends on the kurtosis it ignored. The annual bootstrap assumes nothing about the shape and everything about the sample: it can only reproduce the years that happened, and it cannot produce a 1987 or a 1931 that the nineteen years do not contain. The daily bootstrap breaks the clustering — it draws March 2020's days scattered among 2017's — and so understates the probability of a bad year even while it has the right daily tails, because the bad years are the ones where the bad days arrived together.
None of the three is the answer. The defensible figure is a range — 3% to 5% on this sample, higher if the sample were longer and contained the 1930s — and the width of the range is itself the finding.
Monte Carlo, properly named
The second and third methods are Monte Carlo simulations: repeated random draws to approximate a distribution too awkward for a formula. The method is only as good as its draws. A simulation that draws from a normal distribution has assumed away the tails; one that draws from history has assumed the future contains only the past; one that draws days independently has assumed away clustering. Every Monte Carlo result should arrive with one sentence naming what it drew from.
In the data
The annual returns are twenty December closes from /eod/SPY.US?from=2006-12-01&to=2025-12-31&period=m&fmt=json; the daily returns are the twenty-year pull. The resampling needs no endpoint, and a hundred thousand draws is a second of computation.
Try it now
- Reproduce the annual bootstrap in a spreadsheet or a few lines of code: the nineteen returns from the year-end closes below, a hundred thousand draws with replacement, the fraction below −20%. Then remove 2008 from the nineteen and rerun. Write both fractions; the second is what the method says about a sample that happens to contain no crash — and every sample that has not yet had one looks like that.
Measure calendar 2008, then calendar 2022. 3. A risk report says "probability of a 20% loss next year: 2%, by Monte Carlo". What is the first question to ask?