What did the dataset forget, and what did it know too early?
A statistic is only as honest as the sample it was computed on, and market samples have two characteristic dishonesties: they forget the companies that failed, and they include information that had not yet been published on the date it is used. Both inflate every mean and shrink every tail. This lesson is what each one does to a number, with the endpoints that repair them.
Survivorship: the mean of the ones that lived
Build a list of today's S&P 500 constituents and compute their average return over twenty years. The list contains the companies that grew into the index or stayed in it, and not the ones that were removed — Lehman, Enron, the hundreds of quieter delistings. The average is the average of survivors, and it is higher than any investor could have earned, because no investor in 2006 could have held only the companies that would still exist in 2026.
The size of the effect is not small. The derivatives course's constituents lesson and the backtesting course both measured it; Shumway's 1997 finding, that omitting delisting returns overstated the small-cap premium materially, is the factor course's version. In this course's terms: a survivor-only sample has a higher mean and a thinner left tail than the market that produced it, because the failures it dropped were the large negative returns. What it does to the spread depends on the sample — a survivor universe can be a volatile one — so the two effects to count on are the mean and the tail.
The repair is a point-in-time universe: the list of what existed on each date, with the returns of the ones that then disappeared, including the delisting return — often near total loss. /exchange-symbol-list/US?delisted=1 is the list of what is gone; the price-data course's delistings lesson is how to use it; /symbol-change-history is the trail of renames that otherwise splits one company's history into two.
Look-ahead: the number that was not yet public
A dataset of quarterly earnings is keyed by the quarter it covers. The accounts for a quarter ending in December are published in February. A statistic that uses the December figure on a December date has used a number nobody had, and any relationship it finds between that number and the next month's return is partly the market learning the number.
The same leak lives in prices. The adjustment loop rewrites adjusted_close backwards every time a dividend or split occurs, so today's adjusted series encodes events that had not happened on the dates it describes. For return statistics that is what you want — a return should include the dividend — but for a signal computed on a date, the price known on that date was the raw close, and the reproducibility lesson is about which of the two a query returns tomorrow.
The repair is a filing date: use each number from the day it became public, which /fundamentals reports beside each statement, and rank on prices as of the close before the trade, which the leaking bar enforces.
Why statistics and not just backtests
Both biases are usually taught as backtest errors. They are sample errors, and they distort a mean, a correlation or a volatility computed for any purpose. A correlation matrix built on today's constituents is the correlation matrix of a portfolio that was never available; a volatility or a correlation computed on survivors describes a universe nobody could have held, and can sit on either side of the true figure. The honest sample is the one that includes the dead.
In the data
/exchange-symbol-list/US?delisted=1&fmt=json returns the tickers that no longer trade, with their last dates; /symbol-change-history?from=…&to=… the renames; /fundamentals/{ticker} carries a filing date beside each statement. A point-in-time universe is those three joined by date, and it is the only universe on which the statistics in this course are unbiased.
Try it now
- Both lists are too long for a page, so here are their row counts, taken on 28 September 2026:
| Call | Rows |
|---|---|
/exchange-symbol-list/US?fmt=json (live) |
51,003 |
/exchange-symbol-list/US?delisted=1&fmt=json |
60,303 |
Write the ratio of delisted to all rows; that is the fraction of the listed US universe a survivor-only sample never sees. The delisted list is not all companies: on the same day 33,112 of its rows were Common Stock and 19,018 FUND, with ETFs, preferreds, units and warrants making up the rest. Say which types your own statistics would need to keep, and what filtering both lists to them would do to the ratio.
2. The index whose current members are the usual survivor sample:
- A line each: what survivorship does to a mean and to the left tail, and why its effect on the standard deviation is not fixed.