Contents Lesson 16 of 16

4 min read · practitioner

Statistics for Market Data checkpoint — the number, with its window

Course capstone. Four units ago a return statistic was a number on a factsheet. It is now a moment with a shape, a count against a curve, a sample with a standard error, a dependence that changes with the regime, and a window that has to be stated.

The course in one architecture

  • Analyse returns, never prices — prices trend, carry corporate actions and correlate with a clock at 0.95; returns are close to stationary and comparable (Unit 1).
  • Daily returns are fat-tailed and their size clusters — 83 days beyond three sigma in twenty years against 13.6 predicted, a 22-sigma day in 1987, monthly volatility from 4% to 88% inside an "18%" average (Unit 1).
  • A mean is unknowable and a volatility is knowable — nineteen years give a 4-point standard error on a 12% mean; a t-statistic of 3 is the bar once the number of things tried is counted; a survivor-only sample flatters the mean and the left tail (Unit 2).
  • Correlation is a window's average and rises in the tail — stocks and bonds flipped from −0.46 to +0.34; two random walks correlate at −0.42 on nothing; diversifiers read 0.69 in 2017 and 0.96 in March 2020 (Unit 3).
  • The API's two columns, two calendars and one rolling window each produce a different number — and the usable one arrives with its window stated (Unit 4).

The sentence, decoded

"The fund has 18% volatility, a beta of 1.2, and a −0.3 correlation to bonds." Run it through the instruments. 18% — over which window, and was March 2020 in it? Beta 1.2 — with what R², over which years, and how much of the fund is that number? −0.3 to bonds — measured before or after 2022, and what was it on the six weeks that mattered? Every number in the sentence has a window behind it now, and none of the numbers is false. They are undated.

Where this connects

The price-data course is where the columns and calendars come from; the backtesting course is where these statistics become a strategy's report card, with the survivorship and look-ahead repairs as its test suite; the factor course's zoo is the t-statistic of 3 in its native habitat; and the portfolio-theory course's free lunch is a correlation, which this course has shown to be a window's average that rises in a crisis.

Checkpoint

The exam draws on all four units. The bar: a return series in front of you, and the ability to say what its four moments are, how many of its extremes a normal model would have predicted, what its mean's standard error is, and over what window every one of those numbers holds.

Before you sit it

Each of these is a minute at your desk. Any one that is not names the lesson to reopen first.

Try it now

  1. Write the one-sentence version of each unit from memory — four sentences, your pocket card. Do this before opening anything.
  2. Then read the card of a symbol this course never measured, /eod/GLD.US?from=2016-09-02&to=2026-09-03&fmt=json, computed on adjusted closes on 28 September 2026:
Statistic GLD.US, 2 Sep 2016 to 3 Sep 2026
Daily returns 2,513
Mean +0.052% a day
Standard deviation 1.029% a day
Skewness −0.58
Excess kurtosis 7.21
Days beyond 3σ (±3.09%) 38
One-year returns, Sep to Sep ten, mean +13.9%, standard deviation 18.7%
Correlation with SPY.US, all days 0.11
Correlation with SPY.US, days SPY moved beyond ±2% 0.19

Finish the card yourself: the normal distribution's prediction for the count beyond three sigma in 2,513 days, the standard error of the mean one-year return and its 95% band, and the annualised volatility. Then write one sentence on each unit's question, every number with its window. 3. Finally, one measurement to close the course:

Interactive line chart: GLD.US (MAX)

Measure the sharpest fall on it, and write which of this course's lessons predicts that a normal model would have called it impossible.