Contents Lesson 8 of 16

5 min read · professional

How do you check whether a risk model is working?

A risk model makes a testable prediction. A 99% one-day VaR claims that losses will exceed the threshold on about 1% of days. That claim can be checked against what actually happened — and checking it is called backtesting.

This is the discipline that separates a risk number from a decoration.

Counting exceptions

An exception (or breach) is a day where the realised loss exceeded that day's VaR forecast. The test is arithmetic:

  • Over 250 trading days (about one year), a 99% VaR model predicts 2.5 exceptions.
  • A 95% model over the same period predicts 12.5.

Now compare to what you observed. Zero exceptions in a year is not a triumph — it suggests the model is too conservative and the firm is over-reserving. Twelve exceptions on a 99% model is not bad luck; it is evidence the model is wrong.

The traffic light

Supervisors formalised this into a traffic-light test for a 99% model over 250 days:

Exceptions Zone Reading
0–4 Green Consistent with a working model
5–9 Yellow Questionable — investigate, higher capital multiplier
10 or more Red The model is presumed deficient

The zones exist because 250 days is a small sample and randomness alone can produce four exceptions from a perfectly calibrated model. The bands separate "unlucky" from "broken" — with statistical honesty about the grey area between them.

Not just how many — when

Two models can both record five exceptions and be in completely different health.

Model 1: five exceptions scattered across the year, no two adjacent. Model 2: five exceptions inside a single fortnight.

Model 1 looks like a correctly calibrated model meeting normal randomness. Model 2 is clustering, and clustering means the model failed to react to a change in regime — it kept reporting calm-market risk into a storm. Formal tests of independence (alongside tests of the exception count) exist for exactly this reason: a good model should breach unpredictably, not in bunches.

A worked check

A desk runs a 99% one-day VaR. Over the past 250 days it recorded 7 exceptions, six of which fell in one three-week period.

Reading: the count alone lands in the yellow zone — more breaches than expected. The clustering is the more informative signal: it points to a model whose volatility estimate updates too slowly, so it stayed low while realised volatility jumped. That's a diagnosis of the model, not of the market. It is also a description of what happened, not a prediction of what the desk should do next — that decision belongs to the people accountable for the capital.

Backtesting Expected Shortfall is harder

VaR backtesting is clean because the prediction is binary: breached or not. Expected Shortfall predicts the average size of breaches, so testing it requires comparing magnitudes across a handful of tail days — a much noisier exercise. This is the acknowledged cost of the regulatory move to ES: a better measure that is harder to validate. Practitioners typically backtest VaR at several confidence levels as a proxy for the ES model's health.

The habit

Any risk number you cannot backtest is an opinion. The first question to ask about any model — yours or a vendor's — is not "what does it say?" but "how often was it wrong last year, and were the wrongs bunched together?"

In the data

The realised returns the exceptions are counted against usually come from adjusted closes, and an adjusted close is not a fixed number. The S&P 500 fund's last session of 2025:

Live API response: pm3 spy 2025 12 31 close

The traded close will never change. The adjusted figure for the same day already sits below it, because every dividend paid since has been folded back into the history, and it will move again after the next one. A backtest rerun a year later works from slightly different past returns, so it can land on a different exception count with nothing having changed in the model or in the market. Record the date the data was taken along with the result.

Try it now

  1. The full history is below — enough for a multi-year backtest by eye. Estimate a 99% one-day threshold from a single calm year, the way a fixed-window model would, and write the number down.
Interactive candles chart: SPY.US (MAX)
  1. Now walk forward. Year by year, Measure the sessions that breached your threshold and count them. Compare your count per year to the 2.5 a 99% model expects over 250 days, and place each year in the traffic-light table.
  2. Mark where your exceptions fell on the calendar. Scattered or clustered? Write one sentence on what the pattern says about the model's responsiveness — and note that the rolling version of this test, which re-estimates the threshold every day from the previous 250, is the version a risk desk actually runs. It is a spreadsheet exercise rather than a chart one, and it exists precisely to fix what you just watched go wrong.