What did 2008 teach risk managers about their own models?
By 2007 the machinery was in place everywhere. Every major bank ran daily VaR. Regulators built capital rules on it. Boards reviewed it. Risk was, in an important sense, considered a solved measurement problem.
Then the measurement failed — and failed in ways that were structural rather than accidental. The post-mortem reshaped the discipline, and its lessons are the reason this course exists in its present form.
Failure 1 — the models were calibrated on the good years
Most VaR models in 2007 used a lookback window of one to four years. Those years were the Great Moderation: low volatility, rising asset prices, benign credit. A model fed calm data reports calm risk.
When conditions broke, banks recorded far more VaR exceptions — days where the actual loss exceeded the model's threshold — than the models allowed. A 99% model should be breached about 2.5 times in 250 trading days. Several institutions recorded exceptions in the double digits during the crisis period. The models weren't slightly off. They were describing a market that no longer existed.
Failure 2 — the feedback loop nobody modelled
This is the deepest lesson, and it is about reflexivity: the measurement changes the thing being measured.
Trace the loop. Volatility is low, so VaR is low. Low VaR means positions consume little risk budget, so leverage rises across the system. High leverage means that when prices fall, margin calls force selling. Forced selling raises volatility. Rising volatility raises VaR. Higher VaR breaches limits, which forces more selling.
The risk model that permitted the leverage on the way up demanded the deleveraging on the way down — and everyone's models demanded it simultaneously, because everyone's models were similar. Risk measurement became a channel of contagion. This is often called the volatility paradox: markets are most dangerous precisely when the measured risk is lowest.
Failure 3 — correlations were assumed stable
Portfolios were diversified according to correlations estimated in normal conditions. In the crisis those correlations moved sharply toward one. Assets that had spent years behaving differently fell together. Diversification, as measured, evaporated exactly when it was needed. Unit 3 gives this its own lesson.
Failure 4 — liquidity was not in the model
VaR assumes you can exit at the observed market price. In late 2008, for many instruments, there was no market price because there was no market. A model that quietly assumes a buyer exists is not measuring the risk that no buyer exists. Unit 4 takes this apart.
What changed afterwards
The response was not to abandon models but to surround them:
- Stressed VaR (Basel 2.5, 2009) — run the VaR model on a historical window of significant financial stress, in addition to the recent window, and hold capital against both. A direct fix for calibrating on calm data.
- Expected Shortfall replaced 99% VaR in the Fundamental Review of the Trading Book (2016, revised 2019), to capture tail severity.
- Mandatory stress testing became a supervisory centrepiece — annual, scenario-driven, publicly reported in several jurisdictions.
- Liquidity requirements were introduced as separate standards rather than being folded into market-risk models.
The lesson under the lessons
Every one of those fixes is an admission of the same thing: a risk model is a map drawn from past terrain. It is genuinely useful for navigating terrain that resembles the survey. It is silent about terrain that has changed, and it is dangerously quiet about the possibility that the act of navigating is itself reshaping the ground.
The professional posture that came out of 2008 is not "use better models." It is: use models, know their assumptions by name, and always ask a second question the model cannot answer.
Try it now
- The chart below will show you any window you ask it for, which is exactly the danger. Navigate to 2005–2007, measure a dozen sessions, and estimate a 95% threshold from that calm stretch alone — the way a model fitted in 2007 would have.
- Now move to 2008 and count how many sessions breached the threshold you just built. Compare that count with the roughly 5% the model implied. A model is not wrong because it was badly built; this one was built correctly, on the wrong three years.
- Measure a typical session in 2006 and a typical session in late 2008. The ratio between them is the size of the regime shift a fixed-window model missed. Then write one sentence naming which of the four failures above your own numbers make most visible. Observation only — nothing here suggests what any portfolio should hold.
Unit done. Next: what to do when you accept the model cannot see the future.