The Vault

RISK IS NOT VARIANCE | Episode 4 – Risk Is Not Variance

Shuffle forty years of returns like a deck of cards. Every statistic stays identical. The worst loss moves 50 points.

The Evidence Base:
Over 640,000 daily observations across 68 global futures markets, spanning more than four decades, from September 1984 to July 2026.
One consistent method throughout.

Shuffle the deck

Take every daily return the S&P 500 has produced over forty years. Ten thousand cards, each printed with a number.

Now shuffle them.

Don’t add a card. Don’t remove one. Don’t change a single number on any of them. Just deal the same deck in a different order.

Here is what happens to the statistics.

The average return does not move. Neither does the volatility. Nor the skew, nor the fatness of the tails, nor the Sharpe ratio, nor a single correlation with a single other market. Every familiar number stays exactly where it was.

None of them contains any information about order. They are computed by adding things up, and addition doesn’t care what sequence you add in.

Now here is what happens to the investor.

We dealt the deck three thousand times. The worst loss ranged from -33.6% to -83.8%.

Fifty percentage points of capital, separating histories that every risk model on earth regards as identical. One of those orderings leaves an investor bruised. Another closes the fund.

That is the whole argument. The rest of this episode is detail.

Every number on a risk report survives a shuffle. Investors do not.

Variance measures how widely returns scatter around their average. Risk is the possibility that you are not there at the end. Different questions. Only the second has a floor.

Ruin is a property of a sequence, and the sequence is the one thing a risk model cannot see.

The history that actually happened lost 62.8%, toward the ugly end of the range. It could have been kinder. It could have been a great deal worse.

How long does ruin take?

Ask a risk manager to picture a tail event and you will get a crash. One black morning. A number so far from the mean that the model calls it impossible.

We measured it instead. For each of the 68 markets we located the worst peak-to-trough loss in its history and asked a simple question. How long did it take?

The median answer is 1,949 trading days.

Seven and three quarter years. The middle half of the sample runs from 859 days to 4,074 days, and only 9% of these episodes resolved inside a single year.

A crash is over in a fortnight. At the median, the damage had another seven years to run.

Read that again, because it undoes the picture almost everybody carries. The event that did the most damage to the average futures market was not a crash. It was a grind. It arrived slowly enough that there was time to rationalise it at every stage, and it took most of a decade to finish.

Nothing in that requires a fat tail. Nothing in it requires a single extraordinary day. A market can destroy you at an entirely ordinary pace, and most of them did.

This is why the shuffle matters so much. A crash is a large number, and a large number shows up in the variance. A seven-year decline is a thousand unremarkable numbers arranged in an unlucky order, and there is no moment in the standard toolkit at which that arrangement becomes visible.

There is a link here to our earlier work. In Fractals of Finance we established persistence across all 68 markets using Hurst exponents. A persistent series does not oscillate around a mean. It travels. What we measured there predicts exactly what we have measured here. Long directional grinds are what persistent markets produce, and long directional grinds are what empty accounts.

That is not a coincidence of two projects. The Hurst exponent is a statistic that a shuffle destroys, which is exactly why it can see what the risk report cannot.

Fragility accumulates

A dam does not fail on the day it fails. Water finds a hairline crack. Frost widens it over a winter. The structure carries its load for another decade and shows nothing on the surface. The collapse takes four seconds. The weakening took thirty years.

A forest fire is the same shape. Undergrowth builds season by season until combustible material forms an unbroken path from one side of the valley to the other. Nothing visible happens during the accumulation. Then a spark arrives, and the spark is the least interesting part of the story.

Which tells us what kind of object risk is. It is not a quantity present in the market on a given afternoon, waiting to be measured. It is a chain of contingent steps, each one conditional on the step before it. Remove one link and the chain breaks. Assemble them all and the outcome becomes hard to escape.

A snapshot cannot see this. It measures one link and reports the strength of one link.

The models do multiply probabilities, and that is not the error. The error is which probabilities they multiply. So we went back to the 68 markets and measured the difference.

Count the days when a market moves more than three standard deviations. Across the sample they arrive about once every seventy-five trading days. Rare, and that is what the model is built to expect.

Now count them again, but only in the week following one that has already happened. The rate is one day in eleven.

Seven times more likely, because the first one happened. The effect appears in all 68 markets, without a single exception.

Then shuffle the deck again. Same returns, same number of extreme days, dealt in a different order. The lift disappears completely. In the shuffled histories an extreme day is exactly as likely after another extreme day as at any other time.

So the clustering was never in the numbers. It was in their arrangement.

An independence assumption computes the chance of three bad things happening as the product of three unconditional numbers, and three small numbers multiplied together produce something vanishingly small. That is where the impossible day comes from. Reality multiplies the second event given the first, and the third given both, and those numbers are not small. They are seven times larger, because the earlier event created the conditions for the later one.

Note the symmetry with what compounding does to capital. Wealth is multiplicative through time, which is why the average return is not the return you receive. Fragility is multiplicative through time as well, which is why the measured risk is not the risk you carry. Both errors come from the same place. Both come from treating a sequence as though it were a sum.

An honest correction

Earlier in this research we expected to show that volatility tells you nothing about how badly you can be hurt. That was wrong, and we are withdrawing it without qualification.

Across the 68 markets, a market’s standard deviation explains about three quarters of the variation in its worst drawdown. Silver draws down harder than short-term interest rates, and its standard deviation says so.

We tried hard to break that result and could not. Cut the drawdown window out of the sample, so the crisis cannot help predict itself, and the relationship barely moves. Measure volatility over a three-year window that closes before the drawdown even begins, which is the only way an investor could ever use it, and it holds at the same level. Restrict the measurement to each market’s calmest third of days and it still holds.

Standard deviation is a real guide to the scale of damage a market is capable of. We expected that claim to fall apart under testing and it did not.

What it cannot do is tell you where inside the range you will land. The shuffle test holds volatility fixed at a single value and the outcome still swings by fifty points. It cannot tell you when. And it cannot tell you what will happen to you, as opposed to what will happen to the market, which turns out to be a different question altogether.

Episode Three showed that the worst losses to a leveraged book arrive out of quiet markets, because volatility-based position sizing builds the position during the calm. That result stands, and we have tested it again here.

Ask the same question of the market rather than the book and the answer inverts. Drawdowns begin, on median, at the 65th percentile of their own market’s volatility history. Worst single days begin higher still. Two thirds of the time, the market was already louder than usual before the damage started.

Both are true, and together they say something neither says alone. The market’s worst day and your worst day are not the same day. The shock comes out of turbulence. The position it lands on was built in the quiet.

Volatility gives you the neighbourhood. It cannot give you the address, and the neighbourhood is fifty points wide.

What the squaring costs you

The squaring treats a 30% gain and a 30% loss as identical contributions to risk. In the arithmetic of survival they are nothing of the sort.

Losses compound against you at an accelerating rate, and at the end of the road sits a door that only opens one way. There is no matching door on the upside. Nobody was ever forced out of the market for making too much money.

A measure that weights both ends the same way is describing a person who can keep betting no matter what happens. That person does not exist.

There is a second bill, and this one arrives as arithmetic rather than philosophy.

Modern Portfolio Theory maximises expected return for a given level of volatility. Expected return here means the simple average of the yearly returns.

You do not get the simple average. You get the compounded return, which is what is left in the account. And the gap between them is created by volatility itself.

Look at silver. The model is shown an asset returning 5.4% a year. The investor who actually held it kept 0.8%.

Volatility ate 4.6 of those 5.4 percentage points and never sent an invoice.

Across the 68 markets the relationship is close to mechanical. The more a market moves, the more this costs you.

So the optimiser is minimising a quantity that is quietly consuming the quantity it is maximising. It sees the gross figure and never the net.

The average of many lives, and the one you get

This is the deepest version, and the profession took longest to notice it. Ole Peters has done the rigorous work, and what follows is his argument rather than ours.

Imagine a bet with a positive expected value. A thousand people take it once each. On average the group does well.

Now take the same bet, and have one person repeat it a thousand times, staking their whole pot each round.

These are not the same experiment. They do not have the same answer. The second person can hit zero and stop. The group simply averages over whoever happened to survive.

Volatility, expected return, correlation, the Sharpe ratio, and every statistic in the standard toolkit describe the first experiment. They tell you what happens on average across many parallel lives.

Nobody lives many parallel lives. You get one, in one order, and if it touches zero then the fact that things went well in the other simulations is of no help at all.

You do not need a thought experiment for this. One has been running in public for a century.

An index level is an average taken across the companies present at that moment. It is a snapshot of a population, and it tells you how that population is doing on aggregate today.

The companies inside it are not living that average. They are living their own paths, one at a time, in order, with a floor beneath each of them.

Picture it drawn out.

Reading the figure. Each faint line is one investor’s capital, running from the day they entered to the day they stopped. Grey lines are still holding. Red lines end at a cross on the ruin floor. Entries are staggered, because in life nobody starts on the same morning. The drift and volatility used here are not exotic. They describe a single volatile commodity, or an ordinary market held with leverage, and the ensemble average rises throughout.

The vertical black line is a moment of measurement. It is where a risk report gets written, an index level gets quoted, or a study calculates the return to holding this market.

Sixty-two investors exist at that line. Forty-nine have already gone, thirty of them wiped out. Twenty-nine have not yet arrived. Every statistic computed at that line is computed from the sixty-two.

The right-hand panel shows the same moment as a distribution. Grey is what the calculation can see. Red is everyone who was there and is now gone, plotted at the value they finished on. The red mass sits at the bottom, piled against the ruin floor, entirely outside the blue band of ordinary conditions.

Forty-three of the hundred and eleven people who entered this market are below the lower edge of that band. Thirty-two of them are in no calculation at all.

Now compare the two numbers the moment produces. The average available at the line is 1.47x. The market looks like it has paid. The median outcome across everyone who ever entered is 0.65x. The typical participant lost a third of their capital.

Both figures are correct. They are computed from the same market on the same day. The first describes a population at an instant. The second describes what happened to people.

Note also where the average sits relative to the survivors themselves. It is above most of them, carried there by a handful of large winners. Even among those who lasted, the average is not the experience.

Hendrik Bessembinder ran the same accounting on real companies. Across roughly 25,300 US companies since 1926, four out of every seven returned less over their lifetime than one-month Treasury bills. The best-performing four percent produced the entire net gain of the American stock market. Everything else, taken together, matched cash. The single most common lifetime outcome for an individual stock is a loss of one hundred percent, and the median company stays listed for seven and a half years.

Read those facts side by side. The index compounded handsomely. The typical constituent did worse than a savings account and then disappeared.

Both statements are true. They are answers to different questions, and only one of them is the question an investor is asking.

Note what the index does that no investor is permitted to do. When a constituent fails, it is removed and replaced with something healthier. The index cuts its losers and rides its winners, by rule, forever, and it never has to fund a margin call while it waits. It is immortal by construction. You are not.

An ensemble average describes a population at an instant. Wealth is built along a path through time, and the path is where ruin lives.

Our own test makes it concrete. Take the S&P’s real returns, borrow to double your exposure, and call it ruin if you lose half your money. Across reshuffles of those identical returns, ruin occurs in 100% of orderings.

The distribution never changed. The volatility never changed. What changed is that a human being had to walk through the returns one day at a time, in order, with a floor beneath them.

Mathematics averages over every possible future. You have to survive the one that arrives.

The trap, stated plainly

The industry adopted variance because it can be added, optimised and computed, and those three properties are what let you build a portfolio model, a capital rule and a performance ranking. The price of that convenience is a number that is silent about order, blind to the difference between profit and loss, symmetric where reality has a floor, and descriptive of an average life that nobody actually lives.

And the thing it is meant to protect you from does not look the way it is imagined. It is not a morning. It is seven years.

Variance is a property of a sample. Ruin is a property of a sequence.

None of this would matter if the errors were small. Episode Three is a catalogue of how large they get.

The capital destruction does not happen because anybody ignored the risk model. It happens because they used it.

 

 

Next, an addendum to this episode: The Unwritten Sequence.

This argument leaves one objection standing, and it is the objection a careful reader raises first. If every shuffled ordering finishes in the same place, why should you care about the path?

Three interactive experiments answer it. The first concedes the point completely: ten portfolios, one deck, identical statistics to fifteen decimal places and an identical closing value. The second shows what happens the moment a single pound crosses that path, and how the decisions you make on the only evidence you will ever have turn one asset into wildly different amounts of money. The third shows what you can build that does not depend on which ordering you got, and why not one number on the risk report will ever recommend it.

Then, in Episode Five: what to do instead, and the five things we got wrong.

Richard Brennan writes on systematic trading, complex adaptive markets, and the philosophical foundations of trend following at atstradingsolutions.com. His books include The Fractals of Finance, Complex Adaptive Markets, Carved by Impossibility and The Aussie Turtles Trend Following Guide.

Want to explore why structure exists at all?

Carved by Impossibility: What Remains When Everything Else Is Eliminated

The book explores the architecture of constraint, emergence, and reality itself, and what it means for how we understand markets, life, and the universe.

Available now on Amazon in paperback, hardcover, and Kindle.

Want the theoretical foundation for why markets adapt?

Complex Adaptive Markets: How Living Systems Shape Finance

The book explores the full architecture of feedback, emergence, and adaptive behaviour in financial markets, and what it means for how we trade, invest, and understand risk.

Available now on Amazon in paperback, hardcover, and Kindle.

Want the theoretical foundation for why trend following works?

The Fractals of Finance: Determinism, Adaptation and the Geometry of Markets

The book explores the full architecture of feedback, fat tails, and fractal structure in financial markets, and what it means for how we trade, invest, and understand risk.

Available now on Amazon in paperback, hardcover, and Kindle.

Want a practical field manual for trading trends and capturing outliers?

The Aussie Turtles Trend Following Guide: A Field Manual for Hunting Outliers adapts the timeless principles of the original Turtle traders into a systematic, rules-based approach for modern markets. Co-authored with Adam Havryliv.

Available now on Amazon in paperback, hardcover, and Kindle.

Share this post:

Facebook
LinkedIn
X