Five episodes, 68 futures markets, 42 years. An empirical audit of the number modern finance calls risk.
A Series Synopsis
A five-part empirical investigation into whether the number modern finance calls risk actually measures the thing investors experience.
The premise
Modern finance measures risk by squaring.
To work out how much an investment moves around, you take each day’s return, measure how far it landed from the average, and square that distance. Sum the squares, average them, take a square root. The result is standard deviation, and the profession calls it risk.
That single operation is load-bearing across almost the entire discipline. It is the risk term in Modern Portfolio Theory. It is the bottom half of the Sharpe ratio. It is the beta in the Capital Asset Pricing Model. It is the volatility input to Black-Scholes. It is the engine inside Value at Risk, which regulators wrote into banking law.
This series asks a simple question and then answers it with data: if variance really measures risk, what should we observe in real markets?
We tested it across 68 global futures markets, 647,629 market-days, from September 1984 to July 2026. Currencies, government bonds, metals, grains, energy, equity indices. One consistent method applied to all of it.
The central claim, that risk and variance are different things, is not original. Mandelbrot made it in 1963. Taleb has spent thirty years making it. Ole Peters has done the rigorous mathematical work. What this series adds is the evidence, at a scale and with a consistency that has not been assembled before.
It also publishes, in full, the five hypotheses we held going in that the data destroyed.
The arithmetic, and an honest defence of it.
Squaring does not merely notice an unusual day. It magnifies it, and the magnification compounds with distance. A day three times larger than normal counts nine times as much. A day five times larger counts twenty-five times as much.
Across the 68 markets, the consequences are stark. The busiest 1% of days carry 20% of the entire risk number, against 8.5% in a bell curve. One afternoon per calendar year is doing a fifth of all the work.
The signature test uses two rulers. Standard deviation squares the distances. Mean absolute deviation does not. In a true bell curve their ratio settles at exactly 1.2533, a mathematical fact rather than an estimate. Anything above that means more extreme days than the bell curve permits.
Sixty-eight markets. Sixty-eight results above the line. Not one exception, across four decades, spanning asset classes with almost nothing in common.
Then the episode does something unusual for the genre. It argues the other side, properly.
In a bell curve the square is not merely acceptable. It is provably optimal. A normal distribution is completely described by two numbers, so measuring sigma measures everything there is to know. Three quarters of the variance sits within two standard deviations of the mean, so the extremes are a rounding error. And Fisher proved in 1920 that under exact normality, standard deviation extracts more information per observation than any alternative. Our own simulation confirms it: on clean data the square wins.
The trap is in how narrow that domain turns out to be.
John Tukey asked in 1960 how far you can drift from normality before the optimality reverses. The answer is roughly one contaminated observation in three hundred. Not a margin of safety. A knife-edge.
Real markets sit at one in 243 on the four-sigma days alone. And those days, which a bell curve says should carry a tenth of one percent of the risk number, actually carry 12.1%. One hundred and seven times the load the model assumes.
The closing insight is the one that makes the square dangerous rather than merely wrong. The property that makes it powerful and the property that makes it fail are the same property. Quadratic weighting extracts maximum information from clean data, which is exactly why Fisher could prove it optimal, and amplifies contamination by the square of its distance, which is exactly why it collapses when data is not clean. You cannot keep one and discard the other.
What the square costs a manager whose living depends on rare events.
Squaring destroys the sign. A 5% gain and a 5% loss enter the risk calculation as identical contributions, because both squares are positive. The formula has no way of knowing whether a day was the best of your career or the worst.
For most investors that is a curiosity. For an outlier hunter it is the central fact of professional life.
Using a deliberately unoptimised 50/200 moving average crossover, 1970s technology that anyone could have run in 1990, the series follows 3,696 trades over 34 years. The system won 38.5% of the time, so it was wrong roughly six times in ten.
The best 170 trades, 4.6% of the total, produced 100% of the net profit. The other 3,526 netted to zero.
Then the episode withdraws a claim that circulates widely in trend following circles, one we had told ourselves.
The claim: big winners inflate volatility, volatility sits under the Sharpe ratio, so the scoreboard punishes you for your best trades. Delete the winners and the score improves.
It is false. Removing the best trade lowers the score in 68 of 68 markets. Removing the hundred best days drops the portfolio Sharpe from 0.98 to 0.44. Anyone who tells you great trades hurt your Sharpe ratio has not run the numbers.
The truth is quieter and considerably worse.
October 2008, the best month in the system’s entire history, delivered 3.4% of all the money it ever made. Its effect on the headline Sharpe ratio was to move it from 0.914 to 0.931. An improvement of 1.8%.
The measure does not reject the outlier. It pays you about two cents on the dollar for it, because the same enormous return that made the money also inflated the volatility figure it is divided by.
And the ruler moves while you are being measured. The system’s true long-run Sharpe is 0.98. Across rolling three-year windows, which is the horizon on which nearly every capital decision gets made, the measured score ranged from -0.86 to +2.87. 17.1% of windows reported a negative score on a strategy that demonstrably makes money.
Following the square into the machinery.
Six places where standard deviation is load-bearing, each tested against the same 68 markets, each asking: what should we see if the model is right?
Value at Risk. A 99% VaR promises to be breached one day in a hundred. Observed breach rate: 1.81%. Nearly double, with 67 of 68 markets breaching too often. When the limit does break the average loss is more than a third larger than the limit, and the worst breach in four decades was eleven times it.
The leverage trap. This is the one that kills. Position size is set by dividing a target by recent volatility, so when volatility falls, the position grows automatically. Nobody decides anything.
The defence is real, and we already knew it. Our earlier Fractals of Finance work had shown that volatility clusters, so we expected calm to predict calm, and it does, with shocks about three times less likely after a quiet stretch. While the calm holds, the tripled position is working for you, which is exactly why volatility targeting outperforms in quiet markets. The exposure is at the transition. When the calm breaks, and it always does, the returning volatility arrives while you are carrying the largest position the quiet ever talked you into. Calm shrinks the eventual shock to about a third of its size while more than tripling your position, and the second effect is larger.
The result: 52 of 68 markets suffered their worst levered loss coming out of a quiet market, not a stressed one. The worst days arrive out of below-average volatility, while the book is oversized. Quiet going in, and then the turn.
The specimen case is the Japanese Government Bond market in the summer of 2018, the quietest major market on earth, with trailing volatility of 0.9% a year. The model authorised 16.4 times leverage, and while the calm held that position was profitable. Then, on 1 August, the Bank of Japan loosened its grip and the calm broke. The bond moved six tenths of one percent, and nine percent of the capital was gone. The market barely moved. The leverage did the killing, and the long silence had loaded the gun.
Diversification. Correlation rises in a crisis, in every sector, without exception. Across the full universe it more than triples, and roughly half your diversification evaporates at the only moment you needed it.
Beta. Silver’s beta averages 2.04 and has ranged from 0.20 to 7.34 across individual years. 23 of 68 markets have a beta that changed sign from one year to the next. A number that reverses direction depending on your sample is not describing the asset.
Black-Scholes. This is where the mathematics moves furthest away from experience.
The model does not merely say extreme days are unlikely. It says they cannot happen. The S&P’s October 1987 move was 24 standard deviations, which Black-Scholes rates as occurring once every 10¹²⁶ years. The universe is about 10¹⁰ years old. 67 of 68 markets have produced a move the model calls impossible.
Manager selection. One review in six, conducted on a sound strategy, concludes it destroys value.
The unifying pattern: in every case the model is calibrated on the ordinary and then applied to the extraordinary. These are not faults in the implementations. They are the assumption, arriving on schedule.
The thesis, proved with a deck of cards.
Take every daily return the S&P has produced over forty years. Shuffle them. Add nothing, remove nothing, change no number. Deal the same deck in a different order.
Every statistic a risk model can see survives untouched. The mean does not move. The variance is identical to fifteen decimal places. Skew, kurtosis, Sharpe ratio, every correlation with every other market: unchanged. None of them contains any information about order, because they are computed by adding things up, and addition does not care about sequence.
Deal the deck three thousand times, and the maximum drawdown ranges from -33.6% to -83.8%.
Fifty percentage points of capital, separating histories that every risk model on earth regards as identical. One ordering leaves an investor bruised. Another closes the fund.
What does not move is the endpoint. Deal the same deck any way you like and it compounds to the same closing value, because multiplication does not care what order you multiply in. That holds for exactly as long as nothing responds to the path while it unfolds. The moment a dollar is withdrawn, a position is sized, or a decision is taken on the recent record, the endpoint stops being order-blind too.
The episode also makes an honest correction. We expected to show that volatility tells you nothing about how badly you can be hurt. That is too strong. Volatile markets really do have bigger drawdowns, and sigma explains about three quarters of the variation between markets. What it cannot do is locate you inside the range, and the range is fifty points wide. Between markets and within one market are different questions, and sigma answers only the first: across a single market’s own possible orderings it does not vary at all. Volatility gives you the neighbourhood. It cannot give you the address.
Three further failures follow. Squaring treats a 30% gain and a 30% loss identically, but losing 50% requires making 100% to recover, and at the end of the road sits a door that opens only one way. The optimiser maximises the arithmetic mean, but you receive the compounded return, and volatility eats the difference: silver’s arithmetic return was 5.4% a year while holders kept 0.8%.
And finally the deepest version, following Peters. Every statistic in the standard toolkit describes what happens on average across many parallel lives. Nobody lives many parallel lives. You get one, in one order, with a floor beneath you. Under 2x leverage on real S&P returns, ruin occurs in 100% of reshuffles of the identical return series.
The objection, answered in three experiments.
Ten orderings of thirty years of S&P returns produce one identical risk report and one identical closing value, so the fair objection is that the path was scenery. It is not. The endpoints coincide only because nobody touched the portfolio for thirty years, and no investor, fund or allocator has ever met that condition.
Draw an income of 4.8% a year, identical to the dollar in all ten orderings, and closing capital runs from 941 to 6,735 per thousand committed while the tearsheet reports 8.55% a year to every one of them. Redeem the manager the first month the trailing three-year Sharpe ratio goes below zero, a rule no committee would call reckless, and it fires in ten cases out of ten, with closing capital from 768 to 8,367.
And nobody decides on the thirty-year number, because it does not exist yet. Across every possible ten-year holding period in the same series, nearly one investor in eight lost money on an asset that compounded at 8.55% a year. Not because the asset failed. Because of where they came in.
The third experiment asks what kind of intervention repairs a path. A symmetric volatility target improves every number on the report, volatility, Sharpe, Sortino, Value at Risk and worst month, and leaves the investor with forty percent less money. An asymmetric rule that estimates nothing, holding full exposure only while the index sits above its level ten months earlier, cuts the worst loss from -52.6% to -20.4%, raises the compound return, and holds up on a holdout back to 1871. Scramble the order and its entire advantage disappears, which locates that advantage in precisely the dimension no additive statistic is measured along.
The correct solution is the one that reports badly.
This is one index, measured monthly, inside a series that rests everywhere else on 68 markets and 647,629 market-days. It is offered as a demonstration rather than a survey. The three experiments run in the page, so every figure quoted here can be recomputed by moving the settings.
The prescription, and the ledger of failures.
This episode contains the series’ most important contribution: not a belief overturned, but a mechanism isolated and measured.
The standard defence of Modern Portfolio Theory says fat tails are a single-market problem, and that spreading money across enough markets makes them average out. This rests on a real theorem: add enough independent random things and you converge on a bell curve. In real markets it is false, and we did not need the research to know it. Our own design already assumes crisis correlation: small risk per stream, ensembles structured to hold correlation down, and no reliance on breadth for tail protection. What the research adds is the proof.
We built portfolios of increasing size two ways. The real portfolio used actual data, with markets free to crash together. The independent portfolio used identical markets with identical fat tails, but with each history shuffled separately so nothing crashed simultaneously. Same ingredients. The only difference is co-movement.
The independent portfolio behaves exactly as the textbook promises. Kurtosis collapses from 9.0 to 0.10 as markets are added. Diversification cures fat tails perfectly.
The real portfolio does not. Kurtosis falls briefly, then climbs back to 7.7. A fully diversified 68-market book ends up nearly as tail-heavy as a single market, and worse than a five-market book.
Diversification thins the body. It does not thin the tails.
The reason is that extreme events do not cancel. They arrive together. Diversification strips out the ordinary noise while leaving the extraordinary events untouched, and what remains is a thinner body attached to the same tails, which is a more extreme distribution rather than a less extreme one.
But the tail cuts both ways, and this is what redeems diversification for a trend follower. A fat-tailed distribution is fat on both ends, so a wider sample reaches further into the winning tail as well. This overturns a promise buried in Modern Portfolio Theory, which sells diversification as variance reduction bought at the cost of return. In a leptokurtic world, held at constant risk, the sign is backwards. Widen the book from one market to 68 at a fixed portfolio volatility and the compound return climbs monotonically, from under 2% to nearly 8%, because the single best trade reaches from around 22% to 44% while the worst stays capped by per-trade sizing. Diversification is not the return-diluting hedge the textbook describes. Held at constant risk it is powerfully return-accretive. It just is not a shield: the correlations still converge in the crash.
This also explains why the obvious prescription fails. We tested swapping standard deviation for mean absolute deviation, downside deviation, and expected shortfall. Changing the ruler does almost nothing. MAD, which Episode One praised, produced a marginally worse result. On a diversified book the ruler is not the binding constraint, because the tails are correlated, they survive the diversification, and changing how you measure a fat tail does not make it thinner.
We were fully prepared to recommend a better estimator as the conclusion of this series. The data refused it.
What did work was a hard leverage cap, which improved the worst loss on every book structure tested. It works not by measuring risk more accurately, but by breaking the automatic link between a quiet market and a large position, which is precisely the Japanese bond failure.
Underneath the individual fixes is a single design philosophy, and it is the mirror image of the Gaussian one. A mean-variance investor holds the market to capture its average, exposed to the whole distribution. A trend follower built for fat tails does the opposite: refuse the dull middle by only entering on a qualifying trend; manufacture positive skew by cutting losers with strict stops, which in our data turns a market skew of essentially zero into a realised +0.7; pyramid into major moves through an ensemble of different structures rather than through leverage, so exposure scales with the significance of the move; and size every entry small against realised capital so none of this risks principal. Not one of these is visible inside a Gaussian frame. They are what you build once you stop believing it. Our separate Geometry of Wealth work found the same signature from the outside, in real manager track records: 37 of 41 long-run trend followers produce positive skew, and the best tenth of months carries roughly 90% of terminal wealth.
The five hypotheses that died
- Deleting winners would improve the Sharpe ratio. It lowers it, 68 of 68.
- Volatility misses real risk entirely. It explains about three quarters of drawdown variation between markets.
- The long-run volatility estimate is visibly unstable. It is not. It settles. The instability lives at the three-year horizon, which is more damning.
- We thought volatility clustering would protect the risk manager. It does not. Calm does predict calm, exactly as our earlier work found, but the models respond to the calm faster than the calm protects them.
- A better ruler is the fix. It is not.
We also published our data hygiene log in full, including the admission that an earlier draft reported sigma explaining only 41% of drawdown variation. That figure was an artifact of corrupted source data. The correct figure is 76%, it is considerably less flattering to our argument, and it is what we published.
What the series amounts to
The square is not a mistake. It is a theorem with a domain of validity, and Fisher proved it optimal inside that domain. The domain is the bell curve, and the neighbourhood around the bell curve where the theorem still holds is roughly one contaminated observation in three hundred. Real markets sit at one in 243.
We are outside the domain. Every model in Episode Three is outside the domain. The profession has spent seventy years building an increasingly sophisticated apparatus on a foundation whose assumptions gradually became accepted as reality.
The practical conclusions are specific.
Cap your leverage in absolute terms. Any model that permits sixteen times exposure because a market has been quiet has confused the absence of recent evidence with the absence of risk.
Diversify for the net, not the tail. Breadth catches more of the rare, large moves that pay for everything. It does not save you in the crisis, so it cannot be the last line of defence.
Grow by widening, not by scaling. When capital increases, add uncorrelated return streams at the smallest workable size rather than enlarging existing bets. Leverage buys return at a worsening exchange rate; a wider book does not.
Enter small, let winners run. The large position worth holding is the one the market built from a winning trade’s own unrealised profit, not the one leverage built from your principal. Realised capital stays defended by small entries and stops; open profit is what rides the trend. Harvest beneficial volatility; never the adverse kind.
Measure the path, not the distribution. Drawdown and the sequence of losses are what an investor experiences.
Size for survival, not for score. The floor at zero has no counterpart on the upside.
Remember where the money comes from. A framework designed to suppress variation is designed to suppress the thing that pays you.
Not all of these carry equal weight. The leverage cap and the widening result are tested and they replicate. The rest describe how we build. They follow from the evidence and from principle, and we have not run them to ground the way we ran the cap.
A note on method
The trading system throughout is a deliberate strawman: a 50/200 moving average crossover with no filters, no tuning, and no optimisation. If the argument only worked on a finely tuned proprietary system it would be an argument about that system. It is an argument about the measuring instrument, and the instrument behaves the same way no matter what you point it at.
Every figure and statistic in the series is drawn directly from the underlying market data, computed by one consistent method across all 68 markets. Where the evidence forced us to revise a claim, or corrected an earlier draft, we have said so plainly in the episodes themselves.
Data: 68 global futures markets, 647,629 market-days, September 1984 to July 2026. Intellectual debts: Mandelbrot (1963), Fisher (1920), Tukey (1960), Taleb, and Peters (2019).
The road we took
We began with a single arithmetic operation and a question about its domain. The square is optimal, provably, inside a bell curve, and the neighbourhood where that proof survives is narrower than almost anyone using it has been told.
From there the audit widened. Episode Two asked what the square costs the person whose living depends on rare events, and found a measure that pays two cents on the dollar for the best month of a career. Episode Three followed the same operation into six pieces of working machinery and found each one calibrated on the ordinary and then pointed at the extraordinary. Episode Four dealt the same returns in a different order and produced fifty percentage points of difference in a history that every risk model on earth reports as identical. The addendum closed the last exit, showing that the endpoint only survives reordering while nobody withdraws, sizes, or decides. Episode Five went looking for the fix, expected to find a better estimator, and did not.
Look back along that sequence and something becomes visible that no single episode states. Each one was measuring a property of a distribution. The last one found the thing that was never in a distribution to begin with.
Risk lives across time, not within it
Risk is not a property of a moment. It is a property of a sequence.
A dam wall does not fail on the day it fails. It fails across the seasons of seepage that came before, each one carrying away a little more material, none of them visible as an event. A forest does not burn because of a spark. It burns because of the brush that accumulated through every season the fire did not come. The spark is only the last term in a long sum.
Engineered systems supply the same pattern, and they keep better records.
Chernobyl did not begin on 26 April 1986. The reactor carried a design flaw that made it unstable at low power, and the flaw had been documented years earlier and never passed to the men working the night shift. The safety test had been postponed, rescheduled, and handed to a crew who had not planned it. Beneath all of it sat nearly a decade of uneventful operation at the station, which is the most persuasive evidence a system can offer that nothing is wrong. Each quiet year lowered the perceived risk and left the fragility exactly where it was, or worse. The explosion was the last term in a long sum.
The shape repeats wherever anyone looks for it. The seals on the Challenger’s boosters had been eroding flight after flight, and each safe return was read as evidence that the erosion was tolerable rather than as a warning nobody had collected yet. Fukushima Daiichi was designed in the 1960s against the largest wave in the local record. New analysis some eighteen years before 2011 pointed to a wave near sixteen metres at that site, and little was done. The wave that arrived was about fifteen.
In every one of these the instruments were reporting on the interval, and the risk was living across the intervals.
Markets work the same way. The quiet years are not periods of low risk. They are periods when risk is being accumulated. Positions grow because nothing has punished them. Leverage rises because the recent record says it is affordable. Hedges lapse because they have cost money for eleven quarters. The fragility that produces the drawdown is built during the calm, by the calm, and it is invisible to any statistic computed inside the calm. That is the Japanese bond market in one sentence, and it is also 2008, and 1998, and every episode of the same shape still to come.
Stated that way it sounds like a truism. Risk is the progressive accumulation of fragility. Nobody who has watched a wall seep or a hillside dry out would call that a discovery. Yet finance has declined to see it, and the reason is not stupidity. It is that fragility accumulating over seasons does not reduce to a closed form, and a closed form is what the discipline rewards. So the profession chose the equations it could solve over the process it could observe, and it has been solving them ever since, elegantly, in a world that does not exist.
This is where the instruments go silent. Standard deviation takes a set of returns and gives back a number. Value at Risk does the same. Sharpe does the same. Each one is computed on a distribution, and a distribution has no memory. Shuffle the sample and every figure is unchanged. That is not a flaw in the implementation. It is what those statistics are. They can describe a bag of outcomes. They cannot describe a process that accumulates.
Which returns us to contingency. The record we have is one realised path, and the industry treats that path as the measurement. It is not the measurement. It is one draw from a system whose capabilities were never fully expressed. The relevant question was never what the market did. It was what the market was capable of doing, and what it had been quietly preparing to do while the numbers looked benign.
This series does not claim to have solved that. It claims only to have said out loud what the instruments cannot say.
Variance is permutation invariant. Risk is not. Therefore variance is not risk.
Richard Brennan writes on systematic trading, complex adaptive markets, and the philosophical foundations of trend following at atstradingsolutions.com. His books include The Fractals of Finance, Complex Adaptive Markets, Carved by Impossibility and The Aussie Turtles Trend Following Guide.
Want to explore why structure exists at all?
Carved by Impossibility: What Remains When Everything Else Is Eliminated
The book explores the architecture of constraint, emergence, and reality itself, and what it means for how we understand markets, life, and the universe.
Available now on Amazon in paperback, hardcover, and Kindle.
Want the theoretical foundation for why markets adapt?
Complex Adaptive Markets: How Living Systems Shape Finance
The book explores the full architecture of feedback, emergence, and adaptive behaviour in financial markets, and what it means for how we trade, invest, and understand risk.
Available now on Amazon in paperback, hardcover, and Kindle.
Want the theoretical foundation for why trend following works?
The Fractals of Finance: Determinism, Adaptation and the Geometry of Markets
The book explores the full architecture of feedback, fat tails, and fractal structure in financial markets, and what it means for how we trade, invest, and understand risk.
Available now on Amazon in paperback, hardcover, and Kindle.
Want a practical field manual for trading trends and capturing outliers?
The Aussie Turtles Trend Following Guide: A Field Manual for Hunting Outliers adapts the timeless principles of the original Turtle traders into a systematic, rules-based approach for modern markets. Co-authored with Adam Havryliv.
Available now on Amazon in paperback, hardcover, and Kindle.