The Vault

THE GEOMETRY OF WEALTH | Episode 3 of 15: The Sharpe Mirage

Why the most trusted metric in finance is optimised for a world that does not exist, and how it hides the geometric superiority of the processes that protect compounding

The most widely used risk metric in finance penalises you equally for making too much money as for losing it. And that is not even its biggest problem.

The Sharpe ratio is the industry’s universal language of risk-adjusted performance. It is the first number on every fund fact sheet. It is the gatekeeper through which institutional allocators filter thousands of managers down to a shortlist. It is the metric by which pensions, endowments, sovereign wealth funds, and consultants evaluate whether a return stream is “good” relative to the risk taken. It is, by any measure, the single most influential number in investment management.

It is also, for a significant category of investment processes, precisely wrong.

In the first two episodes of this series, we established that wealth compounds multiplicatively, that variance is a tax on compounding, and that this tax is collected convexly through drawdowns. We showed that the depth of drawdowns, not the average of returns, is the dominant determinant of terminal wealth. A logical corollary follows: the metric by which you evaluate a return stream should be sensitive to drawdowns, to asymmetry, and to the difference between upside and downside volatility. The Sharpe ratio is none of these things. It is a metric designed for a symmetric, Gaussian, additive world. We do not live in that world. And the distortion it introduces is not random. It is systematic, and it is directional: it penalises exactly the return profiles that geometric compounding rewards.

Three Assumptions, Three Failures

The Sharpe ratio is calculated as the portfolio’s excess return over the risk-free rate, divided by the standard deviation of those returns. The formula is simple. The assumptions embedded in it are not, and each of them fails in ways that matter for real wealth.

Assumption One: Returns are additive. The denominator of the Sharpe ratio, standard deviation, is a measure defined in additive space. It assumes that a +10% return and a −10% return are equally sized events that cancel each other out. In a multiplicative world, they do not cancel. A −10% return followed by a +10% return leaves the investor at 99 cents, not $1.00. Over long horizons, the gap between additive and multiplicative accounting grows exponentially. The Sharpe ratio evaluates returns as though the investor were adding chips to a pile. Real investors are compounding capital, and compounding is multiplication, not addition. This mismatch means the Sharpe ratio is measuring performance in a coordinate system that does not correspond to how wealth actually grows.

Assumption Two: Returns are normally distributed. Standard deviation is the correct measure of dispersion for a Gaussian distribution. For distributions that are fat-tailed, skewed, or both, it is misleading. Financial returns are empirically fat-tailed (extreme events occur far more frequently than the normal distribution predicts) and frequently skewed (the tails are not symmetric). The Sharpe ratio uses a Gaussian ruler to measure a non-Gaussian world. It underestimates the probability of extreme events, both good and bad, and it cannot distinguish between a distribution that is dangerously fat on the left (prone to catastrophic losses) and one that is productively fat on the right (prone to windfall gains). Both look like “high volatility” to the Sharpe ratio. They are not the same thing.

Assumption Three: Upside and downside volatility are equally dangerous. This is the most consequential failure. Standard deviation treats every deviation from the mean identically, whether the deviation is a gain or a loss. In the Sharpe ratio’s view of the world, a month that delivers +8% is exactly as “risky” as a month that delivers −8%. This is mathematically coherent within a Gaussian, additive framework. It is geometrically absurd. Episode 2 demonstrated that losses and gains are not mirror images: a 50% loss requires a 100% gain to recover. Upside volatility is wealth. Downside volatility is destruction. A metric that cannot distinguish between them is a metric that cannot evaluate compounding.

The Sharpe ratio does not measure risk. It measures dispersion. In a world where upside dispersion is wealth and downside dispersion is destruction, treating them as equivalent is not simplification. It is distortion.

The Mirage in the Data

The theoretical failures of the Sharpe ratio would be an academic curiosity if they did not produce real distortions in real data. They do. And the distortion is devastating.

Here are the Sharpe ratios for our three benchmarks over the 26-year period from January 2000 to January 2026:

Read the Sharpe column. The S&P 500 and the TTU Trend Following Index are, according to the Sharpe ratio, essentially identical investments. 0.457 versus 0.468. A consultant evaluating these two return streams on the basis of Sharpe ratio would conclude that they offer roughly the same risk-adjusted performance. They would be correct in a Gaussian, additive, symmetric world. In the actual world, they are categorically different animals.

The S&P 500 is negatively skewed (−0.509): its tail risk is concentrated on the left, in catastrophic losses. Its maximum drawdown required a 104% recovery. It spent 21% of the 26-year period more than 20% below its peak.

The TF Index is positively skewed (+0.230): its tail risk, such as it is, leans toward the right, toward outsized gains. Its maximum drawdown was 21.0%, requiring only a 27% recovery. It spent two months of 26 years below 20%.

The Sharpe ratio looked at these two profoundly different return streams and declared them equivalent. It did so because upside volatility in the TF Index (which is a geometric virtue) offset the lower downside volatility (which is also a geometric virtue), producing a standard deviation that is only modestly lower than the S&P 500’s. The TF Index’s geometric superiority is entirely invisible to the Sharpe ratio because the qualities that make it superior, positive skew and drawdown containment, are precisely the qualities that the Sharpe ratio is not designed to detect.

The S&P 500 and the TF Index have nearly identical Sharpe ratios. One suffered a 51% drawdown. The other never breached 21%. The metric cannot tell them apart. This is not imprecision. It is blindness.

The Upside Penalty

The distortion becomes even clearer when we decompose volatility into its directional components. Standard deviation treats all dispersion equally. But what happens when we separate the volatility that comes from positive months from the volatility that comes from negative months?

The S&P 500 has more downside volatility than upside volatility. Its up/down ratio is 0.81, meaning its bad months are more violent than its good months. This is the structural fingerprint of negative skew: the left tail is fatter than the right. In geometric terms, this is the worst possible configuration, because the violent downside months inflict convex damage on the compounding base that the milder upside months cannot repair.

The TF Index has the opposite profile. Its up/down ratio is 1.29: its good months are more volatile than its bad months. The right tail is fatter than the left. This is the structural fingerprint of positive skew, and it is precisely the configuration that geometric compounding rewards. The violent months are the profitable months. The calm months are the losing months. The asymmetry works for the investor rather than against them.

The Sharpe ratio combines these two volatilities into a single number and divides the excess return by it. By doing so, it penalises the TF Index for its upside volatility. The outsized positive months that drive geometric wealth creation are counted as “risk” in the Sharpe denominator. The metric treats a +8% month as identically risky to a −8% month. In a multiplicative world, the +8% month is a gift to the compounding engine and the −8% month is a wound. Sharpe cannot distinguish between a gift and a wound.

What the Distributions Reveal

The skew numbers in the table above are not isolated data points. They are population-level findings that describe a structural property of the trend following process.

Chart 6: Monthly return distributions: S&P 500 (negatively skewed, −0.509) and TTU TF Index (positively skewed, +0.230). The shape of the distribution determines whether the Sharpe ratio is a useful measure or a misleading one.

Chart 6 overlays the monthly return distributions of the S&P 500 and the TF Index. The shapes are different in ways that matter enormously for compounding but are invisible to the Sharpe ratio. The S&P 500’s distribution has a fat left tail: it produces catastrophic months more frequently than a normal distribution would predict, and those months inflict convex damage on the compounding base. The TF Index’s distribution has a fatter right tail: its outlier months lean positive. These outlier months are exactly the events that drive long-term geometric wealth, and the Sharpe ratio penalises the strategy for producing them.

This is not an artefact of the TF Index specifically. It is a characteristic of the process.

Chart 7: Skewness across all 41 trend following managers with 20+ year track records. 37 of 41 (90%) are positively skewed. The S&P 500 is negatively skewed (−0.509). Positive skew is a structural output of the trend following process.

Chart 7 shows the distribution of skewness across all 41 long-term trend following managers. The result is stark: 37 of 41, or 90%, are positively skewed. The median skewness is +0.286. The S&P 500, at −0.509, sits on the opposite side of zero from virtually the entire trend following population. Only four managers have negative skew, and their skew values are close to zero.

Positive skew is not a coincidence. It is a mechanical consequence of the trend following process: cutting losses caps the left tail, riding trends extends the right tail. The result is a distribution that leans positive, that produces its outlier months on the upside rather than the downside. This is the geometric ideal. And the Sharpe ratio, by treating upside and downside dispersion as equivalent, structurally penalises the strategy for possessing it.

The Metric Shootout

The consequences of this distortion become concrete when we ask a practical question: how does the choice of metric change the verdict on trend following?

Only 9 of 41 long-term trend following managers beat the S&P 500 on Sharpe ratio. Evaluated by the industry’s standard metric, trend following looks like a strategy where fewer than one in four managers can match a passive index. This is the number that consultants, pension boards, and institutional gatekeepers would see.

22 of 41 long-term trend following managers beat the S&P 500 on MAR ratio. Evaluated by the metric that measures what actually matters for geometric wealth, CAGR per unit of maximum drawdown, more than half of the trend following population outperforms the index. The metric determines the verdict. And the industry uses the wrong metric.

Chart 8: The metric shootout: 41 trend following managers ranked by Sharpe ratio versus MAR ratio. Berkshire Hathaway and S&P 500 marked for reference. 9 managers beat the S&P on Sharpe. 22 beat it on MAR. The metric determines the verdict.

Chart 8 makes the comparison visual. The same 41 managers, ranked by two different metrics, produce two different pictures of the trend following universe. By Sharpe, the majority of trend followers appear to underperform the S&P 500. By MAR, the majority outperform it. The same data, the same managers, the same returns. Only the lens has changed. And the geometric lens, the MAR ratio, reveals a reality that the Sharpe ratio obscures.

The divergence is largest for the most positively skewed managers. The managers whose return profiles are most geometrically favourable, the ones with the fattest right tails and the most contained left tails, are precisely the managers that the Sharpe ratio undervalues relative to the MAR ratio. This is not a coincidence. It is a direct consequence of the Sharpe ratio’s inability to distinguish between upside and downside volatility. The more asymmetric the return stream, the more the Sharpe ratio distorts the evaluation.

Consider a specific case. Mulvaney Capital’s Global Diversified Programme has the highest CAGR in the dataset at 18.03%, a Sharpe ratio of 0.578, and a skewness of +0.59. By Sharpe, it ranks 4th among the 41 managers. By MAR, it drops to 9th, penalised by its 60.9% maximum drawdown. But within the Sharpe framework itself, the metric is already undervaluing Mulvaney’s actual contribution to geometric wealth. Mulvaney’s violent upside months, the months that contributed most to its extraordinary compounding, are counted as “risk” in the Sharpe denominator. The metric punishes the strategy for the very characteristic that made it the best compounder in the dataset.

The Metrics That Respect Geometry

If the Sharpe ratio is the wrong tool, what are the right tools? There are several, and each captures a different dimension of geometric efficiency that the Sharpe ratio ignores.

The MAR Ratio (CAGR / Maximum Drawdown) measures the return per unit of worst-case drawdown. It directly quantifies the relationship that Episode 2 showed to be the dominant driver of terminal wealth: how much compounding the process delivers relative to its deepest damage to the compounding base. The S&P 500’s MAR is 0.157. The TF Index’s is 0.359, more than double. Berkshire’s is 0.240. Where the Sharpe ratio declared the S&P and TF Index equivalent, the MAR ratio reveals the TF Index to be geometrically more than twice as efficient.

The Sortino Ratio replaces total standard deviation with downside deviation: it penalises only the volatility that hurts. This single modification corrects the most damaging of the Sharpe ratio’s failures. The S&P 500’s Sortino is 0.659. The TF Index’s is 0.772. Berkshire’s is 0.923. The Sortino ratio reveals a gap between the S&P and the TF Index that the Sharpe ratio hides, because it stops penalising the TF Index for its upside months. It is not a perfect metric, and it does not directly measure the convex cost of drawdowns, but it is a significant improvement on the Sharpe ratio for any positively skewed return stream.

Skewness and the Shape of the Distribution are not performance metrics per se, but they are diagnostic tools that the Sharpe ratio ignores entirely. Knowing that a return stream is positively skewed tells you that its outlier months lean toward wealth creation rather than wealth destruction. Knowing that 37 of 41 trend following managers are positively skewed, while the S&P 500 is negatively skewed, is information of first-order importance for anyone whose goal is geometric wealth. No single number captures this. The shape of the distribution is a richer source of information than any ratio, and the Sharpe framework discards it entirely.

These metrics share a common principle: they respect the multiplicative, asymmetric, non-Gaussian nature of real returns. They do not treat upside and downside as equivalent. They do not assume a bell curve. They measure performance in the coordinate system that actual wealth inhabits, not the idealised Gaussian plane that the Sharpe ratio was designed for.

Why the Industry Clings to the Wrong Metric

If the Sharpe ratio is this misleading, why does the industry continue to use it? The answer is not ignorance. The answer is infrastructure.

The entire apparatus of institutional investment management is built around the Sharpe ratio and its close relatives. Mean-variance optimisation, the framework that determines how most pension funds and endowments allocate capital, uses the Sharpe ratio as its objective function. A portfolio is “optimal” in mean-variance space when it maximises the Sharpe ratio. Consultants rank managers by Sharpe. Due diligence questionnaires ask for Sharpe. Regulatory frameworks reference Sharpe-adjacent metrics. The infrastructure is so deeply embedded that changing it would require rebuilding the analytical foundations of institutional capital allocation.

The result is a self-reinforcing system. Managers are evaluated on Sharpe. Managers therefore optimise for Sharpe. Optimising for Sharpe means smoothing returns, suppressing upside volatility, and avoiding the lumpy, positively skewed return profiles that geometric compounding actually rewards. The metric does not just measure badly. It incentivises the wrong behaviour. A manager who delivers steady 6% returns with low volatility will score better on Sharpe than a manager who delivers 8% returns with occasional explosive positive months. The first manager is easier to market. The second manager compounds more wealth. The industry’s metric system selects for the first and penalises the second.

This is the deeper consequence of the Sharpe mirage. It is not merely that one number is imprecise. It is that an entire ecosystem of evaluation, allocation, and incentive is calibrated to a metric that is structurally blind to the properties that matter most for geometric wealth. Trend following, with its positive skew, its crisis alpha, and its drawdown containment, is precisely the class of strategy most undervalued by this ecosystem. The mirage is not a technical footnote. It is a structural reason why the investment industry systematically underweights the processes that compound most efficiently.

The Running Ledger

Our $100,000 continues. This episode adds the risk metrics that define how the industry evaluates each benchmark, and how the evaluation changes when the right metrics are used.

Read the table across the Sharpe row: 0.457, 0.545, 0.468. The TF Index and S&P 500 look nearly identical. Now read the MAR row: 0.157, 0.240, 0.359. The TF Index is more than double the S&P 500 and 50% higher than Berkshire. The Sortino tells the same story as MAR, with less drama: 0.659, 0.923, 0.772. The skew column explains why: the S&P 500’s negative skew and the TF Index’s positive skew produce categorically different geometric properties that the Sharpe ratio cannot detect and the other metrics can.

The choice of metric is not a technical detail. It is the lens through which the entire investment industry evaluates performance. And the industry’s default lens is the one that happens to be blind to the exact properties that geometric compounding rewards.

The Bridge

We have now established three layers of the wrong mathematics. Arithmetic averages overstate wealth. The convex cost of drawdowns is systematically ignored. And the industry’s dominant performance metric is structurally blind to the properties that matter most for geometric compounding.

But why do these failures persist? They persist because the mathematical framework that underpins them, the Gaussian model of financial returns, assumes a world of smooth, bell-shaped distributions where extreme events are vanishingly rare. In this world, the arithmetic average is a reliable guide. Drawdowns are moderate and symmetric. Standard deviation captures all relevant risk. The Sharpe ratio is sufficient.

We do not live in that world. Financial returns are not normally distributed. They are fat-tailed: extreme events occur orders of magnitude more frequently than the Gaussian model predicts. Black Monday, the dot-com crash, the Global Financial Crisis, the COVID collapse. These are not anomalies. They are structural features of the system. And they change everything about how wealth is built and destroyed.

In Episode 4, we will measure the distance between the Gaussian fantasy and the fat-tailed reality, and show why the world that actually exists is the world for which trend following was built.

Data and Sources

All performance data from the NilssonHedge Trend Following Performance Database (January 2000 to January 2026). All returns are net of management and performance fees. Sharpe ratios calculated using annualised excess returns over an assumed 2% risk-free rate, divided by annualised standard deviation. Sortino ratios use downside deviation only (returns below risk-free rate). MAR ratio is CAGR divided by absolute maximum drawdown. Skewness is the Fisher-Pearson standardised moment coefficient computed from monthly returns. Upside and downside volatility are annualised standard deviations computed separately on positive and negative monthly returns. All statistics are computed from monthly data and annualised where appropriate. The 9/41 and 22/41 comparisons use the S&P 500 Total Return Sharpe (0.457) and MAR (0.157) as thresholds.

Want the theoretical foundation for why markets adapt?

Complex Adaptive Markets: How Living Systems Shape Finance

The book explores the full architecture of feedback, emergence, and adaptive behaviour in financial markets, and what it means for how we trade, invest, and understand risk.

Available now on Amazon in paperback, hardcover, and Kindle.

Want the theoretical foundation for why trend following works?

The Fractals of Finance: Determinism, Adaptation and the Geometry of Markets

The book explores the full architecture of feedback, fat tails, and fractal structure in financial markets, and what it means for how we trade, invest, and understand risk.

Available now on Amazon in paperback, hardcover, and Kindle.

Want a practical field manual for trading trends and capturing outliers?

The Aussie Turtles Trend Following Guide: A Field Manual for Hunting Outliers adapts the timeless principles of the original Turtle traders into a systematic, rules-based approach for modern markets. Co-authored with Adam Havryliv.

Available now on Amazon in paperback, hardcover, and Kindle.

Share this post:

Facebook
LinkedIn
X