SMOKIN' ACES·Research

It returned 616% with a Sharpe of 1.16. We did not deploy it.

A Weinstein Stage-2 trend sleeve backtested at 616% total return, Sharpe 1.16, profit factor 3.30 and a deflated Sharpe of 0.9997. Every number is real. We shelved it, and the three reasons are more useful than the headline.

We backtested a Weinstein Stage-2 trend sleeve — buy strength above the 30-week moving average, ride the markup phase, exit on the stage change. Here is what came back:

metrictrend sleeveSPY buy & hold
Total return616.35%309.98%
CAGR21.76%
Annualised Sharpe1.160.886
Max drawdown-25.43%
Trades1,301
Profit factor3.304
Average win+33.47%
Average loss-7.58%
Deflated Sharpe (daily)0.9997

Double the market. Better Sharpe. Profit factor above three. A deflated Sharpe of 0.9997, which is the significance gate we hold everything else to.

We shelved it. Here is why, in order of how much each reason cost us to learn.

1. PBO 0.83

The Probability of Backtest Overfitting asks a question a Sharpe ratio cannot: if I pick the best configuration in-sample, how often does it underperform out of sample?

We ran combinatorially-symmetric cross-validation across 252 configuration combinations. The answer came back 0.833.

Five times out of six, the variant that looked best in the backtest was worse than the median variant out of sample. That is the signature of a strategy whose parameters have been fitted to the history rather than to the market — and it is entirely compatible with a spectacular headline return, because the headline is computed on the same history the parameters were fitted to.

A high Sharpe tells you the past was profitable. PBO tells you whether your choice of configuration will survive.

2. It was not a hedge, it was more of the same

The reason we tested it at all was diversification. Our buy list is mean-reversion — buy oversold dips. A trend sleeve buys strength. Intuitively they should be uncorrelated, and a low-correlation sleeve is worth more than a high-return one.

Measured over 2,520 days:

vs SPYvs our mean-reversion buy list
full variantcorr 0.699, beta 0.73corr 0.451, beta 0.284
moving-average corecorr 0.805, beta 0.768corr 0.530, beta 0.306

Correlation of 0.45 to 0.53 against the thing it was supposed to hedge, and 0.70 to 0.81 against the index. That is not a diversifier. It is a leveraged, more volatile expression of the same market exposure we already had.

And this replicated. We independently tested a Gaussian-channel breakout sleeve on the same question and it came back the same way: market-correlated, not a hedge. Two unrelated trend-following systems, two independent tests, the same answer.

Trend-following stock sleeves come back at roughly market beta. If you are
hunting a low-correlation hedge for a mean-reversion book, the trend-following
family is the wrong place to look. We stopped looking there.

3. The survivorship trap, and where it hides

Our equity warehouse is built from currently-listed names. Companies that were delisted are not in it.

For a mean-reversion strategy that is a mild inconvenience. For a trend strategy it is fatal, and the reason is specific: the sleeve's return is dominated by a handful of enormous multi-year winners. The average win is +33.47% against an average loss of -7.58% at a 42.8% win rate — that is a fat right tail carrying everything.

The right tail is the strategy. And the right tail is exactly what a current-listings warehouse over-represents, because the names that compounded for a decade are by definition still listed.

The bias lands precisely on the part of the distribution the edge lives in.

The trap inside the significance test

This is the part worth taking away even if you never trade a stage system.

We computed the deflated Sharpe two ways and got two different answers:

The trade-level number saturates. A long-hold trend sleeve's trade distribution is a few gigantic multi-baggers on top of many small losses — the median trade is a small loss — and the DSR formula, fed that shape, returns approximately 1.0 even at a per-trade Sharpe near 0.3.

The deflated Sharpe assumes something close to a periodic return series. Feed it a lottery-ticket distribution and it reports confidence it has not earned.

Compute the DSR on daily returns, not trade P&L, for any strategy with a
fat right tail. And when two significance calculations disagree by that much,
the disagreement is the finding — one of them is being fed the wrong shape.

The daily figure was also the one consistent with the PBO, which is how we knew which to trust.

What we actually took from it

The framework itself is not the problem. Stage analysis is a sound reactive model: trend-following and drawdown-avoidance are among the better-supported ideas in the literature, and Weinstein's four-stage cycle is a clear way to teach them.

What we could not establish is that our mechanised version of it was an edge rather than a fitted curve on a survivorship-inflated sample of one long bull regime.

So it stayed in research. We built the missing relative-strength pillar properly anyway — because the measurement was worth having even though the answer was no — and it added approximately zero Sharpe over the moving-average core.

A 616% backtest that we did not deploy is not a failure. It is the system working. The failure would have been shipping it because the number was impressive.