SMOKIN' ACES·Research

Our backtest says 1.37. Live says 0.96. The comparison is a trap.

Our Wyckoff signal tracker backtests at profit factor 1.37 across 95,319 trades and returns 0.96 live. The obvious conclusion is that the backtest overstated by 43%. We checked, and the two numbers are not measuring the same thing.

We run a Wyckoff signal tracker separately from our buy list. Here are its two numbers, and they disagree:

closed tradeswin rateavg winneravg loserprofit factor
Backtest95,31945.3%+10.65%-6.45%1.37
Live forward2,54943.5%+8.26%-6.64%0.96

A profitable backtest and a losing live record. The obvious headline writes itself: backtests overstate, ours overstated by 43%, here is your cautionary tale.

We were going to publish that. Then we checked whether the two numbers describe the same thing, and they do not.

The two columns nobody looks at

average days heldstill open
Backtest15.123 of 95,342 (0.02%)
Live7.0967 of 3,516 (27.5%)

Two structural differences, either one of which is enough to break the comparison.

The live trades are measured over less than half the holding period. If the strategy's return accrues across roughly fifteen days, a record whose closed trades average seven days is not a worse version of the backtest. It is a different, shorter trade.

That shows up exactly where you would expect. The average winner fell from +10.65% to +8.26%, a 22% reduction, while the average loser barely moved, -6.45% to -6.64%. Losses hit their stop at a fixed distance regardless of time. Winners need time to run, and these were not given it.

And 27.5% of the live sample has not resolved at all. The backtest is essentially fully closed at 0.02% open. So the live profit factor is computed on a censored sample, and worse, one censored by speed. The trades that have closed are the ones that resolved fastest: quick targets and quick stops. The grinding middle is still open and contributes nothing to the number.

What the arithmetic says each factor is worth

Decomposing the gap from 1.37 to 0.96:

So the collapse is driven mostly by winners being smaller, not by being wrong more often. The win rate moved 1.8 points. The winner size moved 22%.

That points at the holding period rather than at signal quality, which is a testable claim rather than a comforting one, and it is the next thing we will measure.

What we can and cannot conclude

Can: the live record is currently below break-even at 0.96, across 2,549 closed trades. That is a real number about real signals and we are not hiding it.

Cannot: that the backtest was inflated by 43%. To claim that we would need the live trades measured on the same holding rule with the same censoring, and they are not. The honest statement is that we do not yet know how much of the gap is overfitting and how much is measurement.

Cannot, either: that the live number will improve once the 967 open positions resolve. It might. Censoring by speed cuts both ways and we have not measured which way it leans here. Assuming it resolves in our favour would be the same error in the opposite direction.

Why this matters more than the number

Comparing a backtest to a live record is the single most common test in quantitative trading, and it is almost always run wrong, because the two datasets are produced by different processes and nobody checks the join.

Ours differ in holding period and in completion rate. Yours might differ in survivorship, in fill assumptions, in fee treatment, or simply in the fact that the backtest closed every position while the live book has not.

Before you compare a backtest to a live record, check that they measure the
same thing: sample size, holding period, completion rate, and what happens to a
trade that never resolves.

Our published rule, from testing 21,191 book rules to zero survivors, is that an apparent edge is a confound until a tighter control has failed to remove it. This is the same rule pointed the other way: an apparent failure is also a confound until you have checked the instrument.

We published 0.96 because it is what our system currently returns. We declined to publish "the backtest overstated by 43%" because we have not earned that sentence yet.