How Our Backtest Judge Caught Itself Out, and What We Fixed
A strategy judge is only worth something if its verdicts can be checked, including against itself. On 11 October 2026, while re-reading our own reports, we found a false sentence in the Lab's judge: the SPA test said a strategy "beats buy-and-hold significantly" when it was doing worse than buy-and-hold. Here is the defect, its cause, the fix, and the measurements that verify it.
This is research on past public data. Nothing here is investment advice.
What happened
That same day, the Lab was judging five classic trend-following strategies walk-forward over 2019-2025 (full results: five famous trend-following strategies judged across several crypto cycles). Four of them ended clearly below buy-and-hold: their mean per-bar difference was negative, with a t statistic between -2.16 and -2.58.
Yet each of these four reports displayed:
SPA p = 0.02: the setting frozen on training beats buy-and-hold significantly.
The displayed p-values ranged from 0.5% to 2% (2%, 1.1%, 0.5% and 0.82%). A strategy below the benchmark cannot beat it "significantly". The sentence was wrong.
The SPA test, in two sentences
The SPA test (Superior Predictive Ability, Hansen 2005) answers one question: does the best of the strategies tested really beat the benchmark, or is it the kind of gap chance produces when several rules are tried? It compares the observed statistic with thousands of resamples of the data (bootstrap); the p-value is the share of resamples that do at least as well as the observed one.
The cause
Hansen's test truncates the observed statistic at zero: T = max(0, t). If the strategy is below the benchmark, t is negative and T equals 0. The resampled statistics T* are truncated too, so always greater than or equal to 0. By definition, the share of T* greater than or equal to 0 is 1: the p-value must be 100%.
Our estimator counted T* strictly greater than T instead of greater than or equal. With T = 0, it only counted the rare draws where the strategy moved ahead of the benchmark, hence a tiny p-value and a "significant" sentence. A > where a ≥ belonged.
The defect affected every single-candidate SPA test in the Lab: single split, proof pooled across markets, fixed rule, portfolio and walk-forward. It also affected the multi-variant training SPA when all variants are below the benchmark.
What was fixed
- The estimator now counts T* ≥ T, as in Hansen's definition. A statistic truncated at zero gives p = 100%. When the observed statistic is positive, counting "≥" or ">" is the same: the test's ability to detect a real strategy does not change.
- Defensive reading of old reports: wherever a verdict or a sentence reads this p-value, the Lab checks the observed statistic. If it is zero or negative, the p-value read is 100%, even in a report written by an earlier version. An old report can no longer display "significant" below the benchmark.
- Old receipts remain verifiable: each receipt is re-verified with the estimator it declares. Fixing does not erase history.
- Risk against buy-and-hold is now reported fold by fold in walk-forward (maximum drawdown, Calmar, Ulcer, share of rallies and declines captured, time in position).
The fix is in version 245 of the Lab, currently being delivered. If you use an earlier version: a "significant" SPA sentence on a strategy that does worse than buy-and-hold should be ignored.
Measurements, before and after
A fix gets measured. We replayed the same computations, bit for bit, before and after.
| Check | Before | After |
|---|---|---|
| Synthetic series "strategy below buy-and-hold" (t from -11.8 to -54.2) | p = 0% | p = 100% |
| The 4 walk-forward reports of 11 October (same chain, same t) | p = 2%; 1.1%; 0.5%; 0.82% | p = 100% for all four, sentence "compatible with chance" |
| Pure noise, 3 scenarios × 100 runs: SPA ≤ 10% while the strategy is below the benchmark (single split) | 3; 9; 82 | 0; 0; 0 |
| Same, walk-forward | 8; 100; 88 | 0; 0; 0 |
| Pure noise: "survives" verdict | 0 everywhere | 0 everywhere (unchanged) |
| Power: strategies with a real injected edge, "survives" verdict (single split / walk-forward) | 6 / 8 and 54 / 81 out of 100 | identical: 6 / 8 and 54 / 81 |
Reading: on pure noise, the old computation displayed a "significant" SPA below the benchmark up to 82 and 88 times out of 100 in one scenario. After the fix: zero. And the Lab still detects strategies with a real edge just as often.
Why no false verdict came out
The Lab's final verdict never rests on a single test. To say "survives", it also needs an out-of-sample chain above buy-and-hold, a majority of folds won, a sufficient DSR and at least 30 independent events. Those conditions were already blocking: the four strategies stayed "inconclusive", before and after.
The path to a false "survives" did exist, though. A defensive strategy can end above buy-and-hold in compounded return, because it avoids the big drawdowns, while having a negative arithmetic mean difference per bar. In that case, the faulty SPA could have supplied the missing piece. That path is closed.
One open point remains, stated plainly: a second, fixed-block SPA test has the same form. It is only used for sealed training diagnostics, never for a verdict. Changing it would alter proofs that are already sealed; that decision has not been made yet.
What it changes if you backtest on TradingView
TradingView's Strategy Tester runs a strategy on the displayed history. It gives a profit, a drawdown and a win rate, but no held-out test period and no significance test. When you read a "significance" statistic anywhere, including from us, four reflexes:
- Check the sign before the p-value. A p-value does not say which way the gap goes. If the strategy does worse than buy-and-hold, no p-value makes it better.
- Compounded and average returns say different things. A rule that avoids crashes can end higher in capital while losing on average bar by bar. Look at both.
- Ask for the never-seen period. A result computed on the whole history, after tuning, mostly measures how well it fits the past.
- Require a judge to be tested on noise. A serious judge publishes how often it calls a strategy with no edge "significant", and how often it detects a real one. That table is what lets you trust it, or not.
To have your own Pine script judged on a period it has never seen, the Lab is free and computes the verdict on your PC. The Quick Check gives a first read without installing anything.
Key takeaways
- The Lab's judge wrongly displayed a "significant" SPA (p of 0.5% to 2%) for strategies below buy-and-hold: a
>where a≥belonged. - Fixed at the cause: p = 100% when the observed statistic is zero, including when old reports are re-read.
- Measured: false "significant" results on noise went from 82 and 88 out of 100 to 0, power unchanged, no verdict modified.
- An incorruptible judge is not a judge that is never wrong: it is a judge whose mistakes are visible, measured and fixed in public.
Further reading: backtest engine and Monte Carlo and our methodology.
Research on past public and simulated data. Nothing on this page is investment advice.
⚠️ Disclaimer — This article is for informational and educational purposes only. It does not constitute investment advice or a buy/sell recommendation. Past performance does not guarantee future results. Strategy Arena is an educational simulator with virtual capital. Always do your own research before making investment decisions.