The Multiple-Testing Lottery: What the Deflated Sharpe Ratio Actually Tells You
The 2,255-strategy tournament tested so many combinations that it needed a bar that rose with the size of the search, not a fixed one. The lone survivor missed it by -0.525, and the decisive proof came after: it went from beating buy-and-hold in 3 of 4 assets in validation to 0 of 4 in test.
Buy enough lottery tickets and one will win. That says nothing about that particular ticket's quality. It says something about how many were bought. The trading strategy tournament covered in this series tested 2,255 distinct variants, and before trusting the best result it had to answer an uncomfortable question: is that result better than the best ticket pure chance would produce, testing the same number of times? This piece explains how you answer that question, no formula required to follow it.
Why testing more always finds something that shines
A bar that scales with the search
The experiment computed that threshold with a standard quant-finance formula (the maximum expected value under pure chance, from Bailey and López de Prado), which depends on how many variants were tested and how much the result varies across them. Across the tournament's 2,250 individual variants, using only the 1,750 with a defined result across all four asset categories (the other 500 were excluded for lack of sufficient data, never substituted with a zero), the bar landed at a mean Sharpe of 0.9718. Across all 2,255 (adding the 5 blended combinations), essentially the same number: 0.9720.
The number: 0.4467 against a 0.9718 bar
The one strategy that survived the robustness gate arrived here with a mean Sharpe of 0.4467, a margin of -0.525 against the bar, not even half of what was required. The best blended combination (a portfolio of the two most promising strategies) got closer, 0.549, but still fell short, with a margin of -0.423. Not one of the 2,255 variants, individual or blended, cleared its own bar.
The out-of-sample confirmation
The deflated Sharpe number alone would already have been enough not to declare victory, but the tournament had one more check, and it turned out to be the most decisive of all. The same strategy that beat buy-and-hold in 3 of 4 assets during validation, the partition where the entire search happened, dropped to 0 of 4 in test, the partition nobody had touched before that single shot. That's exactly the behavior predicted by a result indistinguishable from the best ticket in a lottery: it shines where you searched, and vanishes where you hadn't searched yet.
Monte Carlo and bootstrap: two more ways not to fool yourself
The tournament ran two more checks on the test partition, each from a different angle. The Monte Carlo simulation compares each champion's real Sharpe against what a strategy with no real skill would produce on the same history: none of the 6 cleared the 95th percentile of that null distribution, not even the closest one (0.614 against a 0.821 threshold). And bootstrap confidence intervals on the Sharpe ratio, computed asset by asset, cross zero in nearly every case, except, with a real nuance, in the aggregate cross-asset average of the two best candidates, where the interval stays above zero without quite crossing it. Even that exception wasn't enough: between the deflated Sharpe bar, the test-set drop, and the Monte Carlo verdict, no candidate cleared all three at once.
What this means for any result you see online
The next time a backtest boasts an impressive number, the question that actually matters isn't "is this a good number?" It's "how many variants were tested before this one got shown to me?" Almost no backtest circulating online answers that question, and many don't even track it. A strategy screenshot with one Sharpe ratio and one equity curve tells you nothing about how many other parameter combinations were quietly discarded to get there, and there's no way to reconstruct that number after the fact if the author didn't log it in advance.
That's the actual reason a fixed, pre-registered bar matters more than the raw number itself. A Sharpe of 0.4467 would read as a solid result in most backtest write-ups circulating online, with no context about how many attempts it took to find it. Against a search of 2,255 variants, the same number is well short of what pure chance alone would be expected to produce. The number never changed; what changed is whether you know the size of the haystack it came from. The tournament fixed it in writing before computing anything, applied it to its own best candidate with the same rigor, and published the negative result with the same detail it would have given a positive one, which is, at bottom, the reason this series exists.
Nothing in this article is investment advice, and nothing on NodeWitness is. You can read the tournament's full verdict, why the short side loses even while collecting funding, or follow NodeWitness's live Cycle Score, same discipline of publishing what doesn't work too.
Last updated: August 31, 2026