Backtesting Arena

Backtesting Arena

Back to blog

When a Backtest Result Is Just Noise: Four of Our Own Studies That Failed

When is a backtest result a false positive? Four of our own studies failed — against four rules we locked before looking. What each one caught.

Backtesting Arena·August 12, 2026·11 min read·0 views
When a Backtest Result Is Just Noise: Four of Our Own Studies That Failed

The most common question asked of any backtest is whether the result is real or a backtest false positive. Over the past weeks we abandoned four of our own studies. Not because we lost interest — but because in each case a rule fired that had been fixed beforehand. Four studies, four different rules. That is the point of this piece: not the four failures, but the four rules.

One at a time. Each rule gets the case that triggered it.

Rule 1: Race the strategy against a random trader that trades just as often

We wanted to know whether rotating between an altcoin, Bitcoin and dollars leaves you holding more Bitcoin than someone who simply does nothing. The idea is widespread and sounds reasonable: hold the altcoin during altcoin season, Bitcoin otherwise, step aside in a crash.

We measured across nine pairs and five signal-interval combinations — 45 cells — from September 2017 to June 2026. We only evaluate the 30 cells that accumulate at least 30 trades; below 30 we call it an anecdote, not a result. The benchmark is the simplest one available: hold Bitcoin and do nothing. That is 1.00.

The result is a median of 0.86. Anyone running the rotation ends up with roughly 14 percent less Bitcoin than someone who did nothing at all for nine years. Before costs the figure would be 1.36 — fees eat the difference. At a median of 224 switches that is not a side effect, it is the mechanism: the gap between gross and net runs at a median of 20 percent per cell and reaches 64 percent at worst.

One part does hold up: against holding the altcoin itself, the rotation wins clearly, 0.86 against 0.23. But the accurate sentence is uncomfortable — it protects you from the altcoin, less well than never buying it.

Now the rule. In 14 of 30 cells the rotation beats simply holding. Nearly half. You could work with that. Which is why every study runs alongside a placebo: same trading, same number of switches, same fees, same time spent in each state — only the timing is shuffled at random. Two hundred such random runs per cell.

The outcome: at the median, 12 percent of those random runs match what our rule achieves. What got measured was not the rule but simply how long you were exposed to the altcoin. The timing contributed nothing.

And a trap that nearly closed. One row looked excellent: ema_cross on daily bars, median 3.09 — triple the do-nothing result — six of nine pairs above holding, and a placebo value of 0.04. By any conventional reading that is significant. Had we shown that row, it would have become a product page.

Against this runs the final check, the Deflated Sharpe Ratio. It does not ask "is this number good?" but "how many variants did you try before you found it?". We tried 45. Anyone testing 45 combinations and reporting the best will almost always find something. Of the 30 evaluated cells, not one passes this check — including the pretty one.

The full measurement is here: Stacking More Sats With Altcoin Rotation? 45 Tests, One Answer.

Rule 2: Write the rules down before you see the numbers

The Wyckoff Spring is a well-known chart pattern: price drops below a support level, it looks like a downside break, and then it immediately turns back up. The reading is that large players are collecting retail stop orders before the move higher.

The problem with such patterns is not that they do not exist. The problem is that their definition is soft. How far below support? How fast must the return be? What counts as support in the first place? Anyone answering those questions after first looking at the results will always find a combination that works.

So we published our answers beforehand — as a separate piece three days ahead of the result. It contains every threshold, the definition of bull and bear, how days hit multiple times are counted, and the cost assumption. That is the point: the piece is dated and predates the first number.

Then we ran it. Ten pairs, no data gaps, 395 raw events, 332 after de-duplication.

The only effect that rises out of the noise at all points in the wrong direction: five days after a Spring, returns sit about 1.2 percentage points below the general uptrend, and that range lies entirely below zero. Over twenty days it fades to nothing.

The decisive comparison is a different one. We repeated the same measurement with a cheap substitute signal — a plain oscillator, no chart pattern, 312 events. The uncertainty ranges of Spring and substitute overlap at all three horizons. Plainly: the elaborate structural detection delivers nothing a trivial indicator does not.

A third, stricter variant produced 20 events. Twenty. We say nothing about those, regardless of how the result looks.

The full result with all figures.

Rule 3: A cell with three observations does not get computed

Of the four studies this is the only one without its own piece, so here it is in full.

We wanted to know two things. First: does a quiet phase in the Bitcoin price precede a larger move? Second: if so, in which direction — and does that depend on whether we are in a bull or bear market?

The basis was 5,849 trading days from July 2010 to July 2026. A day counted as "quiet" if the last 30 days' movement sat below the average of the last two years — strictly from that day's point of view, with no look into the future. That produced 49 episodes, 47 of them with a complete observation window.

Question one is answered, and the answer is no. At 30, 90 and 180 days after a quiet phase began, movement sat at essentially the same level as at the start. Roughly half the episodes went one way, half the other. Quiet phases are not systematically followed by expansion.

What makes the case interesting is the prelude. Before the study ran, we had eyeballed five examples. In five out of five, calm was followed by a large move. A clear pattern, we thought. Across 47 episodes nothing of it survived. Five cases will nearly always show a pattern — that is not a statement about the market, it is a statement about the number five.

Question two we could not answer, and that is the actual finding. Splitting direction by market phase requires quiet phases in bull markets and in bear markets. We had 34 above the long-term average and 3 below.

Three. Our rule demands 30 per cell, so nothing was computed.

The obvious objection is that our conditions were too strict. We checked: loosening the spacing rule far enough to yield 94 episodes instead of 47 moves the ratio to 69 against 7. There is no setting under which the question becomes answerable — and the reason is not bad luck but the market itself. Quiet phases below the long-term average barely occur. Bear markets are loud.

That is a better result than a number. "Not answerable, and here is the structural reason" can be checked. A number derived from three observations cannot.

Rule 4: First measure how much your own measurement choices wobble

The fourth case differs from the other three, which is why it comes last. Here no study failed. Here the question dissolved.

The starting point: on three-day candles, five strategies produced the best figures. The question was obvious — what is special about three days? Before building a product on it, we wanted to settle one thing.

Three-day candles do not occur in the market. We assemble them from daily candles by grouping three days into one bucket. The only question is where the first bucket starts. In our case it starts where it would have started on 1 January 1970 — the zero point of computer timekeeping. That is a deliberate choice with a good reason: every pair gets identical boundaries. But it is one of three possible alignments, and neither the market nor the method picked it.

So we computed all three. Result: the sign does not flip in a single case — the effect is real, all five strategies stay clearly positive under all three alignments. But the size moves by 6.66 percentage points on average, and for one strategy by 12.30 on a base of roughly 50. Nearly a quarter of the headline depends on where we opened the first bucket. The number of trades stays almost unchanged — the same trades at different prices.

And with that we had, for the first time, a measure of how large differences become purely from arbitrary measurement choices. It settled the original question immediately: the gap between two-day and three-day candles is 4.40 percentage points. It sits below our own noise. There was nothing to explain.

The same measurement series produced one more lesson that has stayed with us. Five strategies showed an 18 to 24 percentage point advantage over simple holding across every view. Split into Bitcoin and the rest it looked different: for Bitcoin, where holding was strongly positive, the same strategies sat 10 to 24 points below it. For the remaining pairs, where holding lost 66 percent a year, the strategy earned approximately nothing — and was therefore 13 points ahead.

The entire advantage was the badness of the altcoins. An advantage over a benchmark is only readable when that benchmark's own value sits next to it. If it reads minus 54 percent, then "plus 23 points" does not mean "good strategy", it means "do not buy this".

What the four cases share

A small sample will nearly always show a pattern. Five hand-picked volatility examples: expansion five times. One of 45 rotation cells: triple the return with a clean-looking randomisation test. Twenty Wyckoff events in the strictest variant: some number will certainly appear.

All three look like findings. None of them is one.

That is why the rule comes before the result each time. Not because we want to be strict, but because a rule set up after looking at the numbers is not a rule, it is a justification.

What we are not claiming

Four points that should not get lost.

Our definition of "quiet" was soft. A day counted as quiet if movement sat below the two-year average — which applies to 51.5 percent of all days. Every second one. The sharper question, whether anything follows unusually quiet phases, we did not answer. It is open, and it would be a new study, not a threshold moved after the fact.

A null result is not proof of absence. It says: with this definition, on this data, in this period, nothing was visible. A different definition may find something else. Which is why the definitions are stated openly.

The rotation is not refuted, it is untested. What we measured is one family of rules on nine pairs. A different signal may come out differently — it would just have to pass the same randomisation test.

We do not claim these four rules are sufficient. They are the ones that have fired so far.

What Backtesting Arena does with this

The four rules are not good intentions here, they are built in. Below 30 trades a result is flagged as an anecdote, not a result. Every return figure gets its benchmark from the same window placed beside it — including when that is inconvenient. And where a question is structurally unanswerable, that is what appears instead of a number.

What that looks like in the product is described in Why Honest Backtesting Looks Different. How we handled a benchmark that blew up in our own face is in Backtest baseline: how our comparison value broke.

Frequently asked questions

What is a placebo test for a trading strategy? You race a random trader against the strategy — one that trades just as often, stays invested just as long and pays the same fees, with only the timing shuffled. If the strategy does not clearly beat that random trader, you measured market exposure, not the rule.

How many trades does a backtest need at minimum? We draw the line at 30 and treat anything below as an anecdote. That is a convention, not a magic number — but it has to be fixed before the measurement, otherwise it becomes a tool for defining inconvenient results away.

What is the Deflated Sharpe Ratio? A correction for how many variants you tried. Test 45 combinations and you will almost certainly find a good-looking one. The DSR discounts for the number of attempts required, and it leaves none of our 45 rotation variants standing.

Why publish failed studies? Because otherwise the selection distorts the result. Showing only what worked produces in readers exactly the reasoning error you are trying to avoid in yourself. And because the rule something fails against is often worth more than the finding would have been.

Does "no effect found" mean there is none? No. It means none was visible with this definition on this data. That is why the definitions are in the text — so someone with a better definition can recompute.

Why is a benchmark mandatory here? Because a return is not readable without one. Plus 23 percentage points against a benchmark sounds good — but if that benchmark stands at minus 54 percent, the statement is not "good strategy", it is "do not buy this asset".


Sources: our own measurements. Rotation study (45 cells, Sept 2017 – June 2026), Wyckoff Spring evaluation (10 pairs, 332 de-duplicated events), volatility episode protocol (5,849 trading days, July 2010 – July 2026), grid phase measurement (BTCUSDT, 1,089 three-day candles, 2019–2026). The measurement protocols underlie the pieces linked above.

This piece describes measurement methods and their results. It is not investment advice and not a recommendation for or against any particular instrument or approach.

Try it yourself

Run the backtest with your own parameters and time ranges.

Run backtest →
📬

Don't miss new blog posts

One short email per new post — strategies, backtests, market analysis. No spam, unsubscribe with one click anytime.

By subscribing you accept our privacy policy. We use Resend for delivery. Double opt-in confirmation required.

Comments (0)

Join free to post comments.

Sign up →

No comments yet. Be the first!