You test a strategy on daily candles, then on 2-day candles, then on 3-day candles. The return rises with every step, and the conclusion suggests itself: the coarser grid smooths out the noise, so take the 3-day candles. The conclusion is wrong, and not because of the strategy, but because of a parameter you never set.
Every strategy description lists its parameters: RSI period 14, moving average 50, stop at 8 percent. What no description lists is the parameter the engine chose for you before you started.
A 3-day grid needs a start day
If an indicator runs on 2-day or 3-day candles, something has to decide which calendar day each candle begins on. There is no natural answer to that. A 3-day grid can start on a Monday, a Tuesday or a Wednesday, and each of the three produces a different chart, different crossovers and different trades on different days.
Most engines, the one on this platform included, anchor the grid to 1 January 1970, the zero point of Unix time. That is a convention, not a finding. One of the possible alignments gets used and the others never appear.
Measured across this platform's own corpus, the choice of alignment alone moves annual return by 6.66 percentage points on average and by 12.30 points at the extreme. Name an edge you would actually put money behind that is bigger than twelve points of annual return.
That is the size of a parameter the strategy description does not list.
What it does to a real comparison
Take the plainest momentum rule there is, an RSI crossing its own moving average, on Bitcoin from the start of 2017 to today. The identical request three times, changing nothing but the candle length.
| Candles | Annual return, gross | Net of fees | Trade events |
|---|---|---|---|
| Daily | 36.1 % | 28.3 % | 530 |
| 2-day | 42.3 % | 38.3 % | 254 |
| 3-day | 47.0 % | 44.6 % | 146 |
A clean ladder. Net of fees it gets steeper still, because the slower grids trade far less. Every instinct says move to 3-day candles.
Every instinct is reading noise. The gap between the ends is 10.9 points. That is above the average effect of alignment alone and below its measured maximum. Inside that band you are not looking at a ranking of intervals but at the Unix offset your data happened to land on.
The second thing the ladder hides
The three runs did not trade the same period. An indicator needs a warmup before it can fire its first signal, and warmup is counted in candles, not days. Fourteen daily candles is two weeks. Fourteen 3-day candles is six.
So the identical request produced three different starting lines. The daily run began trading on 13 September 2017, the 2-day run on 11 October, the 3-day run on 8 November. Fifty-six days apart, from one request, without a hint.
The benchmark moves with it, and this is where the apparent edge comes from. Measured over the window each run actually traded, buy and hold returned 39.88 percent a year for the daily run, 35.27 for the 2-day run and 32.97 for the 3-day run. The autumn of 2017 was the steepest part of the whole sample, and the slower grids sat it out.
| Candles | Trading from | Buy and hold, same window | Strategy minus buy and hold |
|---|---|---|---|
| Daily | 13 Sep 2017 | 39.88 % | −1.65 points |
| 2-day | 11 Oct 2017 | 35.27 % | +3.43 points |
| 3-day | 8 Nov 2017 | 32.97 % | +6.31 points |
The 3-day run beats a yardstick almost seven points lower than the one the daily run had to clear. Still a ladder, but one measured with three different rulers, and that is not a comparison.
Where the evidence thins
There are 530 trade events on daily candles, 254 on 2-day and 146 on 3-day, of which 72 are closed round-trips. The interval that looks best rests on the fewest independent cases. That is the shape of the problem: coarser candles mean fewer signals, fewer signals mean more of the result comes from a handful of moves, and a handful of moves is where luck lives.
None of this makes the 3-day chart wrong. It makes the comparison wrong. Three candle lengths, three alignments you did not pick, three warmup lengths, three windows and three benchmarks, presented as one ranking with a winner.
The objection
"Every engine has a default. If everyone uses the Unix alignment, at least the comparison is fair."
Fair between two people using the same engine, yes. Not between the daily run and the 3-day run inside the same request, because the daily run has no alignment problem at all, only the 3-day run does. And the market never asked which day your grid started on.
Three questions for every interval comparison
When someone shows you a result on 2-day or 3-day candles, ask three questions. Which alignment did the grid use, and did they check the others? Over which window, and does the benchmark cover the same window? How many trades?
If the gap between intervals is under roughly twelve points of annual return, treat daily, 2-day and 3-day as one block rather than a ranking, because that is all the resolution the data has.
The wider lesson costs nothing. Any parameter that was chosen for you is still a parameter, and it still needs a sensitivity check. Defaults are decisions. They are just decisions someone else made, quietly, and never wrote down.
How this was measured
The 6.66 and 12.30 points come from the parameter documentation of the Backtesting Arena connector (field interval), measured across the platform's own corpus. The triple run: strategy rsi_sma, pair BTCUSDT, 1 January 2017 to 6 September 2026, starting capital 10,000, no filters, default parameters. The runs were repeated on 14 September 2026 and returned the same values as on 7 September. The engine flags the differing trading windows itself (benchmark.matches_strategy_window: false) and returns buy and hold over the actual window as a separate field.
Next measurement: the indicator that marked the last two big rallies. The two hits are easy to show. The number that decides whether it is worth anything is the false alarm rate.
Not investment advice, not a recommendation, not a forecast — historical patterns are no guarantee.
Study the Past — Improve your Future 🥋