Backtesting Arena

Backtesting Arena

Back to blog

What 100,000 Backtests Say Actually Works in Crypto — and 4 Questions That Expose a Lying Number

Four findings you can trade on — win rates, timeframes, regime filters, costs — and four questions that expose any performance claim. From 100,000 systematic backtests.

Backtesting Arena·September 7, 2026·6 min read·0 views
What 100,000 Backtests Say Actually Works in Crypto — and 4 Questions That Expose a Lying Number

Every day someone shows you a number. "77% win rate." "Beats the market three years running." "This setup caught the last two rallies." Whether you buy often comes down to whether you believe that one number.

On September 2, the Arena wrote backtest number 100,000 — systematic, across hundreds of coins, every timeframe, net of costs. A dataset this size can do two things no single number ever can: it shows which rules have repeatedly worked in crypto. And it shows exactly how a performance number deceives you — because at 100,000 runs, every one of those tricks happens somewhere on its own, and can be dissected with receipts.

You get both here. First what you can use. Then what protects you.

Four findings you can work with

1. You don't need to be "right most of the time" — stop looking for it. 73.8% of all comparable backtests (at least 5 trades, n = 71,129) have a win rate below 50% — and that's where nearly all the profitable systems live. Only 6.9% of runs clear 70%, usually on a handful of trades. This distribution has held at every one of our milestones, across changing markets and populations. In real data, profitable means: win less often, win bigger. Two takeaways: a strategy with a 40% hit rate is not broken — and anyone selling you a 77% win rate is almost certainly selling one of the four deceptions in part two.

2. The daily candle is the whipsaw trap — mid timeframes are the sweet spot. On daily candles, false signals shred the classics: RSI/SMA has a median CAGR of −2.9% on 1d versus +9.4% on 3-day candles; RSI Overbought/Oversold: 1d −8.9%, 3d +5.5%. "Away from daily" has held since our first milestone. Just as important and less intuitive: the monthly chart is not a safe haven — across all strategies, 1M is the weakest interval in the dataset (49.6% beat buy & hold, median −3.6%). Takeaway: test 2-day and 3-day candles first. And if you run multi-day candles, know which calendar day your grid starts on — that choice alone shifts results measurably; a dedicated piece on that is coming.

3. Market regime beats volatility — and every filter has one job. Among the pro filters, regime filters beat buy & hold most often: Bullmarket Stage 71.6% (n = 17,721), Altcoin Season 62.9%, the ATR volatility filter 61.5%. The ranking "regime before volatility" has reproduced since 40k. But judge each filter by its job: ATR is a drawdown tool — it exists to soften crashes, not to lift hit rates. Measure it by returns and you measure the wrong thing.

4. Turnover eats returns — invisible on average, lethal in particular. After costs, only 1.5% of gross winners flip — sounds harmless. But that 1.5% concentrates almost entirely in high-turnover strategies. Our once second-most-popular strategy (Stochastic RSI, averaging 558 trades per run) sits at 0 of 28 market cells above buy & hold after costs. We shut it off. Takeaway: run every strategy net of fees, and distrust anything that needs hundreds of trades — the average doesn't hurt, but you as an individual take the full hit.

Four questions that expose any number

The best property of 100,000 runs: every trick used to sell performance out there happens somewhere in here on its own — and can then be demonstrated with receipts.

Question 1: "What's in the denominator — and did it change?" Our own headline rate "X% beat buy & hold" ran 64% → 52% → 67% across milestones — without a single strategy getting better or worse. What moved was the asset mix: on crypto, ~65% of runs beat buy & hold; on equities in a decade-long bull market, ~16%. When the platform went crypto-only, the rate jumped. A rate is a property of the sample, not of the strategy. We still publish the wandering number — with exactly this explanation — because we measured it on ourselves, and because it teaches the question you should put to every rate a stranger shows you: which assets? which window? and what changed when the number changed?

Question 2: "Rate or magnitude?" The most vicious case in the dataset: on A2ZUSDC, all three comparable runs beat buy & hold. Their results: −93.7% to −99.1% CAGR. Buy & hold: −99.7%. Beating the benchmark and losing nearly everything — at the same time. It cuts the other way too: on SAHARAFDUSD, RSI/SMA turned a −35% buy & hold into −99.3%. "Beats the market" counts wins, not amounts. Always ask for both.

Question 3: "Are the same assets on both sides?" Compare unfiltered against filtered runs across the whole dataset and the unfiltered ones "win" 79.9% to 65.5%. Are filters useless? The resolution sits in the column such comparisons leave out: the average buy & hold of one group is −31.5%, of the other −6.6% — 25 percentage points of difference in the base populations. Where the benchmark is on the floor, "beating" it is free. What got measured was composition, not filters. Our software now refuses to compute cross-population comparisons entirely; robust filter effects live in the Edge Library, per market×strategy — same population on both sides, with bootstrap intervals.

Question 4: "How many independent cases?" This summer's entry in our curiosities list is genuinely named USELESSUSDT: +1,826% CAGR — from 6 trades. And the platform's reigning CAGR record (197,923.7%, against a 2,238.98% buy & hold) stands on 10 trades on a memecoin that is only in the data because it pumped. Both are worthless, and both look spectacular. Under 30 trades, it's an anecdote — the prettier the number, the more the rule matters. Also ask: how many setups were tried before this one remained? Three confirming runs feel like signal; when we later covered one such "multi-confirmed" candidate systematically with 847 runs, 21% still beat the benchmark. Washing out is the default.

What this means in the product

Verdicts here are never made at the overall aggregate — they are made per cell: minimum sample, net of costs, with multiple-testing correction. And they have consequences: since the last milestone, four of our own strategies have been shut off (the platform's weakest, a phantom edge propped up by dead altcoins, a replaced v1, and the high-turnover strategy above). All stay visible in the library with their numbers and verdicts. For you, curation means: you don't test on material we already know will deceive you.

The platform's patience record, by the way, now belongs to BTCUSDT with Keltner Breakout on daily candles: 9.0 years, 36.6% CAGR — with 30+ trades and no asterisk. That is what a number looks like when it survives all four questions.

See the current strategy×timeframe matrix for yourself →

FAQ

Are the 100,000 real user backtests? No: 97.6% are systematic coverage runs from our pipeline (hundreds of coins × strategies × timeframes, daily), 2,780 come from users. For statistical power that is a feature — systematic coverage carries no selection bias from whatever users happen to find exciting. We label it anyway, because unlabeled distributions read wrong.

Does "beats buy & hold" mean the strategy was profitable? No. On coins with a −99% buy & hold, almost anything beats the benchmark and still loses almost everything (question 2). Rate and magnitude are two different axes — and after costs, some winners flip on top of that.

Which timeframe should I test first? 2-day and 3-day candles. "Away from daily" holds for the classics (RSI/SMA median: 1d −2.9% vs. 3d +9.4%), and the monthly chart is — counter to intuition — the weakest interval in the dataset.

Why did your "beats B&H" rate wander over time? Asset mix, not performance (question 1): crypto-only instead of a mixed universe. The rate within each sub-population was fairly stable — what moved was the blend. Which is exactly why it appears here as a teaching case, not as a success metric.

Where do I find robust statements instead of aggregate rates? In the per-cell views: Strategy Insights (per strategy×timeframe against Avg B&H) and the Edge Library (filter effects with bootstrap intervals, same population on both sides).


Not investment advice, not a recommendation, not a forecast — historical patterns are no guarantee.

Study the Past — Improve your Future 🥋

Try it yourself

Run the backtest with your own parameters and time ranges.

Run backtest →
📬

Don't miss new blog posts

One short email per new post — strategies, backtests, market analysis. No spam, unsubscribe with one click anytime.

By subscribing you accept our privacy policy. We use Resend for delivery. Double opt-in confirmation required.

Comments (0)

Join free to post comments.

Sign up →

No comments yet. Be the first!