Backtesting Arena

Backtesting Arena

Back to blog

The Wyckoff Spring, Backtested: We Locked the Rules First — Here's the Result

We pre-registered a Wyckoff spring backtest: rules locked first, 10 pairs, placebo control. The result is a clean nothing — and short-term even negative.

Backtesting Arena·July 28, 2026·7 min read·0 views
The Wyckoff Spring, Backtested: We Locked the Rules First — Here's the Result

The Wyckoff spring backtest we're presenting here has a property most pattern studies lack: the rules were fixed and public before the first number was computed. No adjusting the pivot length after the result disappointed. That's what makes the answer to "does the Wyckoff spring work?" trustworthy — and the answer is: on this data, with this rule, no.

The framing that belongs in every version of this result: we tested one mechanical implementation of one Wyckoff event, on crypto, on daily bars. This does not refute Wyckoff's theory. It refutes the claim that this one clearly defined pattern delivers a measurable edge.

What the spring is — and why we tested it specifically

Wyckoff's schematic describes how an accumulation phase unfolds: a trading range where large players build inventory before price breaks out. The spring is the most dramatic event inside it: price briefly drops below support — shaking out the last sellers — and reclaims it by the close. The textbook story: a bear trap right before the turn.

We deliberately tested the spring rather than "the whole schematic." Phases are states; they emerge from a path through many thresholds and can't be counted. The spring, by contrast, has a timestamp and a condition — it's countable, and therefore testable. Testability here is a matter of how you cut the question.

There's also a theoretical headwind documented before the test. Osler (2005) shows for FX markets that prices accelerate after breaking through stop-loss clusters — they don't reverse. That's exactly where the spring lives, below support. The documented order structure predicts the opposite of what the spring claims for that zone.

How we tested (and what was fixed in advance)

We published the full pre-registration — every locked number, every interpretation rule — before running the math. The short version:

  • Universe: BTC, ETH, MATIC, BNB, TRX, XRP, EOS, LTC, BCH, VET (each against USDT) — selected by trading volume as of the start of 2020, not by today's size. That's the point: MATIC ranked 3rd back then, VET 10th. Selecting by today's list bakes in survivorship bias.
  • Window: listing start to 2026-06-30, daily bars. No data gaps.
  • The rule, mechanical: a confirmed pivot low (6 bars either side) that sits at the bottom of a 20-bar range; a spring fires when price undercuts that level and reclaims it at the close. Entry at the closing price — not the low, which wasn't tradable.
  • Costs: 0.20% round-trip (0.10% per side), subtracted.
  • Three volume variants: V0 (no filter), V1 (high volume = absorption), V2 (low volume = no supply). The literature is split here, so we test both readings.
  • A placebo as control: a deliberately trivial "spring" defined as RSI(14) crossing 30 from below. No price level, no structure — just a bounce out of oversold. It answers the question no other test asks: does the Wyckoff structure contribute anything a trivial rule doesn't?

We measure the difference from the unconditional base rate: the average price move after a spring, minus the average move across all bars in the same window. A positive return after the spring means nothing if the market was rising anyway.

Fixed before the first number: anything below one standard error counts as no detectable effect — regardless of sign. And all twelve tests (four rules × three horizons) feed a multiple-testing correction (Deflated Sharpe), so "the best of twelve numbers" doesn't pass as a finding.

The result

First the pure price structure with no volume filter — the spring itself:

HorizonEventsDifference vs. base rate≥ 1 std. error?Bootstrap interval
5 days332−1.18%yes[−2.26%, −0.10%]
10 days325−1.06%yes[−2.56%, +0.44%]
20 days325−0.86%nocrosses zero

The only signal that separates from zero is negative: over five days, price after a spring does about a full percentage point worse than after an arbitrary bar, and the bootstrap interval sits entirely below zero. By 20 days the effect fades to nothing. That's the direction Osler's finding predicts — the break tends to continue rather than reverse.

The volume variants don't change the picture:

RuleHorizonEventsDifferenceAssessment
V1 (high volume)20 days197+2.09%interval [−0.72%, +4.84%] crosses zero → not robust
V2 (low volume)all20under 30 events = anecdote, not a result
Placebo (RSI × 30)all312~ 0no effect

V1 has a single positive outlier over 20 days — but the confidence interval reaches well into negative territory; the effect isn't robust. V2 falls below our hard line: fewer than 30 events is an anecdote, not evidence. And the Deflated Sharpe fails across all twelve tests — after correcting for the number of attempts, nothing survives.

The decisive comparison: spring vs. placebo

This is where it gets interesting. If the Wyckoff structure contributes something the trivial RSI cross doesn't, the two should differ. They don't:

HorizonSpring (difference)Placebo (difference)Intervals
5 days−1.18%−0.56%overlap
10 days−1.06%−0.59%overlap
20 days−0.86%+0.53%overlap

At all three horizons the confidence intervals overlap. The elaborate price structure — pivot confirmation, range edge, undercut with reclaim — contributes nothing beyond a simple oscillator cross. And what it does contribute short-term points the wrong way.

The sober takeaway

On this data, with this rule, the Wyckoff spring has no detectable positive edge. The only statistically solid signal is a small negative one over short horizons — consistent with support breaks continuing rather than reversing. After costs and multiple-testing correction, no edge survives in any of the twelve tests.

What this does not mean: that Wyckoff analysis as a whole is worthless. A human with context reads a trading range differently than a mechanical rule. The spring is an event within a larger context no line of code captures. We tested an operationalization, not the idea.

What it does mean: anyone selling the spring as a standalone buy signal — "price undercuts support, reclaims it, get in now" — is selling a pattern that shows no edge on a broad, survivorship-clean crypto dataset and doesn't beat a trivial RSI trigger.

Limits of the test

  • It's one implementation. A different pivot length, a different range window, different volume thresholds would produce different springs — and possibly different numbers. We fixed ours in advance and didn't vary them; the result holds for exactly that choice.
  • Crypto only, daily bars only.
  • The placebo is one trivial alternative explanation, not all of them.
  • Osler's findings come from 1990s FX on minute bars. That order structure behaves the same for crypto daily springs is plausible but not proven here.

What Backtesting Arena contributes here

The appeal of this test isn't the result — it's the order: rules first, numbers second, publication regardless of outcome. That discipline — avoiding look-ahead bias, measuring against the right base rate, correcting for multiple testing, running a control group — is exactly what most self-built backtests miss. If you want to understand why those four steps decide whether a supposed edge is real, the systematic case is in our primer on honest backtesting and the look-ahead problem. The full locked rules of this study are in the Wyckoff spring pre-registration.

FAQ

Does the Wyckoff spring work as a trading signal? In this pre-registered test on ten crypto pairs, the mechanical spring rule shows no detectable positive edge. The only statistically solid result is a small negative effect over five days.

Why is the result negative, not just "zero"? Over short horizons, price after a spring performs measurably worse than after an average bar. That fits research showing prices tend to continue rather than reverse after breaking through stop clusters. Over longer horizons even that effect disappears.

Does this mean Wyckoff analysis is worthless? No. We tested a mechanical implementation of a single event. Discretionary Wyckoff analysis with context is a different thing and wasn't tested here. A negative result refutes the rule, not the theory.

What is the placebo control and why does it matter? A deliberately trivial "spring": RSI(14) crossing 30 from below. It measures pure oversold bounce with no price structure. Because the real spring doesn't beat this placebo, the elaborate structure contributes nothing the simple trigger doesn't.

Why were the rules published in advance? Because you can always make a pattern fit after the fact. If the rule is only fixed after the first result, you're really testing which rule gives the nicest number. Pre-committing rules out that — and the pledge to publish regardless of outcome is what makes pre-registration worth anything.

Why these ten pairs? Selected by trading volume as of the start of 2020, not by today's size. That prevents survivorship bias: picking today's biggest coins implicitly selects the past's winners.

Try it yourself

Run the backtest with your own parameters and time ranges.

Run backtest →
📬

Don't miss new blog posts

One short email per new post — strategies, backtests, market analysis. No spam, unsubscribe with one click anytime.

By subscribing you accept our privacy policy. We use Resend for delivery. Double opt-in confirmation required.

Comments (0)

Join free to post comments.

Sign up →

No comments yet. Be the first!