What we tested — and what broke.

7 studies · 1,192 declared cells · every pre-registration frozen before its runner existed

Most trading education publishes the claim. Almost nobody publishes what happened when someone measured it. These are claims we took from public sources, wrote down in advance how we would test — including what result would make us abandon the idea — and then ran, publishing the answer whichever way it came out.

The hardest part is not the backtest. It is proving the test could have found an edge if one existed. Every study below is run against a placebo built from the same tape and against a planted, known edge; a study that cannot recover the plant is published as unproven, never as a negative. 5 of these 7 clear that bar.

We test methods as they are taught, not the people who teach them. Where a source's edge plausibly lives in judgement a mechanical test cannot reach, the study says so. Full citations are in the underlying reports.

The 9 EMA is special because everyone watches it

Refuted

Widely-viewed 2026 YouTube tutorial (~60,000 views at triage)

"The 5 EMA can produce too many false signals, while the 20 EMA may lag too much" — 9 is the balance point, and because traders, prop firms and algos all watch it, price reacts there.

Over 13 lengths × 2 average types × 3 timeframes × 3 instruments, 9 is the best length in 1 of 18 configurations — chance alone would give about 2. It is statistically indistinguishable from its own neighbours (7, 8, 10, 11) in 14 of those 18, and where it does separate, 2 of the 4 favour the neighbours.

A self-fulfilling level has a testable signature: a discontinuity at exactly 9. Lengths 8 and 10 smooth the same tape almost identically, so anything separating 9 from them is attention rather than arithmetic. The measured curve is smooth and monotone — longer is better, straight through the 20 the source calls "too laggy" — which is the signature of a smoothing parameter and of nothing else. A 9 EMA is also not lower-lag than a 9 SMA: both have their centre of mass at 4.0 bars by construction. It is worse than the plain SMA in 15 of 18 cells.

What this does not claim

The disposition behind the video — fewer indicators, trade with momentum, wait rather than chase — is not tested here and is not refuted. Neither is the claim that professionals use it; prevalence is not performance. What is refuted is everything specific to the number nine.

Readable negative. On the same pipeline, the same engine recovers a planted edge, and returns nothing on pure noise — so this is a measurement of the claim, not of our own blindness.

Measured on
29,858 trades at the source's own settings; gold, Nasdaq and EURUSD; 5 years; two independent vendors
Pre-registered
frozen at 9871a22e before any result existed
Multiple testing
710 declared cells, bar |t| ≥ 3.62

The 9 EMA / VWAP cross scalp pays 3 to 1

Not supported

A proprietary trading firm's published intraday tutorial, 2026

When the 9 EMA crosses above VWAP on a strong stock, you are entering right around VWAP — roughly a 60% hit rate at about 3:1 reward-to-risk.

The win rate broadly survives — 47.6% against a claimed 60%, comfortably above the 25% a driftless 3:1 barrier would give. The payoff does not: the realised reward-to-risk measures 1.03:1, not 3:1. Expectancy lands at −0.026R with a confidence interval straddling zero.

A lagging average cannot fill you where it crosses. By the time a 9 EMA crosses VWAP, price has already travelled — the median fill lands well above the crossing level, which widens the risk leg and shortens the reward leg in the same motion. The nominal 3:1 is measured from the level; the trade is taken from the fill. Note the trap in the obvious follow-up: reward-to-risk falls with entry lag as an arithmetic identity, so that relationship is not evidence about anything — and expectancy is roughly zero at every fill distance.

What this does not claim

This tests the mechanized rule. The prerequisite — "strong stocks pulling back to the 5/10 SMA" — has no mechanizable content: applying it discards 75% of signals and moves the mean by 0.0001R. The setup fires 0.67 times per name per day at the same rate on mega-caps as on momentum names, which means the discretionary selection is doing the work. That skill is real and this study cannot speak to it.

Readable negative. On the same pipeline, a planted edge is recovered at +0.16R, and pure noise returns −0.11R — so this is a measurement of the claim, not of our own blindness.

Measured on
673 trades across 69 symbols in the source's own time window; 60 days; one vendor
Pre-registered
frozen before any result existed
Multiple testing
264 declared cells, bar |t| ≥ 3.34

Buy the lower Bollinger band when RSI is oversold

Refuted

A popular retail trading channel, 2026

In a sideways market, price closing outside the 2σ band with RSI below 25 is a mean-reversion entry back toward the middle band.

Negative on 11 of 11 instruments and in every one of 23 declared cells — and, the sharpest number here, worse than its own drift-matched placebo. The RSI filter adds +0.029R while discarding 84% of trades, which is not a filter, it is a smaller sample.

The cost of the stop the source never states. Stopping at the signal bar's extreme is gross-positive (+0.031R) but pays 0.272R per trade in transaction cost; stopping at the opposite band costs almost nothing and has no gross edge at all (a 54.4% win rate against a 57.8% break-even). No stop distance reconciles the two. There is also a structural problem with the setup's own framing: a 2σ break IS price leaving its range, so "sideways" and "2σ break" compete rather than stack — across 13,989 triggers, exactly zero also satisfied the range, divergence and support/resistance conditions the source shows together.

What this does not claim

This is one exit geometry. The divergence-plus-support/resistance confluence the source also teaches was not testable at the required fidelity and is therefore not refuted. "Sideways" as our own house indicators define it turned out to be unreachable at a 2σ break bar — unreachable is not the same finding as ineffective.

Readable negative. On the same pipeline, planted effects of ±0.25R are recovered with the right sign and in order, and the noise arm is unbiased — so this is a measurement of the claim, not of our own blindness.

Measured on
11 FX pairs and gold; 2016–2025; 40.1 million minute bars; second vendor reproduces the negative on 10 of 11
Pre-registered
frozen at a1c62480 before any result existed
Multiple testing
31 declared cells, bar |t| ≥ 2.62

Enter when price returns to a three-candle zone

Refuted

A trading-education framework taught across several 2026 videos

A three-candle formation marks a zone; when price comes back and touches it, enter in the original direction with a fixed target.

−0.097R over 60,169 trades on 22 FX pairs, and the excess over its own placebo is −0.003 against a bar of +0.20. Out of sample it is positive on 2 of 22 instruments.

It is a cost story, and an unusually clean one: gross expectancy is genuinely positive at +0.025R, and transaction cost turns it into −0.097R. The rule finds something; it does not find enough to pay the spread it must cross to collect it. Five rounds of testing across timeframes did not change that, which is what closed the family.

What this does not claim

Closed for this entry and exit geometry, on these timeframes. Earlier rounds of this same family carried a vendor-comparison defect on daily bars that we retracted rather than quietly corrected — the daily-timeframe vendor agreements from rounds 1 and 2 are withdrawn and are not part of this verdict.

Readable negative. On the same pipeline, a planted edge is recovered at +0.61R on 42,762 fills, and the placebo is drift-matched to the same tape — so this is a measurement of the claim, not of our own blindness.

Measured on
60,169 trades; 22 FX pairs; 2021–2026; 1.5 million 30-minute bars
Pre-registered
frozen at bb929085 before any result existed
Multiple testing
136 declared cells, bar |t| ≥ 3.13

Wait for price to retest the supply or demand zone before entering

Refuted

Standard supply-and-demand curriculum, taught near-universally

A rally-base-rally or drop-base-drop leaves a zone; patience for the pullback into that zone gives you a better price on the continuation.

Every one of 8 pre-declared cells is net-negative on both vendors, none within 0.24R of the bar. But the finding that outlives the verdict is the mechanism: on planted trend worlds the identical zones and stops earn +0.55R entered at the zone's birth and −0.09R entered on the retest.

Waiting for the pullback is not a better price — it is a filter that discards the trades the pattern is about. 27% of zones never retest at all, and those are precisely the strongest continuations; the 73% that do come back are disproportionately the ones where continuation is already failing. The retest entry is structurally anti-selected against the very thing it claims to harvest. This is the most useful thing on this page, because it is a reason rather than a verdict.

What this does not claim

Tested at the 4-hour timeframe on FX and gold. The zone-birth entry that scored +0.55R was measured on planted trend worlds to isolate the mechanism — it is a diagnostic, not a strategy we are proposing, and it has never been gated on real bars.

Readable negative. On the same pipeline, the same zones and resolver recover +0.55R on planted trend regimes, so the pipeline is demonstrably able to see continuation when it exists — so this is a measurement of the claim, not of our own blindness.

Measured on
22 FX pairs and gold; 2023–2026; two vendors
Pre-registered
frozen at e79a397d before any result existed
Multiple testing
8 declared cells, bar |t| ≥ 2.23

Volatility compression predicts a big move within one to two weeks

Not supported

A 2026 YouTube indicator walkthrough

When the compression reading hits an extreme, a large move is coming within one to two weeks — direction unknown — so size up.

The timing content is empty: the median wait from an extreme reading to a big day is 4.5 days, against 6.9 days from any randomly chosen day (p = 0.25). "A big move within one to two weeks" is true essentially always.

The indicator is inverse volatility — median volatility divided by current volatility — which means it can only come down via the move it is said to predict. A demonstration where the line falls as the move arrives is not evidence; it is the construction. Across four arms there was no support for extreme compression producing larger forward moves, and the dose-response ran toward persistence rather than squeeze-transition: the largest forward moves sat in the LOWEST readings on 3 of 4 arms.

What this does not claim

This is UNPROVEN rather than refuted, and the distinction is the honest one: the power check recovered a planted effect only at 1.5× scaling, so effects smaller than that are invisible to this test. One arm reads significantly negative on the shortest panel — do not read that as a tradeable inverse. Separately, generic volatility-targeted position sizing, which the source does not advertise, was positive in 7 of 8 cells and remains open.

Underpowered. We could not establish that this test would have seen an effect of the size claimed, so this is recorded as unproven rather than as a negative. The difference matters and we will not collapse it.

Measured on
BTC and ETH; two vendors each; daily bars
Pre-registered
frozen at 4e7d8b8c before any result existed
Multiple testing
31 declared cells, bar |t| ≥ 2.62

Stand aside before high-impact news

Half true

Near-universal retail risk advice

Cancel or suspend pending orders ahead of CPI, NFP, FOMC and PCE — the release will whipsaw you out.

The premise checks out and the conclusion does not. Volatility really is lower before a release — gold at 0.90× and EURUSD at 0.88× their normal range, both intervals excluding 1.0. But orders that filled into a release averaged slightly BETTER, not worse.

The quiet hour before a release is real, and it is why the advice feels right. What it does not establish is that the release itself is where pending orders lose money. Once the actual fills are measured rather than the volatility around them, standing aside removes trades that were, on average, marginally profitable — so the blackout costs opportunity and buys nothing measurable.

What this does not claim

Measured on pending entry orders on a 4-hour grid, on gold and EURUSD, across CPI, NFP, FOMC and PCE. It says nothing about holding large positions through a release, about slippage on market orders at the print, or about instruments outside this pair.

Underpowered. We could not establish that this test would have seen an effect of the size claimed, so this is recorded as unproven rather than as a negative. The difference matters and we will not collapse it.

Measured on
1,376 fills; 2013–2026; 2,800 scheduled releases
Pre-registered
frozen before any result existed
Multiple testing
12 declared cells, bar |t| ≥ 2.23

Why we publish the negatives

A strategy that survives testing is worth money, so nobody shows you the ones that did not. That makes the published record of any trading educator a survivor list, and a survivor list tells you nothing about the odds.

These studies cost the same to run whichever way they come out, and we register every cell we look at — including the abandoned branches — so the statistical bar rises each time we retest something. That is the mechanism that stops a research programme from eventually finding whatever it went looking for.

Our own systems are held to the same standard, and most of them do not clear it either. The forward record shows what is actually running and what it has done since it started.

These were run on our own initiative. You can commission one on a claim you are weighing up — the verdict is reported as it lands, including when it finds nothing wrong.

Follow the daily outlook: RSS · today's brief