‹ Patterns & Signals Lesson 15 of 16
Contents Lesson 15 of 16

5 min read · professional

How would you test whether a pattern has any edge?

The way out of the trap the last lesson described is counting, with the rules fixed in advance. This lesson is the recipe. It's also the single most transferable skill in this domain — the same procedure evaluates an indicator, a seasonal claim, or anything else somebody tells you "works."

Step 1 — Define the pattern mechanically

If you can't write the rule precisely enough that a computer would identify the same instances twice, you don't have a pattern, you have an impression. "A flag" is not a definition. This is:

A move of at least 15% within 10 sessions, followed by 5 to 15 sessions in which the total high-to-low range is under 40% of the move and no close falls below the move's midpoint.

Every number there is arguable. That's fine — what matters is that they're fixed before you look, so the data can't talk you into adjusting them.

Step 2 — Define the outcome and horizon in advance

"Goes up" is not testable. "The close 10 sessions later is higher than the close on the pattern's final day" is. Pick the horizon first. Testing five horizons and reporting the best one is a different, much weaker exercise — see step 6.

Step 3 — Count every occurrence

All of them. Including the ones that look scruffy, the ones during crashes, the ones you'd have talked yourself out of. This step kills most pattern claims, because most claims are built on a remembered subset rather than a counted population.

Step 4 — Compare against the base rate

This is the step people skip, and it invalidates more results than everything else combined. If your rule is followed by a higher close 56% of the time, you have learned nothing until you know how often that stock closed higher over the same horizon unconditionally. If the answer is 54%, your pattern contributed two percentage points — and the next question is whether two points can be told apart from noise at your sample size.

A worked illustration, rounded: your engulfing rule fires 180 times over 15 years. In 97 cases the close 10 sessions later is higher — 53.9%. The unconditional base rate over the same period is 53.1%. Difference: 0.8 points.

Now size the noise. The standard error of a proportion at n = 180 is roughly √(0.5 × 0.5 / 180) ≈ 3.7 percentage points. Your 0.8-point difference is a fifth of one standard error. It is indistinguishable from nothing, and no amount of staring at the charts will change that.

Step 5 — Split the sample

Build and tune the definition on one period (say 2005–2015). Then run it, untouched, on a period you never looked at (2016–2025). A rule tuned on all the data is fitted to all the data, and flatters itself accordingly. Out-of-sample is where honest results survive and fitted ones die.

Step 6 — Count your attempts, not just your result

If you tested forty variations — different thresholds, horizons, filters — and one produced a striking result, that's roughly what pure chance delivers. This is the multiple comparisons problem, and it is the main reason published market anomalies shrink or vanish after publication. Report how many things you tried. To yourself, at minimum.

Step 7 — Subtract costs

Spread, commission, slippage. Many candidate effects are smaller than the bid-ask spread on the instruments where they appear — which means they exist on the chart and not in reality.

What you'll probably find

Careful researchers who follow this procedure tend to report small, unstable effects that often weaken out of sample or after costs — with a few studies finding non-zero information content in automated pattern definitions, and others finding none. That's what "the evidence is mixed" actually means: not a conspiracy, not a secret, just weak signals in noisy data.

Running the count yourself is the only way to know what you genuinely believe rather than what you've absorbed.

In the data

Counting every occurrence depends on counting the right things. Across more than one stock, the trap is that tickers change. Here are two weeks of renames on US exchanges:

Live API response: ta3 symbol changes october 2022

Each row is one security whose price history carries on under a new ticker. A count keyed on the ticker sees two short histories instead of one long one, and the pattern instances at the seam fall out of both.

Try it now

  1. Write your pattern definition down before you look at anything — mechanically enough that a computer would find the same instances twice — and write your horizon and outcome test beside it. Your input is five years of Apple's daily sessions, each an open, high, low, close and volume, drawn below as candles. Mark the instances you find with a Level so your count is visible.
Interactive candles chart: AAPL.US (5Y)
  1. Count every occurrence and the outcome at your chosen horizon. Then compute the unconditional base rate over the same rows and compare the two. Size the noise before you interpret the difference: the standard error of a proportion at n instances is roughly √(0.25 ÷ n).
  2. Then look at what your count is a count of. If your universe is more than one ticker, renamed companies are a trap: the fortnight of renames in the section above is seven price histories that a count keyed on tickers would split in two. Finally, note how many definition variants you tried before settling. If it's more than one, your result is weaker than it looks, and now you know by how much.