the record, tested

Every claim we make, with its test attached.

A claim without a test does not appear on this page. The numbers are generated from a graded database, current as of 2026-08-29. Grading is conservative: ambiguous calls are scored against us, and the misses are printed beside the hits.

Scope: 133 graded calls the public could actually have followed. 6 further rows are EXCLUDED because we cannot demonstrate they were public - 4 of them scorable - and they are counted neither for nor against us.

Past performance is not indicative of future results. No performance is promised or implied. Directional accuracy is not profitability - a correct direction can still lose money on timing, sizing, or path.

Educational market analysis only. Not investment advice, not a recommendation, not Series 65 licensed.

Directional accuracy 71%, 95% interval 62% to 79%

When a call named a direction, the market subsequently moved that way inside the call's own window 71% of the time across 112 resolved calls. 73 right, 13 mixed (half credit), 26 wrong.

The test: Exact binomial against a 50% coin-flip null: p = 6.6e-06.
Said out loud: 50% is the weak null. Markets drift, and a short-heavy book in a bear year beats 50% partly by regime. So we computed the strong null and then a version with no human judgment in it at all. They are the next two cards, and the last one is the least flattering.

The record beats luck, not just the coin flip

Random timing on this exact book of calls (same instruments, same directions, same horizons; 2,000 simulated portfolios, fixed seed) averages 47.4% directional accuracy, BELOW the coin flip, because the book is short-heavy and shorts fight market drift. The observed record sits above every one of the 2,000 simulated portfolios.

The test: Pre-registered empirical null simulation, null mean 47.4% (sd 4.8%), null maximum 63.7%. On the 102 calls the simulation covers, observed 71.1% at percentile 100.0: one-sided p < 0.0005 (0 of 2000 simulated portfolios reached it). Luck baselines by direction: longs 53.8% (observed 62.8%), shorts 43.7% (observed 76.2%).
Said out loud: These are human grades, and human grades are more permissive than the simulation's mechanical sign test. That asymmetry favors us, so the next card removes the human from both sides.

With no human grades anywhere, the edge shrinks, and one test stops clearing the bar

Scored purely mechanically (the same rule applied to the real calls at their real dates and to random-date windows, no human judgment on either side), the record scores in the mid-50s against a luck baseline near 48%. The gap between the graded headline above and these numbers IS the grading standard, quantified rather than caveated.

The test: Pre-registered symmetric simulation, two rules, both sides mechanical. STRICT close-to-close sign: observed 57.4% against a 47.6% luck baseline, one-sided p = 0.0325 | PATH MFE > MAE: observed 56.4% against a 48.1% luck baseline, one-sided p = 0.0570. Only 1 of the 2 clears the usual 5% bar; PATH MFE > MAE does not, and we report it as unsettled rather than as a win.
Said out loud: Mechanically the shorts carry the entire edge (61.3% observed against a 44.9% luck baseline, n=62), while the longs score 48.7% against a 53.2% luck baseline, BELOW luck at this sample. And of the calls a human graded right, 24% fail the strict mechanical check; of those graded wrong, 8.3% would have passed it. The generosity ran both ways.

The 2022 record is stronger than the recent record

2022 era: 80% over 62 calls. Recent era: 60% over 50 calls. The eras are always reported separately so the strong era never quietly carries the weak one.

The test: Two-proportion z = 2.30, p = 0.021. A real difference at the usual 5% bar.
Said out loud: This cuts against us and is published anyway. The recent record is the weaker one.

Shorts look stronger than longs, reported as a tendency only

Shorts 74% over 68 calls; longs 63% over 39 calls.

The test: Two-proportion z = 1.24, p = 0.21. NOT statistically settled at this sample size.
Said out loud: We say tendency because the test does not let us say fact. When the sample grows, the verdict updates either way.

The headline survives the harshest scoring

Mixed calls scored as wrong: 65%. As half: 71%. As right: 77%. The page quotes the middle.

The test: Sensitivity analysis across all three codings of partial outcomes.
Said out loud: Grading is conservative (ambiguous calls scored against us) and in-house against independent daily price data. Both facts stay attached to every claim.

What we are not claiming

  • That directional accuracy equals profitability. It does not: a correct direction can still lose money on timing, sizing, or path.
  • Any win rate, return, income, or performance outcome, for us or for you.
  • Anything from backtests. Every call in the database was published or spoken on a dated surface before its outcome existed.
  • That the future resembles the past. It may not.

Check the work yourself

None of this is worth anything if you cannot go and look at it.

How to check us: the database, the generator, and the tests are versioned files, and this page's numbers are emitted by that generator rather than typed. Each regeneration states its date above. If a number here ever disagrees with the generated record, the record wins and this page is wrong.

Past performance is not indicative of future results. No performance is promised or implied. Directional accuracy is not profitability - a correct direction can still lose money on timing, sizing, or path.

Educational market analysis only. Not investment advice, not a recommendation, not Series 65 licensed.