Research

Notes from our own data: real-fill backtesting, prop-firm rule math, live bot transparency. We publish the numbers including the ones that hurt.

The size that passes your backtest can bust your account

Your backtest scores PnL — linear, so bigger always looks better. It doesn't score survival, which on a trailing-drawdown account is convex in size. The ratchet keeps parking you back at max-size risk: ~4.75%/night compounds to ~62% bust over 20 nights. The fix isn't a tighter stop — it's a smaller, survival-calibrated size.

June 2026 · series #10 · size & survival · why backtests lie

The 2–30 minute trap: +$3,364 held vs −$4,141 churned, same signal

We pulled 68 real broker fills from one of our own bots. Trades held longer than 30 minutes made +$3,364; the 2–30 minute churn bucket — over half the activity — bled −$4,141 on the identical signal, and fees were 72% of the gross loss. The same bot run with less churn posted PF 3.64 vs 0.90. Holding time, not the signal, decided the month.

June 2026 · series #9 · churn & fees · why backtests lie

Two prop firms, the same $2,000 drawdown, opposite risk

Your backtest measures drawdown on closed trades. One major prop firm checks it on every tick of open profit — the other doesn't. Same $2,000 limit, opposite risk: a backtest can say you survived while the live account was already liquidated. The per-broker math, with our own near-bust numbers.

June 2026 · series #8 · per-broker DD · why backtests lie

Our edge is +EV. A fat tail still cost $920 in a morning.

One of our cleanest signals — an oversold mean-reversion long in the European morning — is genuinely positive-expectancy. A single fat-tail open still took $920 before the stop. +EV is not survival: here's the tail-guard we shipped the same day — decade-gated, then regression-tested so it can't quietly break the edge.

June 2026 · series #7 · live-ops · EU tail-guard

The P&L gap goes both ways

In our first study the broker was worse than the bot by $830. This week, on a live prop account, the broker came in better by $3.52 — the bot under-reported a winning trade. Same blind spot, opposite direction. The trade, to the cent, and why we reconcile every fill against broker truth.

June 2026 · series #6 · broker-truth · live

We tested 30 strategies. Two survived.

A month of running futures edges through a locked train/test gate. Carry, momentum, calendar, opening-range, relative-value, even our own promising finds — nearly all died. Two uncorrelated survivors (mean-reversion + trend) became a book. And simple beat clever, twice.

June 2026 · series #5 · the whole story

Conditioning beats prediction

Almost every raw edge in liquid futures sits at profit factor 1.1. The lever that moves it isn't a better model — it's conditioning. A single "only after a down day" filter took our overnight edge from PF 1.07 to 1.21, nearly tripled the Sharpe, and halved the drawdown. Plus: why stacking correlated edges doesn't compound.

June 2026 · series #4 · the conditioning lever

We tested the internet's favorite stat-arb template. It died in training.

Z-score pairs mean reversion is the most-cloned retail quant strategy. On futures, with a locked train/test protocol, the cleanest pair posted a training profit factor of 0.72 — and the way the others broke (forced rolls, secular trends) is the real lesson.

June 2026 · series #3 · stat-arb autopsy

A profit factor of 1.3 is what a real edge looks like

The edges that survived 18 years of real-fill testing have profit factors of 1.1–1.3 — and the candidate with t = 4.9 in training died out of sample. Why year-count beats magnitude, and why thin edges live or die on prop-firm rule math.

June 2026 · series #2 · edge anatomy

Paper said +$484. The broker said −$346.

We compared bot-reported P&L against broker statements on our own accounts, found an $800+ gap, and traced it to the fill artifact hiding in most retail backtests. Four strategy families died on the way to two thin, real edges. Full numbers inside.

June 2026 · series #1 · fill realism · 143M data points