Trading Bot HQModels & runsTestsResearch

Decision memo: streak_follow_btc_v1 — the second bar-clearing candidate

Date: 2026-07-24. Everything below is research-split evidence; Validation is untouched (9 of 10 looks). Chain of record: R4 rhythm scan → ETH replication scan (both declared-then-run) → execution prereg 2026-07-24-streak-follow-btc-v1.md with ex-ante cost screen → registry run 21 (BAR MET) → Monte Carlo characterization (reports/montecarlo/trades-A1-mc.md). Dashboard: HQ → Models & runs → run 21.

What we have, in one paragraph

After 8 consecutive up-closes on M5, BTC keeps drifting up for ~4 hours. Scan-level: two instruments (BTC 6/6 year-blocks, ETH replicated with dose response and the same asymmetry), full null battery. Execution-level (enter next M1 open, 2×ATR stop, 4h time exit, longs only): net +0.107R/trade at Vantage-class costs, n=487, beats the coin-flip control (+0.024R), 4/6 year-blocks gross-positive — including 2021 (+0.29) and 2022 (+0.23), the two years where displacement's edge was already fading. Ex-ante cost screen passed at 5.3× before the run was registered. One honest miss: H4's drag point-estimate ran 46% high (right-skewed ATR + Jensen's inequality — the convention was kept, the sensitivity reported, nothing re-run).

How it compares to the candidate that failed Validation

disp_follow (run 19 → 20) streak_follow (run 21)
Research net (Vantage) +0.28R (A2, n=282) +0.107R (n=487)
Control gap (net) +0.36R +0.083R
BTC recent blocks (2021/2022 gross) +0.74 / +0.06 +0.29 / +0.23
ETH scan recent years (4h) replicated broadly 2021 −0.12, 2022 −0.67 (negative)
Validation outcome FAIL (direction inverted post-2022) — (undecided)

Two honest readings of the same table: the for case — smaller edge but steadier on BTC into 2021–22, exactly where displacement was already dying; the against case — ETH's most recent two years are negative at the 4h horizon, and run 20 taught us this effect family can invert when the regime changes. I'd put the Validation pass probability at roughly a coin flip, not better.

The drawdown reality (Monte Carlo, block bootstrap, 10k paths)

The observed path (+52.3R total, 31.6R max drawdown) is the median experience of this trade distribution — the p95 path draws a 68R drawdown, p99 89R, and the median path spends 191 of 487 trades under water (~40% of the whole sample). Terminal outcomes range down to −134R: the total depends heavily on a few large winners (max single trade +17R; win rate 31%). Any future deployment sizing must survive the p95 path, not the observed one.

Decision — spend Validation look 2 of 10?

§5.6 is already satisfied for this candidate (ETH scan replication under Amendment A1), so there is exactly one decision. If yes: a fresh Validation prereg (criteria frozen first, same structure as run 20's: proposed net(Vantage) ≥ +0.05R · n ≥ 80 expected ~225 · beats control · maxDD ≤ 2× control's), the frozen spec run once against 2023-01-01→2025-06-30, machinery guard (exact reproduction of run 21) before first contact, zero Vault bars. A fail is final for this spec.

My recommendation: yes — with eyes open. This is what the looks are for, the candidate is the strongest object since run 19, and waiting adds no information the split wouldn't add better. But price in the coin-flip prior above: a second consecutive Validation fail would also be the project working correctly, and would sharpen D2's eventual question (does anything from this technique class survive the current regime?).

Nothing runs until you answer. Reply e.g. "look 2 yes" (or ask for anything deeper).


Addendum 2026-07-25 — the ETH execution cross-check is closed

The planned free de-risking step (execute the spec on ETH's research split) was refused by the ex-ante cost screen: measured Vantage-class ETH costs are 22.6bp RT (~7× BTC — spread-only crypto CFDs, ~$4 on ~$1,860), above the effect's own 17.5bp 4h gross. 0.77× where 3× is required; no stop geometry changes that (2026-07-25-streak-follow-eth-REFUSED.md).

Two consequences for this decision: (1) it now rests on the existing inputs alone — two-instrument scan replication, BTC execution (run 21), and the MC characterization; waiting adds nothing further. (2) If the candidate ever deploys, it is BTC-only at this broker's cost structure — ETH validated the signal, not the harvest. My recommendation is unchanged (yes, coin-flip prior, eyes open), and the decision is now ripe rather than improved by waiting.


Addendum 2 — 2026-07-25: the universality sweep weakens the case; recommendation REVISED

The sweep's pre-declared score-keeping said a patchy streak map (≤2 of 5) "materially weakens the look-2 case." It landed exactly there: 2/5 decisive (BTC, ETH), SOL supporting-only (2.4y of data), BNB a narrow miss (4h CI crosses zero), and XRP INVERTS — CI-clean −0.41 ATR at 4h. Displacement, scanned identically on the same instruments, replicated 4/5. So the streak effect is not asset-class structure; it lives only on the largest leveraged venues, and the XRP inversion demonstrates the family can flip sign across instruments — the same fragility that killed displacement across time in run 20.

Revised recommendation: HOLD the look — one more free step first. My prior drops from coin-flip to roughly 35–40%. Unlike yesterday (when the ETH execution path closed and waiting added nothing), there is now a genuinely informative free step: the regime-conditioning scan (research split, declared next) — when within 2017–2022 does the streak effect pay, against vol-regime state? If the effect concentrates in identifiable regimes, the better object to take to Validation may be a regime-conditioned successor (new prereg, fresh evidence); if it is regime-flat, the original candidate's case partially recovers. Either way the look is spent smarter. Spending 1 of 9 remaining looks at ~35–40% today, when a free scan can move that number this weekend, is worse than waiting a day. If you want to spend it anyway, everything is ready — but this memo now recommends: wait for the regime scan.


Addendum 3 — 2026-07-25 (later): regime scan ran; free evidence is now exhausted; final position

The regime-conditioning scan (2026-07-25-streak-regime-conditioning-scan.md, 24 hypotheses, both instruments) came back ambiguous by the pre-declared rule: no vol band concentrates the effect beyond CI overlap on either instrument — but the partition is underpowered (~160 events/band) and the instruments disagree in shape (BTC monotone stronger-with-vol; ETH low-vol-favoring). It neither convicts the candidate of regime-concentration nor certifies regime-flatness. The one directional whisper is mildly unfavorable: BTC's low-vol band (closest analog to 2023–25) has the weakest point estimate (+0.38 vs +0.74 high-band).

There are no further free evidence steps. Every research-split angle this candidate can be examined from has been: execution (run 21), 5-instrument map, regime bands, MC drawdowns, cost screens. The decision is now purely: spend look 2 at a ~35–40% prior, or hold.

The two honest positions, stated fairly: - Spend (my lean, stated with its reasoning): a fail is genuinely valuable — two independent effect families failing the same 2023–25 holdout would be strong evidence about this entire technique class in the current regime (it sharpens D2 from speculation into data). A pass is the project's first validated edge. Looks don't earn interest by being hoarded: 9 remain against a 2,000-hypothesis program, and this is the strongest object the research split can produce. - Hold (the defensible alternative): ~35–40% is a poor price for a scarce look; the Vantage demo (measured costs — the commission audit suggests BTC net may be understated) and the mentor's data response arrive soon and cost nothing; D4 (2026-10-22) is the natural review point. Nothing decays by waiting except momentum.

Either answer is respectable. Reply "look 2 yes" or "hold until D4" — the machinery for both paths is ready, and nothing runs until you do.