Trading Bot HQModels & runsTestsResearchShadow paper

Scan declaration: mentor-TDE event definition vs our displacement definition

Date declared: 2026-07-27 (this document written BEFORE any implementation or data contact). Layer: 3 (event study, no execution, no Validation/Vault contact). Split: BTCUSDT research split only (2017-08-17 → 2022-12-31, frozen §3). Lineage: mentor-direct — the event thresholds are HIS (questionnaire Q6, first mentor-sourced numbers in the project; phase-2-formalization/MENTOR-ANSWERS-JULY-2026-v1.0.md §2). Parameters that Q6 left qualitative are OURS and marked (ours) below.

Question

Does the mentor's own TDE definition (body-vs-mean-body + wick veto) select a better event population than our range-vs-median definition (disp4/disp5) — better meaning stronger or equally strong signed continuation per event, or comparable strength with materially more events?

Event definitions (frozen)

Base series: M5 bars aggregated from canonical Binance M1. Body = |close − open|; range = high − low.

Hypothesis cells (28 total, charged to the 2,000 ledger at registration)

Signed continuation (event-direction-aligned forward return, close-to-close) at horizons {15m, 30m, 1h, 4h}:

Effect metric: mean signed forward return in event-time ATR14(M5) units and in bp, exactly as prior displacement scans (comparability requirement).

Null battery (Edge Lab standard, all four)

Frequency-matched random events; block bootstrap CI (event-level, block length as in displacement scan config); label shuffle (valid here — direction varies per event); time-shifted placebo. Per-cell shuffle-p and bootstrap CI reported; BH-FDR at 10% project-wide via the experiments ledger.

Pre-declared comparison & interpretation rules

Descriptive (not hypothesis-charged): Jaccard overlap of each TDE event set vs disp4/disp5 event sets on the same split; event counts; median event ATR%.

The mentor definition is declared an improvement only if BOTH: 1. ≥2 horizons in some cell show CI-clear positive signed continuation surviving FDR, and 2. at 1h, per-event magnitude ≥ our disp4's on this split (+0.29 ATR), or magnitude ≥ 0.20 ATR with ≥1.5× disp4's event count (n=517 ⇒ ≥776).

Anything else = "his definition is a variant of the same physics, not an improvement" — and the standing displacement conclusions (crypto-class signal, BTC execution dead on Validation) remain governing. No execution test is licensed by this scan; any follow-on needs its own prereg incl. ex-ante cost screen. Regime blocks (per-year sign stability) reported for context.

What this scan cannot say

Nothing about Validation-era behavior (research split only); nothing about harvestability (layer 3 has no costs); nothing about his discretionary use of the definition (run 18's lesson: mechanization ≠ the man).


RESULT (run 2026-07-27, same day as declaration) — NOT AN IMPROVEMENT

Ledger row 25 (+28 hypotheses → 251/2000); artifacts bot/reports/mentor-tde-scan/results-btcusd.json; detector src/tradingbot/edgelab/tde_events.py (causality-tested); commit 452ab3f; suite 252 green.

Both pre-declared improvement criteria FAIL: 1. No TDE cell shows ≥2 CI-clear positive horizons surviving FDR. Best cells: tde3 1h ≈ +0.05 ATR (shuffle-p .125–.130), tde2.5_vetoon 4h +0.08 (p .145) — nowhere near the project BH-FDR threshold. 2. Best per-event 1h magnitude ≈ +0.05 ATR vs disp4's +0.29 on the same split — a ~6× weaker event population, and its 20–40× larger event count does not rescue it under criterion 2's own alternative clause (magnitude < 0.20 ATR).

Structural reading: his body-vs-mean-body population is near-disjoint from our range-based disp sets (Jaccard ≤ 0.036) and much lower-severity (median event ATR% 0.23 vs 0.25–0.26 at 10–40× the frequency) — it mostly selects ordinary active bars, not the rare displacement outliers that carried the continuation signal. The wick veto ("stiffness") helps slightly and monotonically but never enough. The H4-in-M5 variant (n=623, the most severe population, median ATR% 0.60) shows 4h reversion (−0.30 ATR, boot CI clear of zero, shuffle-p .075 — not FDR-significant, logged not promoted): consistent with the project-wide pattern that extreme-magnitude bars mean exhaustion, not continuation (cf. H4-ignition anti-monotonicity, ledger 9).

Standing conclusions unchanged; displacement definitions of record remain disp4/5. Mentor mechanized-claims scorecard: 0-for-5 → 0-for-6. No execution test licensed. Implementation notes of record: the declaration's time-shifted placebo null was newly implemented for the engine (200 re-anchors, +4h…+1wk); ledger charged as one row with hypotheses=28 per existing schema convention.