Trading Bot HQModels & runsTestsResearch

Pre-registration: fable-mentor-internal — mentor v2 spec transplanted to XAUUSD

Date registered: 2026-07-21, before any test execution on XAUUSD. Lineage: mentor. Run name (Richard's): fable-mentor-internal — the mentor-derived internal branch, as improved by the R1 cost-honesty fixes, tested on the instrument class the mentor's own documents claim (gold/BTC). Program rung R2 of phase-2-formalization/MENTOR-CRITIQUE-AND-IMPROVEMENT-v1.0.md.

This is an instrument transplant, not a rule change. The signal logic is byte-identical to the v2 spec (2026-07-21-expansion-v2-costfix.md): same M5 displacement test, body filter, compression precondition, first-pullback confirmation, cancels, daily cap, H4 alignment, zero-baseline guard, rollover exclusion, stop floor. Only pip-denominated constants are translated to gold units, by the declared rules below. Every number is OURS (his docs contain none); his verified numbers (Gates 0.2/0.4) supersede these when they arrive.

Data (fetched & validated 2026-07-21, commit bf80bc2)

Dukascopy XAUUSD M1 bid/ask, 2024-01-01 → 2026-07-14: 1,114,672 bars/side, 0 duplicate timestamps, 0 crossed bars, gaps map 1:1 to holidays/weekends. Tick volume present but not used by this spec (volume rules need the mentor's usage, Gate 0.4 — nothing is smuggled in here). This is a fresh data segment: iteration budget resets. This registration covers run 1 of 3.

Cost model (the decisions this prereg freezes)

Spread is paid in prices (engine fills at ask/bid from the data — measured median spread 0.567 raw, hour-21 UTC 0.77/2.11 p90). The model fixes:

Parameter Value Basis
pip_size 0.1 USD Standard retail gold pip; 1 std lot = 100 oz ⇒ pip value $10/lot, aligning R math with the majors
pip_value_per_lot $10.0 Follows from the above
commission_per_lot_rt $7.0 Same raw-account placeholder class as the FX model (pessimistic placeholder doctrine, costmodel.py); replaced by broker sheet before capital
slippage_entry_pips 1.0 (=$0.10) Scaled from FX (0.2p) by ~5× — between the measured spread ratio (~8×) and M5-volatility ratio (~14×), leaning pessimistic vs FX but not absurd
slippage_stop_exit_pips 4.0 (=$0.40) Same scaling from FX 0.8p
swap long/short per lot-night −35 / −10 Pessimistic placeholder; ~never triggered (hold_max ≤ 45 min)
session spread table (reporting only) from our measured by-hour table; rollover entry 10.0 pips pessimistic Own measurement, this fetch

Implied round-trip friction ≈ spread 5.7 + slippage 5.0 + commission 0.7 ≈ 11.4 gold pips (~$1.14) on a stopped trade. By the drag law, a 60-pip stop carries ≈ 0.19R drag; the 0.05R eligibility bar needs ≈ 230-pip stops. Declared up front: typical displacement-geometry stops here will likely carry 0.1–0.3R drag — better than GBPUSD's 0.4–0.6R, still above the bar.

Constant translation (declared rules, not tuning)

Walk-forward 12-month IS / 3-month OOS rolling (harness default); IS selection = best net expectancy with ≥ 20 IS trades. Baseline: same coin-flip control as v2 (identical event set, floored 0.5×D stop). Registry: strategy fable_mentor_internal, pair XAUUSD, lineage mentor, kind backtest.

Registered hypotheses

Pass criteria (all four, stitched OOS only — identical to every prior run)

  1. OOS net expectancy > +0.05R
  2. Beats the coin-flip baseline's net expectancy
  3. |OOS max drawdown| ≤ 2 × |baseline max drawdown|
  4. ≥ 100 OOS trades

Expected outcome (honest prior)

FAIL, most likely on criteria 1 and/or 4 — the friction arithmetic above says typical displacement stops still carry 2–6× the eligible drag, and gold's event frequency on 2.5 years may not reach 100 OOS trades. The decision tree is registered now: if H1 holds (gross > 0) with H2 (drag halved), the branch continues toward geometry that clears the bar (wider-stop variants, his real numbers). If gross ≤ 0 on gold too, our formalization of his mechanics is refuted on both instrument classes we can test, and the mentor branch pauses until his verified data (Gates 0.1–0.4) re-anchors the spec.

Iteration budget

Run 1 of 3 on the XAUUSD 2024–2026 segments. No in-sample smoke tests have been run on this instrument. Any rule, grid, or cost-model change after seeing results = new pre-registration. No re-runs of this spec.


Result (filled 2026-07-22, after run 1)

FAIL — on criteria 1 and 4, exactly as the honest prior predicted. All three registered hypotheses HELD. Stitched OOS (2025-04-04 → 2026-06-26), 43 trades: net −0.0404R, gross +0.1002R, win rate 34.9%, max DD −12.79R, WFE −0.372 (metric unreliable near zero — known issue). Criteria: net ✗ (−0.04 ≤ +0.05), beats baseline ✓ (−0.04 vs −0.27), drawdown ✓ (12.79 ≤ 2 × 25.88), sample ✗ (43 < 100). Report: reports/review-fable_mentor_internal.html · registry run 16, fable_mentor_internal@1a29b085, cost xauusd@97dbc9ad, lineage mentor.

Where the run leaves the branch (decision tree, as registered): gross > 0 with drag halved ⇒ the branch continues. The remaining gap is arithmetic, not artifact: net −0.04R sits 0.09R below the bar while drag is 0.14R — so the same signal with drag ≤ ~0.05R (stops ≥ ~230 pips, or cheaper execution) would clear criterion 1 if the gross edge survives the geometry change, which is NOT guaranteed (widening stops shrinks gross-per-R; v2 measured this). Sample is the second binding constraint: 43 trades in ~15 OOS months will not reach 100 without more data or a rule change. Both paths — wide-stop geometry variant, and any sample-widening change — are rule changes requiring a new pre-registration (runs 2–3 of this segment's budget remain).

Iteration budget: run 1 of 3 consumed. In-sample smoke tests: none.