หน้านี้คืออะไร: ก่อนที่เราจะให้ระบบไหนถือเงินจริง เราตรวจสอบตัวเองก่อนว่า "ผลตอบแทนที่เจอ เป็นของจริง
หรือแค่บังเอิญ" — เพราะเราทดสอบกลยุทธ์ไปหลายร้อยแบบ ถ้าเลือกแบบที่ backtest ออกมาดีที่สุดโดยไม่เช็คอะไรเลย
โอกาสสูงว่าจะได้ "ของปลอม" ที่ดูดีเพราะบังเอิญ ไม่ใช่เพราะมันเวิร์คจริง หน้านี้ใช้สถิติสองตัว (Deflated
Sharpe Ratio และ Probability of Backtest Overfitting) ที่พัฒนาโดย López de Prado เพื่อตรวจว่าผลที่เห็น
รอดจากการทดสอบแบบเข้มงวดหรือเปล่า ก่อนตัดสินใจให้ระบบไหนถือเงินจริงจริง ๆ — ผลตัดสินใจสุดท้ายของ Nov 2026
จะดูตัวเลขจากหน้านี้เป็นหลักฐานประกอบ
This page is our own audit of ourselves before any system gets real capital: is the edge we found real,
or just the lucky one out of hundreds of variants we tried? These two statistics (both from López de Prado's
work on backtest overfitting) test whether the result survives rigorous scrutiny before we commit real money
to it.
The keystone "which system earns real capital?" test. Over 16 pattern strategy variants
× 212 months (2006-02 → 2026-05). DSR deflates a Sharpe for how many variants were tried + return non-normality;
PBO measures whether "pick the best backtest" survives out-of-sample. Engine: overfitting_stats.py (López de Prado).
Read-only — decision aid, changes nothing live.
Read the metric, not just the colour. These are selective, let-winners-run systems
(few big Webster runners + many −1R stops, plus many zero-trade months). Sharpe structurally penalises that
profile — positive skew and inactivity both lower Sharpe even when per-trade expectancy is strong. So the
per-trade lens is primary (it tests the metric the config was actually selected on: R expectancy + WF
robustness). The monthly lens is shown for completeness and is biased against this style — do not read it as the verdict.
PBO = 0.48
✓ ROBUST — the in-sample best tends to stay above the OOS median 12870 CSCV combinations · S=16 · median OOS rank of the IS-best = 0.53 (0.5 = coin-flip)
The LOCKED Nov 2026 config — distribution-free verdict (PRIMARY)
The valid test for these fat-tailed systems. p_overfit = probability the expectancy is search-luck given 16 variants were tried (bootstrap max-expectancy null; <0.05 green = survives the search). CI = bootstrap 95% interval on expectancy; excluding 0 (green) = a real edge. Bar to clear (null max-expectancy q95) = +0.529R.
Locked component
Expectancy
Rank
Skew
p_overfit
Bootstrap 95% CI
first_pullback_webster_pt
+0.820R
4/16
+5.96
0.002
[+0.45, +1.22]
failed_reentry_partial_2r_ma21
+0.297R
11/16
+3.06
0.331
[+0.16, +0.44]
Why not DSR? The Deflated Sharpe assumes near-normal returns; at skew +5 (the let-winners-run profile) its correction is outside its valid regime and reports ≈0 for edges that are demonstrably real by bootstrap. DSR per-variant is still shown in the ranking table below for contrast — but it is the WRONG tool for this style and should not drive the decision.
THIRD LENS — tail-dependency / edge-concentration (skew-aware)
p_overfit and PBO confirm the edge is real and persistent — neither asks whether the whole
20-year edge rests on a handful of monster trades that may never recur. That is the specific risk of a
let-winners-run profile at high skew. top-3 % = share of total net R from the 3 biggest trades (lower = broader).
Gini = inequality of the winning trades (0 even … 1 all-in-one). Drop-best-k = bootstrap expectancy
after deleting the k largest winners; green = the reduced edge's 95% CI still excludes zero. An edge that survives
deleting its best 5 trades over 20 years is structurally robust — the skew is a few huge winners sitting ON a
genuinely positive base, not standing IN for it.