walk-forward-validation
Walk-forward validation framework for trading strategies and ML models with time-series-aware splits, overfit detection, and regime-aware validation
By agiprolabs · 514 installs
npx skills add agiprolabs/claude-trading-skills --skill walk-forward-validation
Source repository · Upstream listing
Walk Forward Validation
Walk forward validation framework for trading strategies and ML models. Standard cross validation (k fold, random splits) fails catastrophically for financial time series because it introduces lookahead bias and ignores autocorrelation. This skill covers proper time series validation techniques including rolling and expanding windows, purged cross validation, combinatorial purged cross validation (CPCV), and overfit detection metrics.
Why Standard Cross Validation Fails
Standard k fold CV assumes data points are independent and identically distributed (IID). Financial time series violate both assumptions:
1. Lookahead bias — Random splits let the model train on future data and predict past data, artificially inflating performance.
2. Autocorrelation — Adjacent observations are correlated. A random split that puts Monday in test and Tuesday in train leaks information.
3. Regime dependence — Markets shift between regimes. A model trained on a bull market and tested on a bull market tells you nothing about bear market performance.
4. Label overlap — If labels are computed over windows (e.g., 24h forward return), adjacent train/test samples share label computation periods, leaking information.
Walk Forward Framework
Rolling Window (Fixed Train Size)
The train window has a fixed size and slides forward in time. This is preferred when you believe older data is less relevant (common in crypto).
Parameters:
train size : Number of bars/days in the training window
test size : Number of bars/days in the test window
step size : How far to advance between folds (often equals test size )
Expanding Window (Growing Train)
The train window starts at the beginning and expands forward. This uses all available historical data, which helps when data is scarce.
Parameters:
min train size : Minimum training samples before first fold
test size : Fixed test window size
step size : How far to advance between folds
Choosing Between Them
Factor Rolling Expanding
Data recency Prioritizes recent data Uses all history
Regime changes Better adapts to new regimes May dilute recent regime
Sample size Fixed, may be small Grows over time
Crypto preference Preferred for < 6mo horizons Better for regime stable models
Purging and Embargo
Purging
Remove training samples whose labels overlap with the test set's time range. If a label is computed as the 24h forward return starting at time t , any training sample where t + 24h extends into the test period must be purged.
Embargo
Add a buffer gap between the end of training and start of testing to account for serial correlation that purging alone does not eliminate.
Typical embargo sizes:
1 minute bars : 60–240 bars (1–4 hours)
5 minute bars : 12–48 bars (1–4 hours)
Hourly bars : 6–24 bars (6–24 hours)
Daily bars : 2–5 bars (2–5 days)
Crypto rule of thumb : Embargo = 2x the label computation horizon
Combinatorial Purged Cross Validation (CPCV)
CPCV (Lopez de Prado, 2018) generates all possible train/test combinations from N groups while maintaining temporal ordering. This produces far more test paths than standard walk forward, enabling statistical tests for overfitting.
Key properties:
Splits data into N contiguous groups
For each combination of k test groups, the remaining N k groups form the training set
Applies purging and embargo at each train/test boundary
Produces C(N, k) backtest paths (e.g., N=6, k=2 gives 15 paths)
See references/methodology.md for the full CPCV algorithm and formulas.
Overfit Detection
Deflated Sharpe Ratio (DSR)
The observed Sharpe ratio must be adjusted for:
Number of strategies tested (multiple testing)
Non normality of returns (skewness, kurtosis)
Length of the backtest
A DSR below 0.95 suggests the observed performance is likely due to overfitting across the trials tested.
Probability of Backtest Overfitting (PBO)
PBO uses CPCV to measure the fraction of backtest paths where the in sample optimal strategy underperforms the median out of sample. A PBO above 0.50 indicates more likely than not overfitting.
See references/overfit detection.md for complete derivations and implementation details.
Crypto Specific Considerations
1. Shorter windows : Crypto regimes change faster than equities. A 90 day rolling window may be more appropriate than 252 days.
2. 24/7 markets : No weekends or holidays to account for, but funding rate resets (every 8h on perps) create microstructure effects.
3. Survivorship bias : Many tokens delist. Validation must include delisted tokens or at minimum acknowledge this limitation.
4. Liquidity regime shifts : A token's liquidity profile can change dramatically (new CEX listing, liquidity mining end). Train/test splits should ideally not straddle major liquidity events.
5. Data availability : Many tokens have < 1 year of data. Expanding windows with small min train size may be necessary.
Practical Window Sizes for Crypto
Strategy Timeframe Train Window Test Window Embargo
Scalping (1 5min) 3 7 days 1 day 2 4 hours
Intraday (15min 1h) 14 30 days 3 7 days 12 24 hours
Swing (4h daily) 30 90 days 7 14 days 2 5 days
Position (daily weekly) 90 180 days 30 days 5 10 days
Quick Start
Files
References
references/methodology.md — Walk forward theory, window types, purging, embargo, CPCV algorithm with formulas
references/overfit detection.md — Deflated Sharpe ratio, probability of backtest overfitting, multiple testing corrections
references/practical guide.md — Window size selection for crypto, regime considerations, common validation mistakes
Scripts
scripts/walk forward.py — Walk forward validation engine with rolling and expanding windows; demo mode with synthetic data
scripts/overfit detector.py — Deflated Sharpe ratio and PBO computation; demo mode with synthetic backtest results