Current Blog
How to Build a Walk-Forward Test for a Trading Strategy
A walk-forward test trains on past data, evaluates the next untouched period, and repeats that sequence across several decision dates.
By alyc
A walk-forward test recreates a sequence of historical decision points. At each point, the strategy is fitted with data that would have been available then and evaluated on the next untouched period. The clock moves forward and the process repeats. This design shows whether a fixed research procedure held up across several later windows instead of succeeding in one convenient split.
The test needs a written schedule before any result is viewed. That schedule fixes the first decision date, the training span, any gap between training and testing, the test span, the step between folds, and the last eligible date. It also states what gets refitted and what remains frozen. Without that record, a walk-forward label can hide substantial discretion.
Start with the decision the strategy will make
Define the unit of prediction or action before choosing a fold length. A strategy that updates once a month needs a different test window from one that makes an intraday decision. The test period should cover the horizon that matters to the actual rule, including the time needed for the position to enter, mature, and exit.
Set one timestamp for the simulated decision and record the latest data release available at that timestamp. Corporate actions, revised economic series, exchange calendars, and late venue messages can all change after the fact. Look-ahead controls for trading backtests belong inside every fold, not in a separate review after the test has run.
The first training window must be long enough to estimate the strategy without consuming every useful test period. There is no universal ratio. A slow signal may need several market cycles. A short-lived microstructure rule may become misleading when trained on records from a retired venue regime. The choice should follow the mechanism being studied.
Choose expanding or rolling training windows
An expanding window begins at a fixed date and adds new observations before each test. It suits a process that is expected to learn from its full history. A rolling window keeps a fixed length and drops the oldest observations as it advances. It suits a process where older regimes may no longer describe the present market.
Leonard Tashman's review of out-of-sample forecasting tests separates choices such as fixed and rolling origins, updating and recalibration, fixed and rolling windows, and single and multiple test periods. Those choices change the question the test answers. They should appear in the method record rather than emerge from implementation defaults.
Run both window types only when the research plan has a reason to compare them. If the better result determines which design survives, that comparison becomes another model-selection trial. Add it to the trial ledger described in How to Detect Backtest Overfitting.
Keep each test period in the future
Ordinary random cross-validation mixes earlier and later observations. That makes it unsuitable when a model could learn from a record that occurred after the decision it is being asked to predict. The scikit-learn TimeSeriesSplit documentation uses ordered splits where each training set precedes its test set. Successive expanding training sets contain the earlier ones.
A gap may be needed between the last training observation and the first test observation. Scikit-learn exposes a gap parameter that excludes samples from the end of each training set. The correct gap depends on the data. Overlapping labels, delayed publications, execution horizons, and features built from future-return windows can carry information across a fold boundary even when row timestamps appear ordered.
Preprocessing stays inside the fold. Fit scalers, missing-value rules, feature selection, model parameters, and calibration with the training records for that fold. Applying a transformation estimated from the complete history lets later distributions influence earlier decisions.
Match the test horizon to the use case
Rob Hyndman and George Athanasopoulos describe time series cross-validation as evaluation on a rolling forecasting origin. Their time series cross-validation chapter notes that multi-step forecasts require errors at the relevant horizon. A trading test should follow the same principle.
If the strategy holds for five trading days, a one-day test tells an incomplete story. If orders are rebalanced monthly, a test made of isolated daily predictions can overweight neighboring observations that belong to the same position. Define how overlapping positions, repeated signals, and shared market events are counted before aggregation.
The step between folds can equal the test span or be shorter. Short steps create more observations, yet overlapping test periods share outcomes and should not be treated as independent evidence. Keep fold-level results visible so a dense schedule does not create false precision.
Worked example: write a fold ledger in session numbers
Use an illustrative sequence of equally spaced trading sessions numbered 1–100. Each label is a one-session forward return: the label attached to session 39 becomes known at session 40. Fit after session 40; the first fold trains on feature rows 1–38, excludes rows 39–40 as a two-row gap, and tests decisions on rows 41–50. Score the last test label only after session 51.
Fit the second fold after session 50 using rows 1–48, exclude 49–50, and test 51–60; score its last label after 61. A fixed two-row gap is an assumption for this toy example, not a universal rule. Check every label’s actual availability time, including delayed records, before admitting it to training.
TimeSeriesSplit defines its gap in samples, not calendar days. Record the time represented by each row, the fitting cutoff, and the evaluation horizon. Keep preprocessing inside each training fold and reserve a later untouched period if these folds influence model selection.
Recreate trading costs in every fold
Each forward period should use the prices, spreads, volume, fee schedule, and borrow or funding assumptions available for that period. A cost rule calibrated with future fills leaks later execution knowledge into earlier folds. Trading-cost modeling for backtests should be fitted or selected inside the same chronological boundary as the strategy.
Report return after costs, drawdown, turnover, fill rate, exposure, and the number of decisions for each fold. One pooled return can hide a process that worked in one long window and failed in the rest. Show the distribution across folds and identify which market conditions produced the weakest outcomes.
Keep an untouched evaluation after model selection
Walk-forward folds used to choose features, parameters, window length, or approval thresholds belong to model development. They are out of sample for each local fit but no longer untouched once their results influence the design. Reserve a later period for one evaluation after the procedure and its acceptance rule are frozen.
Write that rule before opening the reserved period. It can require acceptable after-cost performance across most folds, bounded drawdown, stable turnover, enough independent decisions, and no dependence on one narrow parameter choice. The exact limits follow the strategy. Their timing is the control.
A fold ledger should preserve the training start and end, data cutoff, gap, test start and end, code revision, data snapshot, parameters selected, cost model, result, and decision made after review. Data lineage makes the schedule reproducible when source data or transformations change.
Know what the test cannot establish
A clean walk-forward result does not prove that future performance will match the historical estimate. The available history may omit the regime that breaks the strategy. Several folds may still contain the same broad market event. Repeated design changes can overfit the walk-forward schedule itself.
The result can support a narrower claim. It can show how a frozen research procedure behaved when it repeatedly trained on earlier records and met later ones. Paper-first testing then adds decisions collected after historical research ended. Keep both records separate so a retrospective test is never presented as live evidence.
Editorial note: This article was prepared with AI assistance and checked against the cited sources. No named human review is recorded.
More blogs to read
Public preview · Join the waitlist




