Mattheus Public Preview

Join Waitlist

Mattheus Public Preview

Current Blog

How to Estimate Confidence Intervals for a Trading Strategy Backtest

A backtest confidence interval estimates how much a strategy statistic could vary under repeated samples, provided the resampling method preserves the dependence in the data.

Published Updated
Blue resampled outcome paths pass through a translucent chamber and spread between warm confidence bounds

By alyc

A backtest confidence interval is a range for a strategy statistic, such as average after-cost return, drawdown, turnover, or a risk-adjusted measure. The range estimates how much that statistic could change across repeated samples from a stated data-generating process. It does not predict the next return, and it does not turn a historical result into a guarantee.

The interval is only as credible as the backtest and the resampling rule behind it. Market observations often share trends, volatility clusters, and overlapping positions. Resampling individual rows as though they were independent can produce a narrow range that understates uncertainty. A defensible workflow names the statistic, preserves the relevant dependence, records every research choice, and reports the method beside the result.

Define the statistic before resampling

Start with the quantity used to approve or reject the strategy. A confidence interval for mean daily return answers a different question from an interval for maximum drawdown. The observation unit also changes the result. Daily portfolio returns, completed trades, and monthly strategy returns do not carry the same dependence or sample size.

Calculate the statistic after applying the same fees, spread, slippage, funding, and fill assumptions used in the backtest. The resampling code should call the complete performance calculation for every replicate. It should not resample gross returns and subtract one fixed cost estimate at the end. How to Model Trading Costs in a Backtest explains why costs must follow the simulated order and market state.

Write down whether the interval is two-sided or one-sided. The NIST explanation of confidence intervals distinguishes a two-sided range from an upper or lower bound. A downside review may need a lower bound for average return or an upper bound for drawdown. Choose that direction before looking at the resampled distribution.

Build a bootstrap distribution

A bootstrap repeats three operations. It draws a resample from the observed records, recalculates the chosen statistic, and stores the result. Repeating the process creates an empirical distribution for that statistic. The NIST bootstrap plot guidance describes the resulting values as a way to inspect the location and variation of a sampling distribution.

The simplest bootstrap samples observations with replacement until the resample has the same length as the original series. This approach is suitable only when the sampled units can reasonably be treated as independent and identically distributed for the question being asked. Trade returns from one position can still share exposure with nearby trades. Daily returns can inherit serial dependence from holding periods, signal persistence, and volatility regimes.

Software can calculate percentile, basic, and bias-corrected and accelerated intervals. The SciPy bootstrap documentation lists all three methods and returns both the bootstrap distribution and its standard error. Record the library version, interval method, confidence level, number of resamples, random seed, and statistic function. Otherwise, another researcher cannot reconstruct the range.

Preserve dependence with blocks

For a return series with meaningful serial dependence, resample contiguous blocks instead of single periods. A block keeps neighboring observations together. The method then assembles a pseudo-series from those blocks and recalculates the statistic.

The stationary bootstrap proposed by Dimitris Politis and Joseph Romano uses blocks of random length. The Stanford technical report record identifies the method and its original authors. The published method was designed to estimate standard errors and confidence regions for parameters based on weakly dependent stationary observations.

Block length is a research choice. Blocks that are too short break the dependence the method is meant to preserve. Blocks that are too long leave few distinct rearrangements and can make the interval unstable. Report results across a small prespecified range of plausible block lengths. Do not select the one that gives the most attractive lower bound.

Stationarity is also an assumption, not a label created by the method. A series that crosses a structural break, venue migration, position-sizing change, or new execution rule may not represent one stable process. In that case, split the analysis by a documented regime, shorten the estimation period for a stated reason, or report that the available history cannot support a stable interval.

Reproducible example: individual days versus three-day blocks

The following Python example uses twelve invented after-cost daily returns. It resamples either individual observations (block length 1) or circular blocks and reports a two-sided 95% percentile interval for mean daily return. It is a demonstration of mechanics; twelve invented observations cannot support a trading conclusion.

With seed 7 and 10,000 resamples, the point estimate is 0.050% per day. Individual-day resampling gives −0.125% to 0.225%; three-day blocks give approximately −0.142% to 0.233%. Block lengths 2 and 4 give −0.150% to 0.250% and −0.100% to 0.200%, respectively. Width is not monotonic in block length, so report the prespecified sensitivity instead of selecting the most favorable interval.

from random import Random
from statistics import mean, quantiles

returns = [.004, .003, .002, -.002, -.003, -.004,
           .005, .004, .003, -.003, -.002, -.001]

def interval(block_length):
    rng = Random(7)
    estimates = []
    for _ in range(10000):
        sample = []
        while len(sample) < len(returns):
            start = rng.randrange(len(returns))
            sample.extend(returns[(start + j) % len(returns)]
                          for j in range(block_length))
        estimates.append(mean(sample[:len(returns)]))
    cuts = quantiles(estimates, n=40, method="inclusive")
    return cuts[0], cuts[-1]

for length in [1, 2, 3, 4]:
    lo, hi = interval(length)
    print(length, f"{100 * lo:.3f}%", f"{100 * hi:.3f}%")
print("mean", f"{100 * mean(returns):.3f}%")
from random import Random
from statistics import mean, quantiles

returns = [.004, .003, .002, -.002, -.003, -.004,
           .005, .004, .003, -.003, -.002, -.001]

def interval(block_length):
    rng = Random(7)
    estimates = []
    for _ in range(10000):
        sample = []
        while len(sample) < len(returns):
            start = rng.randrange(len(returns))
            sample.extend(returns[(start + j) % len(returns)]
                          for j in range(block_length))
        estimates.append(mean(sample[:len(returns)]))
    cuts = quantiles(estimates, n=40, method="inclusive")
    return cuts[0], cuts[-1]

for length in [1, 2, 3, 4]:
    lo, hi = interval(length)
    print(length, f"{100 * lo:.3f}%", f"{100 * hi:.3f}%")
print("mean", f"{100 * mean(returns):.3f}%")
from random import Random
from statistics import mean, quantiles

returns = [.004, .003, .002, -.002, -.003, -.004,
           .005, .004, .003, -.003, -.002, -.001]

def interval(block_length):
    rng = Random(7)
    estimates = []
    for _ in range(10000):
        sample = []
        while len(sample) < len(returns):
            start = rng.randrange(len(returns))
            sample.extend(returns[(start + j) % len(returns)]
                          for j in range(block_length))
        estimates.append(mean(sample[:len(returns)]))
    cuts = quantiles(estimates, n=40, method="inclusive")
    return cuts[0], cuts[-1]

for length in [1, 2, 3, 4]:
    lo, hi = interval(length)
    print(length, f"{100 * lo:.3f}%", f"{100 * hi:.3f}%")
print("mean", f"{100 * mean(returns):.3f}%")

The circular fixed-block method above is not the random-length stationary bootstrap and does not reproduce a full trading engine. It assumes the supplied after-cost return sequence is the resampling unit. When costs, position state, or path-dependent rules change under resampling, recompute them within each replicate. The interval describes this toy sample under this rule; it does not account for earlier strategy selection.

Keep selection outside the interval

A bootstrap interval measures sampling variation under its assumptions. It does not account for every strategy, feature, threshold, or date range tried before the reported strategy was chosen. If the winner emerged from hundreds of trials, the interval for that winner can still look precise because the selection step is missing.

Keep a trial ledger and preserve the full candidate set. How to Detect Backtest Overfitting covers the record needed to separate model selection from later evaluation. The final interval should be calculated on a period that did not determine the surviving specification, or its label should state that it describes development data.

Walk-forward testing offers another view of uncertainty. It shows how the same research process behaved at successive decision dates. Compare the pooled bootstrap interval with fold-level results from a chronological walk-forward test. Agreement does not prove future performance, but a sharp conflict often reveals dependence on one period or one regime.

Read the interval correctly

A 95 percent frequentist confidence interval does not mean there is a 95 percent probability that the fixed true parameter lies inside this realized interval. The coverage statement refers to the long-run behavior of the interval procedure across repeated samples under its assumptions. NIST describes this as the share of repeated intervals that would bracket the parameter.

The width matters as much as the point estimate. A positive mean return with a wide interval crossing zero is weak evidence about the sign of the expected return. A narrow interval built from thousands of overlapping rows may be false precision. Always show the sample span, effective observation unit, dependence rule, and interval bounds together.

Inspect the bootstrap distribution rather than keeping only two endpoints. Strong skew, several modes, or a pileup at a boundary can signal that the statistic or resampling rule needs review. A degenerate distribution is also a warning. SciPy notes that its bias-corrected and accelerated interval can return undefined bounds when all bootstrap statistics are identical.

Store an uncertainty record

The research record should include the data snapshot, code revision, strategy parameters, cost model, statistic definition, observation unit, resampling method, block rule, block-length sensitivity, confidence level, interval method, number of resamples, random seed, point estimate, bounds, and distribution summary. Link that record to the exact backtest output.

Data lineage for trading research keeps the input and transformation history attached to the result. That link matters when corrected market data or a revised cost model changes the interval later.

An interval remains historical evidence. It cannot cover a market regime absent from the sample, a future venue failure, or a rule change that alters execution. Collect forward observations after the research procedure is frozen. Paper-first evaluation provides a separate record of decisions made after historical testing ended. Keep that record outside the bootstrap sample until the review date specified in advance.

Editorial note: This article was prepared with AI assistance and checked against the cited sources. No named human review is recorded.

Explore Markets with Infinite Context.

Explore Markets with Infinite Context.

Browse the public preview. Join the waitlist for gated trading features.

Browse the public preview. Join the waitlist for gated trading features.

Browse the public preview. Join the waitlist for gated trading features.

Public preview · Join the waitlist

Dashboard interface preview
Dashboard interface preview