Why backtest matters
It can test logic, sensitivity, turnover, and failure modes, but attractive historical results are easily manufactured through overfitting or biased data.
How it is applied
Define the investment rule before examining results, then simulate it using point-in-time data, realistic rebalance dates, investable securities, fees, spreads, market impact, taxes where relevant, and constraints that could actually have been followed. Reserve later observations for validation and compare the strategy with a suitable benchmark and simple alternatives.
Portfolio example
A researcher tests a monthly value strategy from 2005 to 2025. Delisted companies remain in the universe, accounting data becomes available only after its publication date, and each trade pays estimated costs. Parameters are chosen on 2005 to 2016 data, checked on 2017 to 2020, then evaluated untouched from 2021 onward.
How to interpret it
A useful backtest shows how a rule might have behaved under historical conditions, not what an investor actually earned. Examine return sources, turnover, drawdown, capacity, subperiods, and sensitivity to modest parameter changes. Performance that disappears after costs or depends on one narrow setting is weak evidence.
Limitations and common misconceptions
Look-ahead bias, survivorship bias, data mining, overfitting, stale prices, and unrealistic fills can create impressive but untradeable results. Historical regimes may not recur, and crowded adoption can erode an anomaly. Statistical significance does not establish an economic rationale, while repeated experiments make conventional confidence measures misleading. A credible report preserves the original hypothesis, data vintage, code version, parameter choices, rejected tests, and out-of-sample result. It should disclose whether the backtest is gross or net and distinguish simulated performance from a live track record. Stress tests and paper trading are useful bridges, but neither guarantees future results. Useful robustness checks include changing the start date, rebalance frequency, portfolio breadth, cost assumption, and signal definition within economically defensible ranges. Results should also be segmented by bull and bear markets, inflation regimes, volatility, geography, and liquidity. If one episode supplies most excess return, readers need to see it. Comparing the proposed rule with a simpler version helps reveal unnecessary complexity. The final report should state how many variants were tried, because selecting the winner from hundreds of tests gives a very different level of evidence from testing one pre-registered hypothesis.
Sources and further reading
- Equity Valuation: Applications and ProcessesCFA Institute