Every backtest overstates live results. Always. The gap between them is called backtest-to-live decay, and it comes from five sources: spread, slippage, look-ahead bias, curve fitting, and execution quality. A backtest showing +0.5R per trade often produces +0.25R live. Discount every backtest by 30–50% before you trust it.
Why this closes Block 7
Block 7 built the entire risk framework: 1% per trade, stop placement, portfolio heat, expectancy, sample size. All of that only matters if the strategy producing those numbers is real — not an artefact of a curve-fitted backtest.
This lesson is the reality check. It is the difference between a strategy you can trade for a decade and a strategy that looked great on a spreadsheet and died in month two.
T
Written by the Trade To The Top team|Reviewed 30 September 2026
Backtest methodology cross-checked against Evidence-Based Technical Analysis (Aronson), Advances in Financial Machine Learning (López de Prado), and the backtest-vs-live decay research published in the algorithmic trading literature. Slippage and spread modelling verified against broker execution studies and MiFID II best-execution reporting standards.
You built a strategy. You backtested it over 500 trades. Expectancy: +0.5R. You go live. Fifty trades later, expectancy is +0.15R. Did your edge disappear, or was it never as big as the backtest said? Almost always, the second one. This lesson explains exactly where the missing 0.35R went — and how to build a backtest that does not lie to you.
Key takeaways
Every backtest overstates live performance. There are no exceptions.
Five sources of decay: spread, slippage, look-ahead bias, curve fitting, execution quality.
Discount every backtest by 30–50% before you trust the expectancy.
Use realistic costs. Backtests that assume zero spread or zero commission are fiction.
No look-ahead bias. You cannot use the close of a candle to enter at its open.
Curve fitting is silent. A strategy that was optimised on the same data it was tested on is not a strategy. It is a memory.
Walk-forward testing is non-negotiable. Optimise on data, then test on data the strategy has never seen.
Sample size still applies. 50 live trades is not evidence. 150+ is.
Live execution differs from paper execution. Spreads widen, orders slip, and emotions intervene.
The right benchmark is a realistic backtest, not a perfect one.
Here is the reality. No backtest replicates live trading. The backtest uses historical data, clean entry prices, no emotional pressure, and no real execution. Live trading uses today's data, real fills, real slippage, and real emotions. The two will never match.
The question is not "will there be a gap?" The question is "how big is the gap, and how much can I trust the backtest?"
BACKTEST VS LIVE · THE EXPECTANCY DECAY
Same strategy · 500 backtest trades · 150 live trades · five sources of decay
Five sources of decay. Same strategy. Backtest +0.50R. Live +0.15R.
Common mistake
Running a backtest that uses the candle close as the entry price. You cannot fill at the close of the candle that generated the signal. By the time the candle closes, price has already moved. A realistic backtest enters at the next candle's open, or on a limit order at the level — never at the signal candle's close.
Spread and commission
The first source of decay is the most obvious and the most commonly ignored. Backtests that assume zero spread are fiction. Every trade costs the spread. On EUR/USD at 1 pip spread, a 20-pip stop, and 200 trades per year, that is 200 pips of pure cost — two full R of expectancy stolen per year before the first trade.
Cost modelling — what to include
01
Spread: model the typical spread on your broker, not the best-case spread. Use 1.0–1.5 pips on majors, 2.0–3.0 pips on crosses and gold.
02
Commission: if your broker charges per lot, model it. $7 round-turn per lot is standard on ECN accounts.
03
Swap: for trades held overnight, model the swap. On a long EUR/USD it can be negative several pips per night depending on rate differentials.
04
Wider spread during news: if your strategy trades around releases, model spreads 3–10x normal for those bars.
05
Weekend gap: if you hold through the weekend, model a gap cost of 5–20 pips on the Monday open.
Cost model
False result
Realistic result
Spread
0 pips — "enter at mid"
1.0 pip per side on majors
Commission
$0 per lot
$7 round-turn per lot
Swap
Ignored
Modelled per overnight hold
Expectancy impact
Inflated by 0.05–0.15R per trade
Realistic
Slippage
Backtests assume you get filled at the exact price you wanted. Live trading does not. Slippage is the difference between the price you requested and the price you got. It affects entries, stops, and targets — and it always moves against you on the trades that matter most.
The worst case is stop-loss slippage during news. A 20-pip stop on a quiet day fills at 20 pips. A 20-pip stop on NFP can fill at 40. Half the losses you backtested at 1R will actually be 1.5R or 2R live.
SLIPPAGE · WHERE THE EXTRA R DISAPPEARS
Same trade, three execution scenarios · backtest vs live quiet vs live news
Same setup. Same prices. Different execution. From +0.35R to +0.05R.
The backtest assumes perfect fills. The market does not care what the backtest assumes.
Look-ahead bias
Look-ahead bias is using information you would not have had at the time of the trade. It is the silent killer of backtests because it does not produce obviously bad results — it produces obviously good ones.
Common forms:
Entering at the signal candle's close. In live trading, you cannot enter at the close of the candle that just closed. You enter on the next tick, which is the next candle's open — usually worse.
Using the final high or low of a session to determine whether a level was tested. In real time, you did not know where the session would end. You only knew where it was so far.
Backtested indicators that use future data. Some indicators (like smoothed moving averages with look-ahead implementations) use future bars to compute past values. This is a coding error, not a strategy feature.
Bar-magnifier data. Using tick-level data for the exit but bar-level data for the entry. The asymmetry leaks information you would not have had.
The look-ahead illusion
A strategy that appears to catch the exact top or bottom of a range is almost always using look-ahead bias. Nobody catches exact tops and bottoms. If your backtest says you do, the backtest is wrong — not the market. Fix the entry rule to something a human could actually execute in real time.
Curve fitting
Curve fitting is optimising a strategy to fit the past, not to trade the future. It is the single most destructive backtest error and the easiest one to commit without realising it.
The process looks innocent. You test a strategy. It gives +0.2R. You tweak the parameters. Now it gives +0.3R. You tweak more. Now +0.5R. You add a filter. +0.6R. You have just built a strategy that perfectly describes the last 500 bars — and is worthless for the next 500.
CURVE FITTING · PERFECT ON THE PAST, DEAD ON THE FUTURE
Same strategy · optimised data vs unseen data
Perfect on data it was fitted to. Negative on data it has never seen.
Metric
Curve fitted
Robust
In-sample expectancy
Very high (+0.50R or more)
Moderate (+0.20–0.35R)
Out-of-sample expectancy
Negative or near zero
Similar to in-sample
Number of parameters
Many — 5+ tuned inputs
Few — 1–3 total
Sensitivity to parameters
Falls apart with small changes
Stable across a range
Performance across instruments
Only works on one instrument
Works on several
Logic
"It just works in the data"
Clear structural reason
Walk-forward testing
Walk-forward testing is the standard defence against curve fitting. The rule: optimise on one window of data, then test on the next window that the strategy has never seen. Then roll the windows forward and repeat.
Walk-forward testing — the protocol
01
Split the data into windows. For example: 6 months in-sample, 3 months out-of-sample. Then roll forward 3 months and repeat.
02
Optimise on the in-sample window only. Choose parameters based on the 6-month training data. Freeze them.
03
Test on the out-of-sample window. Run the frozen strategy on the next 3 months. Record the results.
04
Roll forward and repeat. New 6-month in-sample, new 3-month out-of-sample. Repeat 5–10 times.
05
Compare in-sample vs out-of-sample expectancy. If OOS is close to IS, the strategy is robust. If OOS is drastically lower, it is curve fitted.
06
The rule of thumb: OOS expectancy should be at least 60–70% of IS expectancy. Below 50%, the strategy does not survive the test.
WALK-FORWARD TESTING · ROLLING WINDOWS
Six in-sample / out-of-sample pairs · each OOS window is data the strategy has never seen
Six rolling tests. OOS expectancy at 80% of in-sample means the strategy is real, not curve fitted.
Worked example — the same strategy, two backtests, two live results
Setup
Range breakout with retest on H1
Instruments
EUR/USD, USD/JPY
Data period
3 years of H1 data, 2023–2026
Trades backtested
480
Live period
6 months, 150 live trades
Backtest A — Naive backtest.
Assumptions: 0 spread, 0 slippage, entry at signal close.
In-sample expectancy: +0.55R
What the trader expected to make: 150 × 0.55 = +82.5R
Backtest B — Realistic backtest.
Assumptions: 1.0 pip spread on majors, 1 pip slippage, entry at next candle open, walk-forward tested across 6 windows.
In-sample: +0.38R. Out-of-sample average: +0.30R.
Realistic expected: 150 × 0.30 = +45R
Live result (same 150 trades).
Realised expectancy: +0.26R
Actual: 150 × 0.26 = +39R
Naive estimate was 111% too high.
Realistic estimate was within 13% of live.
THE REALISTIC BACKTEST WAS THE ONLY USEFUL ONE.
Same strategy. Same market. Same trader. The difference was the cost model and the walk-forward test. The naive backtest was a fantasy. The realistic one was a plan.
The seven backtest rules
The rules — print these
01
Model realistic costs. Spread, commission, and swap. Never zero. Use the typical spread, not the best-case.
02
Model slippage. 1 pip on majors, more on crosses and around news. Assume stops slip against you.
03
Enter at the next candle's open. Not at the signal candle's close. Not at the exact level. At the realistic price.
04
Keep parameters low. One to three tuned inputs maximum. More than five is curve fitting.
05
Walk-forward test. Optimise on one window, test on the next. OOS expectancy should be at least 60–70% of IS.
06
Discount the backtest by 30–50% when setting expectations. If the backtest says +0.40R, plan for +0.20R to +0.28R live.
07
Gather 150+ live trades before concluding. The same sample-size rule from Lesson 48 applies to live results.
When this fails
When this fails
The market regime changed after the backtest. A strategy built for ranging markets will not survive a trending year. Backtest across multiple regimes, not just one.
The backtest data was not clean. Missing bars, incorrect ticks, or bad timestamps produce garbage results. Verify the data source before trusting the results.
The broker's live execution is worse than assumed. Some brokers slip more than others. Some widen spreads more aggressively. Run a 30-trade live test at minimum size before scaling up.
The trader's psychology interferes. A backtest does not hesitate, get greedy, or move a stop. A trader does. Even a perfect backtest will produce worse live results if the trader does not follow it.
The strategy relies on a rare event. A backtest that produces most of its profit in two outlier trades is not robust. Check the profit distribution — if removing the top 5% of trades turns the strategy negative, it is fragile.
The overfitting is undetectable. Some curve fitting survives walk-forward testing. The only defence is simplicity. The more parameters a strategy has, the more likely it is overfitted, even if the tests pass.
If you remember nothing else: the backtest is a hypothesis. The live account is the experiment. Discount the backtest by 30–50% before you commit capital, and gather 150+ live trades before you conclude anything.
In one box
Every backtest overstates. Always. No exceptions.
Five sources of decay: spread, slippage, look-ahead bias, curve fitting, execution quality.
Model spread, commission, swap. Never zero.
Model slippage. 1 pip on majors, more on stops.
Enter at next candle's open. Not signal candle close.
Keep parameters low. One to three tuned inputs.
Walk-forward test. OOS should be 60–70%+ of IS.
Discount by 30–50%. Plan for a fraction of the backtest.
150+ live trades before concluding.
Simplicity is robustness. Complex strategies fail more often.
See it in practice. Our free trading journal computes live expectancy against your backtested expectancy. See exactly how much decay your strategy is experiencing, in real time.
5 questions · immediate feedback · retake any time
Question 01 of 05
What happens to a backtest when you add realistic costs?
Correct: C. Realistic costs reduce expectancy. Spread, commission, and slippage all eat into the edge. Backtests that assume zero costs overstate expectancy by 0.05–0.15R per trade.
Question 02 of 05
What is look-ahead bias?
Correct: A. Look-ahead bias uses information you would not have had at the time. The most common form is entering at the signal candle's close — which you cannot do live.
Question 03 of 05
What is walk-forward testing?
Correct: D. Walk-forward testing optimises on one data window, freezes the parameters, and tests on the next window the strategy has never seen. This is the standard defence against curve fitting.
Question 04 of 05
Your backtest says +0.40R per trade. What should you plan for live?
Correct: B. Always discount a backtest by 30–50% when planning. A +0.40R backtest should be expected to produce +0.20R to +0.28R live.
Question 05 of 05
Why does curve fitting produce dangerous backtest results?
Correct: C. Curve fitting optimises parameters to fit historical data. The result is spectacular on the training data and negative on data the strategy has never seen. Walk-forward testing and parameter simplicity are the defences.
Block 7 complete. Block 8 — Psychology — begins. The journal is where strategy meets behaviour, and where every risk rule is either followed or broken.