1 · The problem
A persuasive backtest can still be wrong
A historical result describes what a particular configuration did
against a particular dataset under particular assumptions. It does not,
by itself, establish what will happen next.
Selection bias
If enough configurations are tested, some will look successful by
chance. Selecting the best result without accounting for the wider
trial population can turn luck into an apparently compelling
strategy.
Overfitting
A configuration can become too closely adapted to the precise
historical period used to find it. It may describe past noise
exceptionally well while failing on new data.
Information leakage
A test becomes unreliable if information that would not have been
available at the decision point accidentally influences the signal,
parameter selection or validation process.
Execution assumptions
Attractive results may disappear when spread, slippage, latency,
financing, fill behaviour and other operating frictions are more
realistically represented.
Regime dependence
A candidate may work only during a particular volatility, trend or
liquidity regime. Aggregate performance can conceal long weak
periods or dependence on a small number of favourable episodes.
Insufficient observations
An extreme profit factor or smooth equity curve can be produced by
very few trades. Small samples provide little evidence about the
range of outcomes the strategy may encounter.
No single metric constitutes proof.
Return, profit factor, Sharpe ratio, drawdown and the equity curve each
describe part of the result. None establishes that the apparent edge is
repeatable, operationally viable or likely to survive new market data.
The hypothesis under examination
The candidate considered here belongs to a generic intraday
mean-reversion strategy family operating in EUR/USD on a 30-minute
signal horizon.
The first-phase implementation is deliberately bounded. Entries are
evaluated from 30-minute market data and, once a position is opened,
static profit-target and stop-loss levels govern the trade. A possible
later research phase may examine tick-based stop management informed by
maximum favourable excursion and maximum adverse excursion evidence.
That later work is not part of the result documented here.
A governed evidence sequence
Lucitech treats optimisation as the beginning of an investigation,
rather than its conclusion. A candidate must move through a recorded
sequence of evidence gates.
01
Discover
Structured Optuna search evaluated through the JForex historical
tester.
02
Validate
Reproduction, out-of-sample testing, population analysis and
robustness challenges.
03
Approve
Evidence is reviewed against explicit gates before a candidate can
progress.
04
Deploy
Controlled shadow-demo observation precedes or accompanies
tightly bounded live operation.
05
Observe
Runtime health, trades, execution parity and evidence drift remain
under telemetry.
Figure 1 — Lucitech’s evidence pipeline records whether a candidate
advances, stops or returns for further investigation at each gate.