When Backtests Agree but the Data Doesn't
Machine learning in algorithmic trading is commonly evaluated after a historical dataset has already been constructed. Considerable attention is then given to model choice, feature selection, chronological train/test separation, walk-forward validation, and tests for future leakage or backtest overfitting. These controls are necessary, but they do not by themselves establish that the model received the same decision-time representation that could have existed in live operation. Vortraq began as a private study of geometric relationships in market behaviour; machine learning was introduced later. Building a historical and live system that produced the same decision-time state proved substantially harder than training the models themselves. The work involved repeated causal, timing, association, reconstruction and state-consistency corrections. Some apparently useful machine-learning relationships weakened as those defects were removed. A traced forming-candle defect, for example, reduced pooled AUC from 0.731 to 0.669 across successive corrections, while top-decile profit factor fell from 1.33 to 1.00. This paper uses observational parity as an umbrella term for a stricter engineering requirement. Causal availability asks whether information could exist at decision time; observation identity asks whether records refer to the same underlying decision; live/offline observational parity asks whether an offline reconstruction reproduces the decision-time representation recorded by the live system. In Vortraq this became a release invariant rather than a one-time certificate: an unexplained deterministic divergence is corrected before the affected implementation is accepted as a new research base, or the change is reverted. A deliberate certified experiment matched 1,055 live candidate observations unambiguously to 1,055 offline reconstructions and produced zero differences across 92 certified fields. Later testing broadened the exercised surface. In a VORT88 comparison spanning 15,729 bars and three simultaneously trading timeframes, 10,925 shared candidates were compared across 123 shared fields. All 1,183,759 individual value comparisons were exactly equal; no tolerance was required. These results establish exact reconstruction only for the tested populations and surfaces. They do not establish profitability, venue execution parity, or universal correctness of logic shared by both paths.
Authors
- Tadas Siksnius
Institutions
- VORtech (Netherlands) (NL)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22760801
- Primary Topic
- Sports Analytics and Performance
- Type
- article
- Field-Weighted Citation Impact
- 0.00