When Backtests Agree but the Data Doesn't

Machine learning in algorithmic trading is commonly evaluated after a historical dataset has already been constructed. Considerable attention is then given to model choice, feature selection, chronological train/test separation, walk-forward validation, and tests for future leakage or backtest overfitting. These controls are necessary, but they do not by themselves establish that the model received the same decision-time representation that could have existed in live operation. Vortraq began as a private study of geometric relationships in market behaviour; machine learning was introduced later. Building a historical and live system that produced the same decision-time state proved substantially harder than training the models themselves. The work involved repeated causal, timing, association, reconstruction and state-consistency corrections. Some apparently useful machine-learning relationships weakened as those defects were removed. A traced forming-candle defect, for example, reduced pooled AUC from 0.731 to 0.669 across successive corrections, while top-decile profit factor fell from 1.33 to 1.00. This paper uses observational parity as an umbrella term for a stricter engineering requirement. Causal availability asks whether information could exist at decision time; observation identity asks whether records refer to the same underlying decision; live/offline observational parity asks whether an offline reconstruction reproduces the decision-time representation recorded by the live system. In Vortraq this became a release invariant rather than a one-time certificate: an unexplained deterministic divergence is corrected before the affected implementation is accepted as a new research base, or the change is reverted. A deliberate certified experiment matched 1,055 live candidate observations unambiguously to 1,055 offline reconstructions and produced zero differences across 92 certified fields. Later testing broadened the exercised surface. In a VORT88 comparison spanning 15,729 bars and three simultaneously trading timeframes, 10,925 shared candidates were compared across 123 shared fields. All 1,183,759 individual value comparisons were exactly equal; no tolerance was required. These results establish exact reconstruction only for the tested populations and surfaces. They do not establish profitability, venue execution parity, or universal correctness of logic shared by both paths.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-15
DOI
https://doi.org/10.5281/zenodo.22760801
Primary Topic
Sports Analytics and Performance
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

When Backtests Agree but the Data Doesn't

Tadas Siksnius
Zenodo (CERN European Organization for Nuclear Research)
Sports Analytics and Performance
article

When Backtests Agree but the Data Doesn't

Tadas Siksnius
article en

Abstract

Machine learning in algorithmic trading is commonly evaluated after a historical dataset has already been constructed. Considerable attention is then given to model choice, feature selection, chronological train/test separation, walk-forward validation, and tests for future leakage or backtest overfitting. These controls are necessary, but they do not by themselves establish that the model received the same decision-time representation that could have existed in live operation. Vortraq began as a private study of geometric relationships in market behaviour; machine learning was introduced later. Building a historical and live system that produced the same decision-time state proved substantially harder than training the models themselves. The work involved repeated causal, timing, association, reconstruction and state-consistency corrections. Some apparently useful machine-learning relationships weakened as those defects were removed. A traced forming-candle defect, for example, reduced pooled AUC from 0.731 to 0.669 across successive corrections, while top-decile profit factor fell from 1.33 to 1.00. This paper uses observational parity as an umbrella term for a stricter engineering requirement. Causal availability asks whether information could exist at decision time; observation identity asks whether records refer to the same underlying decision; live/offline observational parity asks whether an offline reconstruction reproduces the decision-time representation recorded by the live system. In Vortraq this became a release invariant rather than a one-time certificate: an unexplained deterministic divergence is corrected before the affected implementation is accepted as a new research base, or the change is reverted. A deliberate certified experiment matched 1,055 live candidate observations unambiguously to 1,055 offline reconstructions and produced zero differences across 92 certified fields. Later testing broadened the exercised surface. In a VORT88 comparison spanning 15,729 bars and three simultaneously trading timeframes, 10,925 shared candidates were compared across 123 shared fields. All 1,183,759 individual value comparisons were exactly equal; no tolerance was required. These results establish exact reconstruction only for the tested populations and surfaces. They do not establish profitability, venue execution parity, or universal correctness of logic shared by both paths.

Zenodo (CERN European Organization for Nuclear Research)
VORtech (Netherlands) (NL)
Peace, Justice and strong institutions
Openalex Percentile: Top 5%
Sports Analytics and Performance
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.