Leakage-Free Operational Earthquake Forecasting: A Multi-Region CSEP Testbed with Pre-Registered Negative Results and the Horizon-Dependent Value of Geodetic Context

Operational earthquake forecasting (OEF) issues calibrated conditional probabilities of future seismicity; the Epidemic-Type Aftershock Sequence (ETAS) model is its de-facto benchmark, and under fair, prospective, CSEP-style testing no machine-learning temporal point process has been shown to beat a well-fit ETAS. We describe a complete OEF system and, on it, a leakage-free, multi-region CSEP testbed designed to test four questions empirically. The system fits a regime-tiled space-time ETAS with a full hygiene pipeline (rolling magnitude of completeness, moment-magnitude homogenization, dual-catalog declustering, propagated uncertainty), against a mandatory adaptive smoothed-seismicity Poisson null and a transparent Reasenberg-Jones fallback, calibrated by isotonic regression with an epistemic-plus-aleatory uncertainty triad, and it emits both gridded and catalog-based forecast representations so that over-dispersion is accounted for in scoring. Skill is established only by winning CSEP comparison tests (information gain per earthquake, IGPE, in nats) against both baselines, under a strict forecast clock that guards against five leakage modes, with adoption rules fixed before each run and shuffled-label negative controls. Four findings result. (i) ETAS adds skill over the stationary null only where there is triggering to exploit (Japan +0.072 nats at one day; low-seismicity interiors exactly 0.0), quantifying the high-versus-low-seismicity bias. (ii) A convex log-score-optimal stack of ETAS-family variants earns a significant global gain (+0.011 over 751 events) that does not generalize to the canonical active margins, so the pre-registered adoption rule rejects it; a temporally adaptive variant that appears to rescue it is a multiple-comparison artifact. (iii) The binding one-to-seven-day consistency failure is count over-dispersion, not spatial shape: a catalog-based number test passes a window the Poisson number test rejects, and the frozen-intensity forecast systematically under-counts by about 28 percent because it omits within-window secondary triggering. (iv) The value of a GNSS-strain geodetic context covariate, added to a Hawkes-structured neural point process, is horizon-dependent: it does not beat ETAS at the one-to-seven-day operational horizon (mean IGPE -0.053 over eight weekly windows) but beats it robustly at thirty days (global +0.115 over 2166 events, positive in 9/10 windows and in every high-seismicity region), because the calibrated model deploys a time-flat geodetic background rather than triggering. We conclude that base tiled ETAS is at or near the practical ceiling for mean-rate one-to-seven-day IGPE over global M>=5, and that a geodetic covariate belongs in a longer, background-dominated outlook. This is an independent research and education tool; it is not an operational alarm system and must not be used for life-safety decisions. Code, configurations, provenance manifests and artifacts (MIT): https://github.com/fsantibanezleal/CAOS_SEISMIC . Static forecast viewer: https://seismic.fasl-work.com .

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-18
DOI
https://doi.org/10.5281/zenodo.21508362
Primary Topic
earthquake and tectonic studies
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Leakage-Free Operational Earthquake Forecasting: A Multi-Region CSEP Testbed with Pre-Registered Negative Results and the Horizon-Dependent Value of Geodetic Context

Felipe Santibañez-Leal
Zenodo (CERN European Organization for Nuclear Research)
earthquake and tectonic studies
preprint

Leakage-Free Operational Earthquake Forecasting: A Multi-Region CSEP Testbed with Pre-Registered Negative Results and the Horizon-Dependent Value of Geodetic Context

Felipe Santibañez-Leal
preprint en

Abstract

Operational earthquake forecasting (OEF) issues calibrated conditional probabilities of future seismicity; the Epidemic-Type Aftershock Sequence (ETAS) model is its de-facto benchmark, and under fair, prospective, CSEP-style testing no machine-learning temporal point process has been shown to beat a well-fit ETAS. We describe a complete OEF system and, on it, a leakage-free, multi-region CSEP testbed designed to test four questions empirically. The system fits a regime-tiled space-time ETAS with a full hygiene pipeline (rolling magnitude of completeness, moment-magnitude homogenization, dual-catalog declustering, propagated uncertainty), against a mandatory adaptive smoothed-seismicity Poisson null and a transparent Reasenberg-Jones fallback, calibrated by isotonic regression with an epistemic-plus-aleatory uncertainty triad, and it emits both gridded and catalog-based forecast representations so that over-dispersion is accounted for in scoring. Skill is established only by winning CSEP comparison tests (information gain per earthquake, IGPE, in nats) against both baselines, under a strict forecast clock that guards against five leakage modes, with adoption rules fixed before each run and shuffled-label negative controls. Four findings result. (i) ETAS adds skill over the stationary null only where there is triggering to exploit (Japan +0.072 nats at one day; low-seismicity interiors exactly 0.0), quantifying the high-versus-low-seismicity bias. (ii) A convex log-score-optimal stack of ETAS-family variants earns a significant global gain (+0.011 over 751 events) that does not generalize to the canonical active margins, so the pre-registered adoption rule rejects it; a temporally adaptive variant that appears to rescue it is a multiple-comparison artifact. (iii) The binding one-to-seven-day consistency failure is count over-dispersion, not spatial shape: a catalog-based number test passes a window the Poisson number test rejects, and the frozen-intensity forecast systematically under-counts by about 28 percent because it omits within-window secondary triggering. (iv) The value of a GNSS-strain geodetic context covariate, added to a Hawkes-structured neural point process, is horizon-dependent: it does not beat ETAS at the one-to-seven-day operational horizon (mean IGPE -0.053 over eight weekly windows) but beats it robustly at thirty days (global +0.115 over 2166 events, positive in 9/10 windows and in every high-seismicity region), because the calibrated model deploys a time-flat geodetic background rather than triggering. We conclude that base tiled ETAS is at or near the practical ceiling for mean-rate one-to-seven-day IGPE over global M>=5, and that a geodetic covariate belongs in a longer, background-dominated outlook. This is an independent research and education tool; it is not an operational alarm system and must not be used for life-safety decisions. Code, configurations, provenance manifests and artifacts (MIT): https://github.com/fsantibanezleal/CAOS_SEISMIC . Static forecast viewer: https://seismic.fasl-work.com .

Zenodo (CERN European Organization for Nuclear Research)
Open University of Cyprus (CY)
Good health and well-being
earthquake and tectonic studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.