Replicability of agent-based simulations: a case study with a SEIRS epidemiological model

Reproducibility is not only a cornerstone of the advancement of science, but also of scientific credibility. However, it remains a persistent challenge across many disciplines, and it is now referred to as the “reproducibility crisis.” This issue is particularly relevant in computational modeling. In this paper, we want to know if the usage of different tools and implementation choices can lead to divergent outcomes. We investigate the replicability of a SEIRS epidemiological model implemented via agent-based modeling (ABM) across seven different platforms and programming languages, namely C++, Julia, Python, NetLogo, GAMA, Cormas, and PythonPDEVS. Each implementation was based on a common formal specification using the Overview, Design concepts and Details (ODD) protocol and developed independently by different modelers. Our results show that while all implementations qualitatively reproduce the expected dynamics of the model, a sharp infection peak followed by damped oscillations, they produce statistically significant differences in key metrics such as peak amplitude and timing. These differences are also large in magnitude: the variation between implementations is about 7.6 times larger than the stochastic variation within an implementation (across its 30 replications). Because each implementation was produced by a different modeler on a different platform, this variance reflects the combined effect of the tool and the modeler, which are confounded in the present design and cannot be separated. A complementary analysis of several independent developers working in a single language (C++ or Java) shows that the modeler alone induces statistically significant and large differences, confirming that the individual developer is an important source of divergence; the specific contribution of the technology therefore remains open. We conclude that even when a model is formally specified, independently developed ABM implementations can yield divergent numerical results. A controlled ablation traces much of this divergence to two under-specified execution details, namely how the residence times are discretised into whole days and whether the state-transition test is strict; specifying them in the ODD would remove much of the divergence. Our findings highlight the need for rigorous reproducibility and replicability practices, and for model specifications precise enough to pin down such implementation choices.

Authors

Institutions

Publication Details

Journal
Complex & Intelligent Systems
Published
2026-10-09
DOI
https://doi.org/10.1007/s40747-026-02546-3
Primary Topic
Simulation Techniques and Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Replicability of agent-based simulations: a case study with a SEIRS epidemiological model

Étienne Delay, Clément Foucher, Raphaël Duboz, Claude Mazel et al.
Complex & Intelligent Systems
Simulation Techniques and Applications
article

Replicability of agent-based simulations: a case study with a SEIRS epidemiological model

Étienne Delay, Clément Foucher, Raphaël Duboz, Claude Mazel, Benjamin Antunes, Paul-Antoine Bisgambiglia, David R.C. Hill, Arthur Scriban, Hanae Havion, Oleksandr Zaitsev
article en

Abstract

Reproducibility is not only a cornerstone of the advancement of science, but also of scientific credibility. However, it remains a persistent challenge across many disciplines, and it is now referred to as the “reproducibility crisis.” This issue is particularly relevant in computational modeling. In this paper, we want to know if the usage of different tools and implementation choices can lead to divergent outcomes. We investigate the replicability of a SEIRS epidemiological model implemented via agent-based modeling (ABM) across seven different platforms and programming languages, namely C++, Julia, Python, NetLogo, GAMA, Cormas, and PythonPDEVS. Each implementation was based on a common formal specification using the Overview, Design concepts and Details (ODD) protocol and developed independently by different modelers. Our results show that while all implementations qualitatively reproduce the expected dynamics of the model, a sharp infection peak followed by damped oscillations, they produce statistically significant differences in key metrics such as peak amplitude and timing. These differences are also large in magnitude: the variation between implementations is about 7.6 times larger than the stochastic variation within an implementation (across its 30 replications). Because each implementation was produced by a different modeler on a different platform, this variance reflects the combined effect of the tool and the modeler, which are confounded in the present design and cannot be separated. A complementary analysis of several independent developers working in a single language (C++ or Java) shows that the modeler alone induces statistically significant and large differences, confirming that the individual developer is an important source of divergence; the specific contribution of the technology therefore remains open. We conclude that even when a model is formally specified, independently developed ABM implementations can yield divergent numerical results. A controlled ablation traces much of this divergence to two under-specified execution details, namely how the residence times are discretised into whole days and whether the state-transition test is strict; specifying them in the ODD would remove much of the divergence. Our findings highlight the need for rigorous reproducibility and replicability practices, and for model specifications precise enough to pin down such implementation choices.

Complex & Intelligent Systems
Centre National de la Recherche Scientifique (FR), Centre de Coopération Internationale en Recherche Agronomique pour le Développement (FR), Université Toulouse III - Paul Sabatier (FR), Université de Corse Pascal Paoli (FR), Laboratoire d'Analyse et d'Architecture des Systèmes (FR), Université de Montpellier (FR), Franche-Comté Électronique Mécanique Thermique et Optique - Sciences et Technologies (FR), Sorbonne Université (FR), Institut National de Recherche pour l'Agriculture, l'Alimentation et l'Environnement (FR), Animal, Santé, Territoires, Risques et Ecosystèmes (FR), Institut de Recherche pour le Développement (SN), Unité de Modélisation Mathématique et Informatique des Systèmes Complexes (FR), Institut de Recherche pour le Développement (FR), Sciences pour l'Environnement (FR), Université de Toulouse (FR), Cheikh Anta Diop University (SN)
Openalex Percentile: Top 10%
Simulation Techniques and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.