Heavy-tail-aware representation learning and dynamic Bayesian state modelling to derive an operational proxy definition of problem gambling risk from routine online gambling data

Abstract Background: Problem gambling causes harm, but operational identification often relies on heuristic thresholds or sparse manual reviews. Routine online gambling logs are heavy-tailed and temporally structured, complicating risk definition and early detection. Methods: We analysed de-identified records from an online gambling operator across four streams (transactions, bets, sessions, payments). Time series were summarised into leakage-audited 30-day windows with heavy-tail-aware exceedance frequency and magnitude features. Window embeddings were learned using a hierarchical conditional variational autoencoder: a teacher trained on responsible-gambling proxy signals, then a student fine-tuned on sparse manual analyst assessments on the training split only. To address missing-not-at-random assessments, backlog-aware label inference conservatively augmented training data. Dynamic regimes were inferred from embeddings using a regularised Gaussian hidden Markov model, yielding a three-class operational proxy definition. Agreement with analyst assessments and early-warning utility under explicit capacity constraints were evaluated on held-out labels. Results: Balanced accuracy on labelled test windows ranged from 0.38 (transactions) to 0.62 (payments), with best macro-averaged F1-score in bets (0.54). Under a capacity-constrained top-10-per-week queue, escalation detection ranged from 0.39 (sessions) to 0.62 (bets), with median lead times of 42–290 days. Conclusions: Heavy-tail-aware representations combined with dynamic regime modelling can derive an auditable operational proxy definition of gambling-related risk from routine data and support realistic, capacity-constrained monitoring.

Authors

Institutions

Publication Details

Journal
EPJ Data Science
Published
2026-09-04
DOI
https://doi.org/10.1140/epjds/s13688-026-00698-3
Primary Topic
Gambling Behavior and Treatments
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Heavy-tail-aware representation learning and dynamic Bayesian state modelling to derive an operational proxy definition of problem gambling risk from routine online gambling data

Sam Andersson, Philip Lindner, Olof Molander, Helga Westerlind et al.
EPJ Data Science
Gambling Behavior and Treatments
article

Heavy-tail-aware representation learning and dynamic Bayesian state modelling to derive an operational proxy definition of problem gambling risk from routine online gambling data

Sam Andersson, Philip Lindner, Olof Molander, Helga Westerlind, Keenan Lyon, Timo Koski, P. Carlbring
article en

Abstract

Abstract Background: Problem gambling causes harm, but operational identification often relies on heuristic thresholds or sparse manual reviews. Routine online gambling logs are heavy-tailed and temporally structured, complicating risk definition and early detection. Methods: We analysed de-identified records from an online gambling operator across four streams (transactions, bets, sessions, payments). Time series were summarised into leakage-audited 30-day windows with heavy-tail-aware exceedance frequency and magnitude features. Window embeddings were learned using a hierarchical conditional variational autoencoder: a teacher trained on responsible-gambling proxy signals, then a student fine-tuned on sparse manual analyst assessments on the training split only. To address missing-not-at-random assessments, backlog-aware label inference conservatively augmented training data. Dynamic regimes were inferred from embeddings using a regularised Gaussian hidden Markov model, yielding a three-class operational proxy definition. Agreement with analyst assessments and early-warning utility under explicit capacity constraints were evaluated on held-out labels. Results: Balanced accuracy on labelled test windows ranged from 0.38 (transactions) to 0.62 (payments), with best macro-averaged F1-score in bets (0.54). Under a capacity-constrained top-10-per-week queue, escalation detection ranged from 0.39 (sessions) to 0.62 (bets), with median lead times of 42–290 days. Conclusions: Heavy-tail-aware representations combined with dynamic regime modelling can derive an auditable operational proxy definition of gambling-related risk from routine data and support realistic, capacity-constrained monitoring.

EPJ Data Science
National Institute of Economic Research (SE), Stockholm University (SE), Korea University (KR), Karolinska Institutet (SE), Stockholm Health Care Services (SE), KTH Royal Institute of Technology (SE)
Stockholms Universitet
Peace, Justice and strong institutions
Openalex Percentile: Top 82%
Gambling Behavior and Treatments
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.