SemioKAT-FL: explainable, background-invariant, and privacy-preserving video seizure detection with a Kolmogorov–Arnold Transformer and federated learning

Abstract Automatic detection of epileptic seizures from video promises non-invasive, continuous patient monitoring, yet web-sourced seizure corpora are dominated by background confounds : recording-specific cues (studio, home, clinic) that are spuriously correlated with the seizure label. Models that appear highly accurate under conventional random splits may therefore be recognizing studios rather than seizures , and they remain opaque and dependent on centralizing sensitive patient video. We present SemioKAT-FL, a framework that addresses all three issues jointly. Spatial features from an ImageNet-pretrained MobileNetV2 encoder are temporally modeled by a Kolmogorov–Arnold Transformer (KAT) that couples multi-head self-attention with learnable B-spline (KAN) feed-forward layers; a temporal-attention pooling layer yields per-second saliency; and an adversarial background-invariance head, driven by a gradient-reversal layer, penalizes any ability to recover the source recording. Training uses Federated Averaging so that raw video never leaves a site. On a corpus of 1,148 clips from 275 source recordings evaluated under leakage-controlled subject-wise (GroupKFold) cross-validation, we first show that a faithful re-implementation of the reference method drops from 99.5% under random splits to 94.9% subject-wise, and that a background-only model still reaches 90.4%—direct evidence of confounding. The adversarial head substantially reduces a source-recovery probe from 70% to 50% (chance $$\\approx 1.3\\%$$ ); the representation retains reduced but non-negligible source information, which we characterize as partial rather than complete invariance. The full calibrated ensemble attains $$96.6\\%\\pm 1.8\\%$$ accuracy (F1 $$=96.8\\%$$ , AUC $$=0.93$$ , expected calibration error 0.02–0.03) under source-level GroupKFold, and the federated model—trained on simulated site partitions without formal differential privacy—matches its centralized counterpart within 0.4%. To our knowledge, SemioKAT-FL is the first video seizure detector that jointly addresses background confounding through adversarial invariance, provides intrinsic explainability, and avoids centralization of raw video through federated training.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-09-22
DOI
https://doi.org/10.1038/s41598-026-71827-1
Primary Topic
EEG and Brain-Computer Interfaces
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SemioKAT-FL: explainable, background-invariant, and privacy-preserving video seizure detection with a Kolmogorov–Arnold Transformer and federated learning

Mukta Dhopeshwarkar, Mohammed Tawfik, Alhasan A. Alharbi
Scientific Reports
EEG and Brain-Computer Interfaces
article

SemioKAT-FL: explainable, background-invariant, and privacy-preserving video seizure detection with a Kolmogorov–Arnold Transformer and federated learning

Mukta Dhopeshwarkar, Mohammed Tawfik, Alhasan A. Alharbi
article en

Abstract

Abstract Automatic detection of epileptic seizures from video promises non-invasive, continuous patient monitoring, yet web-sourced seizure corpora are dominated by background confounds : recording-specific cues (studio, home, clinic) that are spuriously correlated with the seizure label. Models that appear highly accurate under conventional random splits may therefore be recognizing studios rather than seizures , and they remain opaque and dependent on centralizing sensitive patient video. We present SemioKAT-FL, a framework that addresses all three issues jointly. Spatial features from an ImageNet-pretrained MobileNetV2 encoder are temporally modeled by a Kolmogorov–Arnold Transformer (KAT) that couples multi-head self-attention with learnable B-spline (KAN) feed-forward layers; a temporal-attention pooling layer yields per-second saliency; and an adversarial background-invariance head, driven by a gradient-reversal layer, penalizes any ability to recover the source recording. Training uses Federated Averaging so that raw video never leaves a site. On a corpus of 1,148 clips from 275 source recordings evaluated under leakage-controlled subject-wise (GroupKFold) cross-validation, we first show that a faithful re-implementation of the reference method drops from 99.5% under random splits to 94.9% subject-wise, and that a background-only model still reaches 90.4%—direct evidence of confounding. The adversarial head substantially reduces a source-recovery probe from 70% to 50% (chance $$\approx 1.3\%$$ ); the representation retains reduced but non-negligible source information, which we characterize as partial rather than complete invariance. The full calibrated ensemble attains $$96.6\%\pm 1.8\%$$ accuracy (F1 $$=96.8\%$$ , AUC $$=0.93$$ , expected calibration error 0.02–0.03) under source-level GroupKFold, and the federated model—trained on simulated site partitions without formal differential privacy—matches its centralized counterpart within 0.4%. To our knowledge, SemioKAT-FL is the first video seizure detector that jointly addresses background confounding through adversarial invariance, provides intrinsic explainability, and avoids centralization of raw video through federated training.

Scientific Reports
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
EEG and Brain-Computer Interfaces
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.