SemioKAT-FL: explainable, background-invariant, and privacy-preserving video seizure detection with a Kolmogorov–Arnold Transformer and federated learning
Abstract Automatic detection of epileptic seizures from video promises non-invasive, continuous patient monitoring, yet web-sourced seizure corpora are dominated by background confounds : recording-specific cues (studio, home, clinic) that are spuriously correlated with the seizure label. Models that appear highly accurate under conventional random splits may therefore be recognizing studios rather than seizures , and they remain opaque and dependent on centralizing sensitive patient video. We present SemioKAT-FL, a framework that addresses all three issues jointly. Spatial features from an ImageNet-pretrained MobileNetV2 encoder are temporally modeled by a Kolmogorov–Arnold Transformer (KAT) that couples multi-head self-attention with learnable B-spline (KAN) feed-forward layers; a temporal-attention pooling layer yields per-second saliency; and an adversarial background-invariance head, driven by a gradient-reversal layer, penalizes any ability to recover the source recording. Training uses Federated Averaging so that raw video never leaves a site. On a corpus of 1,148 clips from 275 source recordings evaluated under leakage-controlled subject-wise (GroupKFold) cross-validation, we first show that a faithful re-implementation of the reference method drops from 99.5% under random splits to 94.9% subject-wise, and that a background-only model still reaches 90.4%—direct evidence of confounding. The adversarial head substantially reduces a source-recovery probe from 70% to 50% (chance $$\\approx 1.3\\%$$ ); the representation retains reduced but non-negligible source information, which we characterize as partial rather than complete invariance. The full calibrated ensemble attains $$96.6\\%\\pm 1.8\\%$$ accuracy (F1 $$=96.8\\%$$ , AUC $$=0.93$$ , expected calibration error 0.02–0.03) under source-level GroupKFold, and the federated model—trained on simulated site partitions without formal differential privacy—matches its centralized counterpart within 0.4%. To our knowledge, SemioKAT-FL is the first video seizure detector that jointly addresses background confounding through adversarial invariance, provides intrinsic explainability, and avoids centralization of raw video through federated training.
Authors
- Mukta Dhopeshwarkar (ORCID: https://orcid.org/0009-0000-4856-8533)
- Mohammed Tawfik (ORCID: https://orcid.org/0000-0002-1227-387X)
- Alhasan A. Alharbi
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-22
- DOI
- https://doi.org/10.1038/s41598-026-71827-1
- Primary Topic
- EEG and Brain-Computer Interfaces
- Type
- article
- Field-Weighted Citation Impact
- 0.00