Federated knowledge-guided 3D Transformer learning for bone tumor segmentation from partially annotated computed tomography data
Accurate 3D delineation of bone tumors and metastatic lesions from volumetric CT is essential for diagnosis, staging, radiotherapy planning, and treatment monitoring, yet manual contouring is time consuming and variable across observers. Although modern 3D CNN/Transformer segmenters achieve strong centralized performance, they typically assume pooled multi-institution data and consistent label definitions. In practice, clinical imaging archives are privacy-isolated, exhibit pronounced non-IID appearance shift (scanner/protocol and case-mix variation), and provide heterogeneous and partially annotated supervision (label-set mismatch and voxel-level incompleteness), where naive training and FedAvg-style aggregation can induce client drift and systematic false negatives. We propose federated knowledge-guided 3D Transformer segmentation (BoneFormer-KG) framework that explicitly addresses these constraints. BoneFormer-KG adapts a pretrained promptable backbone to volumetric CT using a parameter-efficient 3D pipeline with factorized in-plane/axial patch embedding, a learnable depth positional table, and lightweight 3D adapters to inject through-plane inductive bias while keeping memory and communication costs tractable. To handle label-set mismatch and partial voxel supervision, local optimization uses masked hybrid supervision (Dice + cross-entropy) computed only over annotated voxels and locally available labels, avoiding false-negative gradients on unobserved classes. On the server, we introduce effective labeled-signal weighting to aggregate client updates proportional to supervised voxel evidence, reducing bias under uneven annotation density. To further mitigate heterogeneity and supervision gaps, BoneFormer-KG integrates a coupled global-local distillation strategy: confidence-gated global distillation regularizes missing labels/unlabeled regions using the current global teacher, while class-sparse peer distillation transfers specialist knowledge from label-rich clients to label-poor clients with controlled bandwidth. An Auto Prompt Generator (APG) replaces manual prompting by synthesizing volumetric prompts from learned features for scalable prompt-free inference. We evaluate BoneFormer-KG on three CT datasets under non-IID federated splits and heterogeneous/partial supervision. BoneFormer-KG achieves competitive macro-averaged performance, reaching $$0.793\!\pm \!0.049$$ lesion DSC and $$7.4\!\pm \!2.7$$ mm HD95 on Spine-Mets-CT-SEG, $$0.845\!\pm \!0.065$$ BM DSC with $$5.6\!\pm \!2.1$$ mm BM HD95 (and $$0.964\!\pm \!0.014$$ bone DSC) on BM-Seg, and $$0.672\!\pm \!0.078$$ lesion DSC with $$11.3\!\pm \!4.1$$ mm HD95 and $$0.751\!\pm \!0.069$$ lesion $$F_1$$ on ULS23 Bone VOIs, outperforming the best baseline (UFPS) by 2.5% DSC on Spine-Mets-CT-SEG, 1.7% on BM-Seg, and 2.4% on ULS23 Bone baselines.
Authors
- 徐敬慈
- Mu Hu (ORCID: https://orcid.org/0000-0002-5566-1231)
- Shenyu Wang (ORCID: https://orcid.org/0000-0002-0173-1644)
- Xinchen Li (ORCID: https://orcid.org/0000-0002-4384-3572)
- Yongan Liu
- Ming Ni
- Dan Su
- Tao Yu
- Yuan Hong
- Liying Yang
Institutions
- Ruijin Hospital (CN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1038/s41598-026-71869-5
- Primary Topic
- Medical Image Segmentation Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00