Federated knowledge-guided 3D Transformer learning for bone tumor segmentation from partially annotated computed tomography data

Accurate 3D delineation of bone tumors and metastatic lesions from volumetric CT is essential for diagnosis, staging, radiotherapy planning, and treatment monitoring, yet manual contouring is time consuming and variable across observers. Although modern 3D CNN/Transformer segmenters achieve strong centralized performance, they typically assume pooled multi-institution data and consistent label definitions. In practice, clinical imaging archives are privacy-isolated, exhibit pronounced non-IID appearance shift (scanner/protocol and case-mix variation), and provide heterogeneous and partially annotated supervision (label-set mismatch and voxel-level incompleteness), where naive training and FedAvg-style aggregation can induce client drift and systematic false negatives. We propose federated knowledge-guided 3D Transformer segmentation (BoneFormer-KG) framework that explicitly addresses these constraints. BoneFormer-KG adapts a pretrained promptable backbone to volumetric CT using a parameter-efficient 3D pipeline with factorized in-plane/axial patch embedding, a learnable depth positional table, and lightweight 3D adapters to inject through-plane inductive bias while keeping memory and communication costs tractable. To handle label-set mismatch and partial voxel supervision, local optimization uses masked hybrid supervision (Dice + cross-entropy) computed only over annotated voxels and locally available labels, avoiding false-negative gradients on unobserved classes. On the server, we introduce effective labeled-signal weighting to aggregate client updates proportional to supervised voxel evidence, reducing bias under uneven annotation density. To further mitigate heterogeneity and supervision gaps, BoneFormer-KG integrates a coupled global-local distillation strategy: confidence-gated global distillation regularizes missing labels/unlabeled regions using the current global teacher, while class-sparse peer distillation transfers specialist knowledge from label-rich clients to label-poor clients with controlled bandwidth. An Auto Prompt Generator (APG) replaces manual prompting by synthesizing volumetric prompts from learned features for scalable prompt-free inference. We evaluate BoneFormer-KG on three CT datasets under non-IID federated splits and heterogeneous/partial supervision. BoneFormer-KG achieves competitive macro-averaged performance, reaching $$0.793\!\pm \!0.049$$ lesion DSC and $$7.4\!\pm \!2.7$$ mm HD95 on Spine-Mets-CT-SEG, $$0.845\!\pm \!0.065$$ BM DSC with $$5.6\!\pm \!2.1$$ mm BM HD95 (and $$0.964\!\pm \!0.014$$ bone DSC) on BM-Seg, and $$0.672\!\pm \!0.078$$ lesion DSC with $$11.3\!\pm \!4.1$$ mm HD95 and $$0.751\!\pm \!0.069$$ lesion $$F_1$$ on ULS23 Bone VOIs, outperforming the best baseline (UFPS) by 2.5% DSC on Spine-Mets-CT-SEG, 1.7% on BM-Seg, and 2.4% on ULS23 Bone baselines.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-10-05
DOI
https://doi.org/10.1038/s41598-026-71869-5
Primary Topic
Medical Image Segmentation Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Federated knowledge-guided 3D Transformer learning for bone tumor segmentation from partially annotated computed tomography data

徐敬慈, Mu Hu, Shenyu Wang, Xinchen Li et al.
Scientific Reports
Medical Image Segmentation Techniques
article

Federated knowledge-guided 3D Transformer learning for bone tumor segmentation from partially annotated computed tomography data

徐敬慈, Mu Hu, Shenyu Wang, Xinchen Li, Yongan Liu, Ming Ni, Dan Su, Tao Yu, Yuan Hong, Liying Yang
article en

Abstract

Accurate 3D delineation of bone tumors and metastatic lesions from volumetric CT is essential for diagnosis, staging, radiotherapy planning, and treatment monitoring, yet manual contouring is time consuming and variable across observers. Although modern 3D CNN/Transformer segmenters achieve strong centralized performance, they typically assume pooled multi-institution data and consistent label definitions. In practice, clinical imaging archives are privacy-isolated, exhibit pronounced non-IID appearance shift (scanner/protocol and case-mix variation), and provide heterogeneous and partially annotated supervision (label-set mismatch and voxel-level incompleteness), where naive training and FedAvg-style aggregation can induce client drift and systematic false negatives. We propose federated knowledge-guided 3D Transformer segmentation (BoneFormer-KG) framework that explicitly addresses these constraints. BoneFormer-KG adapts a pretrained promptable backbone to volumetric CT using a parameter-efficient 3D pipeline with factorized in-plane/axial patch embedding, a learnable depth positional table, and lightweight 3D adapters to inject through-plane inductive bias while keeping memory and communication costs tractable. To handle label-set mismatch and partial voxel supervision, local optimization uses masked hybrid supervision (Dice + cross-entropy) computed only over annotated voxels and locally available labels, avoiding false-negative gradients on unobserved classes. On the server, we introduce effective labeled-signal weighting to aggregate client updates proportional to supervised voxel evidence, reducing bias under uneven annotation density. To further mitigate heterogeneity and supervision gaps, BoneFormer-KG integrates a coupled global-local distillation strategy: confidence-gated global distillation regularizes missing labels/unlabeled regions using the current global teacher, while class-sparse peer distillation transfers specialist knowledge from label-rich clients to label-poor clients with controlled bandwidth. An Auto Prompt Generator (APG) replaces manual prompting by synthesizing volumetric prompts from learned features for scalable prompt-free inference. We evaluate BoneFormer-KG on three CT datasets under non-IID federated splits and heterogeneous/partial supervision. BoneFormer-KG achieves competitive macro-averaged performance, reaching $$0.793\!\pm \!0.049$$ lesion DSC and $$7.4\!\pm \!2.7$$ mm HD95 on Spine-Mets-CT-SEG, $$0.845\!\pm \!0.065$$ BM DSC with $$5.6\!\pm \!2.1$$ mm BM HD95 (and $$0.964\!\pm \!0.014$$ bone DSC) on BM-Seg, and $$0.672\!\pm \!0.078$$ lesion DSC with $$11.3\!\pm \!4.1$$ mm HD95 and $$0.751\!\pm \!0.069$$ lesion $$F_1$$ on ULS23 Bone VOIs, outperforming the best baseline (UFPS) by 2.5% DSC on Spine-Mets-CT-SEG, 1.7% on BM-Seg, and 2.4% on ULS23 Bone baselines.

Scientific Reports
Ruijin Hospital (CN)
Good health and well-being
Openalex Percentile: Top 15%
Medical Image Segmentation Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.