Knowledge-based anomaly detection for identifying network-induced shape artifacts

Synthetic data provides a promising approach to address data scarcity for training machine learning models; however, adoption without proper quality assessments may introduce artifacts, distortions, and unrealistic features that compromise model performance and downstream utility for medical imaging research and evaluation purposes. This work introduces a novel knowledge-based anomaly detection method for detecting network-induced shape artifacts in synthetic images. The introduced method utilizes a two-stage approach comprising (i) a novel feature extractor that constructs a specialized feature space by analyzing the per-image distribution of angle gradients along anatomical boundaries, and (ii) an isolation forest-based anomaly detector. This representation captures local shape variations while maintaining global anatomical correspondence regardless of size differences. The isolation forest is trained on patient data to learn normal anatomical shape characteristics and assigns anomaly scores to synthetic images. This method may be used to detect anatomically unrealistic images irrespective of the generative model used and provides interpretability through its knowledge-based design. We demonstrate the effectiveness of the method for identifying network-induced shape artifacts in two synthetic mammography datasets generated by different architectures: a latent diffusion model and StyleGAN2, trained on CSAW-M and VinDr-Mammo patient datasets respectively. Quantitative evaluation shows that the method successfully concentrates artifacts in the most anomalous partition (1st percentile), with AUC values of 0.97 (CSAW-syn) and 0.91 (VMLO-syn). Comparison with anomaly detection based on other standard shape descriptors was used to demonstrate the superiority of the proposed method at isolating subtle shape artifacts along with extreme ones. In addition, a reader study involving three imaging scientists confirmed that images identified by the method as containing network-induced shape artifacts were also flagged by human readers with mean agreement rates of 66% (CSAW-syn) and 68% (VMLO-syn) for the most anomalous partition, approximately 1.5-2 times higher than the least anomalous partition. Kendall-Tau correlations between algorithmic and human rankings were 0.45 and 0.43 for the two datasets, indicating reasonable agreement despite the challenging nature of subtle artifact detection. This method is a step forward in supporting quality assurance practices for synthetic data in research and development contexts, as it allows developers to evaluate synthetic images for known anatomic constraints and pinpoint and address specific issues to improve the overall quality of a synthetic dataset. This method is presented as a research tool for synthetic image quality evaluation. Its application in clinical or regulatory decision-making contexts would require additional validation appropriate to the intended use.

Authors

Institutions

Publication Details

Journal
The Journal of Machine Learning for Biomedical Imaging
Published
2026-09-28
DOI
https://doi.org/10.59275/j.melba.2026-eb33
Primary Topic
Anomaly Detection Techniques and Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Knowledge-based anomaly detection for identifying network-induced shape artifacts

Seyed Kahaki, Elim Thompsond, Ghada Zamzmid, Aldo Badanod et al.
The Journal of Machine Learning for Biomedical Imaging
Anomaly Detection Techniques and Applications
article

Knowledge-based anomaly detection for identifying network-induced shape artifacts

Seyed Kahaki, Elim Thompsond, Ghada Zamzmid, Aldo Badanod, Adarsh Subbaswamyd, Miguel Lagod, Tahsin Rahmand, Rucha Deshpanded, Jana G. Delfinod
article en

Abstract

Synthetic data provides a promising approach to address data scarcity for training machine learning models; however, adoption without proper quality assessments may introduce artifacts, distortions, and unrealistic features that compromise model performance and downstream utility for medical imaging research and evaluation purposes. This work introduces a novel knowledge-based anomaly detection method for detecting network-induced shape artifacts in synthetic images. The introduced method utilizes a two-stage approach comprising (i) a novel feature extractor that constructs a specialized feature space by analyzing the per-image distribution of angle gradients along anatomical boundaries, and (ii) an isolation forest-based anomaly detector. This representation captures local shape variations while maintaining global anatomical correspondence regardless of size differences. The isolation forest is trained on patient data to learn normal anatomical shape characteristics and assigns anomaly scores to synthetic images. This method may be used to detect anatomically unrealistic images irrespective of the generative model used and provides interpretability through its knowledge-based design. We demonstrate the effectiveness of the method for identifying network-induced shape artifacts in two synthetic mammography datasets generated by different architectures: a latent diffusion model and StyleGAN2, trained on CSAW-M and VinDr-Mammo patient datasets respectively. Quantitative evaluation shows that the method successfully concentrates artifacts in the most anomalous partition (1st percentile), with AUC values of 0.97 (CSAW-syn) and 0.91 (VMLO-syn). Comparison with anomaly detection based on other standard shape descriptors was used to demonstrate the superiority of the proposed method at isolating subtle shape artifacts along with extreme ones. In addition, a reader study involving three imaging scientists confirmed that images identified by the method as containing network-induced shape artifacts were also flagged by human readers with mean agreement rates of 66% (CSAW-syn) and 68% (VMLO-syn) for the most anomalous partition, approximately 1.5-2 times higher than the least anomalous partition. Kendall-Tau correlations between algorithmic and human rankings were 0.45 and 0.43 for the two datasets, indicating reasonable agreement despite the challenging nature of subtle artifact detection. This method is a step forward in supporting quality assurance practices for synthetic data in research and development contexts, as it allows developers to evaluate synthetic images for known anatomic constraints and pinpoint and address specific issues to improve the overall quality of a synthetic dataset. This method is presented as a research tool for synthetic image quality evaluation. Its application in clinical or regulatory decision-making contexts would require additional validation appropriate to the intended use.

The Journal of Machine Learning for Biomedical ImagingVol. 2026(MIDL 2025)
Center for Devices and Radiological Health (US)
Openalex Percentile: Top 9%
Anomaly Detection Techniques and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.