Evaluation-design effects and reference-CT-defined HU-class errors in brain MRI-to-CT synthesis on SynthRAD2023: a technical note

To quantify how evaluation design affects brain MRI-to-CT synthesis performance, we analyzed 180 MRI/CT pairs from three centers in the SynthRAD2023 Task 1 dataset. A lightweight 2D U-Net was trained with a common validation-based stopping rule under slice-random, patient-random, center-stratified size-matched patient-random, and three center-held-out designs. Comparisons were descriptive. The primary outcome was whole-mask mean absolute error (MAE), supplemented by MAE within reference-CT-defined Hounsfield unit (HU) classes. MAE (mean ± standard deviation) was 94.4 ± 3.0, 100.8 ± 1.3, and 105.2 ± 2.8 HU for the random designs and 187.4 ± 6.0, 160.6 ± 7.1, and 137.4 ± 5.1 HU for centers A, B, and C held out. Slice-random testing yielded lower MAE than patient-random testing. Center-held-out testing yielded higher MAE than the center-stratified comparator, although institutional and acquisition-domain effects could not be separated. Bone errors were high across designs; adipose-range and air errors were elevated in specific center-held-out tests.

Authors

Institutions

Publication Details

Journal
Radiological Physics and Technology
Published
2026-09-10
DOI
https://doi.org/10.1007/s12194-026-01134-x
Primary Topic
Advanced MRI Techniques and Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluation-design effects and reference-CT-defined HU-class errors in brain MRI-to-CT synthesis on SynthRAD2023: a technical note

Tatsuya Hayashi, Shinya Kojima, Norio Hayashi
Radiological Physics and Technology
Advanced MRI Techniques and Applications
article

Evaluation-design effects and reference-CT-defined HU-class errors in brain MRI-to-CT synthesis on SynthRAD2023: a technical note

Tatsuya Hayashi, Shinya Kojima, Norio Hayashi
article en

Abstract

To quantify how evaluation design affects brain MRI-to-CT synthesis performance, we analyzed 180 MRI/CT pairs from three centers in the SynthRAD2023 Task 1 dataset. A lightweight 2D U-Net was trained with a common validation-based stopping rule under slice-random, patient-random, center-stratified size-matched patient-random, and three center-held-out designs. Comparisons were descriptive. The primary outcome was whole-mask mean absolute error (MAE), supplemented by MAE within reference-CT-defined Hounsfield unit (HU) classes. MAE (mean ± standard deviation) was 94.4 ± 3.0, 100.8 ± 1.3, and 105.2 ± 2.8 HU for the random designs and 187.4 ± 6.0, 160.6 ± 7.1, and 137.4 ± 5.1 HU for centers A, B, and C held out. Slice-random testing yielded lower MAE than patient-random testing. Center-held-out testing yielded higher MAE than the center-stratified comparator, although institutional and acquisition-domain effects could not be separated. Bone errors were high across designs; adipose-range and air errors were elevated in specific center-held-out tests.

Radiological Physics and Technology
Kitasato University (JP), Gunma Prefectural College of Health Sciences (JP), Teikyo University (JP)
Openalex Percentile: Top 11%
Advanced MRI Techniques and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.