From Prompts To Masks: Assessing Gemini Image Models for Tooth Segmentation on Panoramic Radiographs
Introduction and aims Tooth-region segmentation on panoramic radiographs supports annotation, region-of-interest extraction, and dataset triage, but conventional approaches require task-specific training labels. Although vendor-hosted multimodal models can follow natural-language prompts, their pixel-level reliability in dental radiographs remains unclear. This study assessed the accuracy, false-positive behavior, and run-to-run stability of two Gemini Image models for prompt-based tooth-region masking. Methods Ninety-six panoramic radiographs from the Tufts Dental Database (32 permanent dentition, 32 mixed dentition, 32 edentulous) were standardized to 2880 × 1472 pixels. Each image was submitted five times to Gemini 3 Pro Image and Gemini 3.1 Flash Image using a fixed text-to-mask prompt. Permanent- and mixed-dentition outputs were evaluated using mean Dice similarity coefficient (DSC) and intersection over union (IoU); edentulous outputs were evaluated using specificity, false-positive (FP) area ratio, and no-mask rate. Bootstrap 95% confidence intervals and Holm-corrected paired permutation tests were used for inference. Results In the pooled permanent- and mixed-dentition group, both models achieved a mean DSC of 0.883; mean IoU was 0.794 and 0.795 for Gemini 3 Pro Image and Gemini 3.1 Flash Image, respectively. Permanent dentition favored Gemini 3 Pro Image numerically (DSC 0.905 vs 0.896), whereas mixed dentition favored Gemini 3.1 Flash Image (DSC 0.871 vs 0.861). For edentulous radiographs, specificity and FP area ratio were comparable between models, and the no-mask rate was numerically higher for Gemini 3 Pro Image (76.3% vs 63.8%). No comparison remained significant after Holm correction, and k-of-5 analysis revealed variable run-to-run stability. Conclusion Prompt-based tooth-region segmentation yielded clinically plausible outputs without task-specific training, but dentition-dependent performance, target-absent false positives, and limited reproducibility support expert-supervised rather than autonomous use. Clinical relevance Expert-reviewed annotation support, dataset triage, and human-in-the-loop preprocessing are more appropriate than stand-alone clinical segmentation.
Authors
- Naoya Kakimoto (ORCID: https://orcid.org/0000-0003-1274-8991)
- Tsuyoshi Taji (ORCID: https://orcid.org/0000-0002-1286-5357)
- Yuichi Mine (ORCID: https://orcid.org/0000-0002-7057-1955)
- Shota Okazaki
- Tzu-Yu Peng
Institutions
- Hiroshima University (JP)
- Sapporo City University (JP)
- Taipei Medical University (TW)
Publication Details
- Journal
- International Dental Journal
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1016/j.identj.2026.111171
- Primary Topic
- Dental Radiography and Imaging
- Type
- article
- Field-Weighted Citation Impact
- 0.00