From 2D Vision–Language Models to Volumetric Medical AI: Large Language Models and Foundation Models for 3D Medical Imaging

Multimodal large language models (MLLMs) and vision–language models (VLMs) have rapidly entered medicine, demonstrating promising performance in clinical reasoning, radiology report generation, and visual question answering (VQA). However, many current multimodal architectures and pretrained visual backbones remain fundamentally rooted in two-dimensional (2D) image processing, even though major clinical imaging modalities, including computed tomography (CT), magnetic resonance imaging (MRI), optical coherence tomography (OCT), and echocardiography, are inherently volumetric or temporal. This narrative review examines the transition from 2D vision–language systems to volumetric multimodal AI, tracing the evolution from 2D and slice- or projection-based approaches through sequential and video-like methods to three-dimensional (3D) vision foundation models and native 3D VLMs/MLLMs. We examine their representational and computational trade-offs, evaluation gaps, and clinically grounded benchmarks. Approaches differ substantially in how they represent and preserve 3D information. Slice- and projection-based methods offer computational efficiency but may discard spatial context, whereas sequential and native volumetric approaches increasingly model relationships across the full imaging study. Recent 3D foundation models and multimodal systems demonstrate the feasibility of reusable volumetric representations and language-enabled 3D image interpretation, but face barriers in computational cost, training-data scale, evaluation methodology, and clinical reliability. Only 53% of Med-Gemini-3D reports were judged clinically acceptable, and natural language processing (NLP) metrics such as BLEU and ROUGE correlate poorly with diagnostic correctness. True 3D multimodal medical intelligence remains in its early stages. Future progress requires efficient volumetric representation strategies, clinically grounded evaluation frameworks, standardized benchmarks, and robust cross-institution validation.

Authors

Institutions

Publication Details

Journal
Computation
Published
2026-09-13
DOI
https://doi.org/10.3390/computation14090215
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

From 2D Vision–Language Models to Volumetric Medical AI: Large Language Models and Foundation Models for 3D Medical Imaging

Roni Ramon‐Gonen, Haya Engelstein
Computation
Multimodal Machine Learning Applications
article

From 2D Vision–Language Models to Volumetric Medical AI: Large Language Models and Foundation Models for 3D Medical Imaging

Roni Ramon‐Gonen, Haya Engelstein
article en

Abstract

Multimodal large language models (MLLMs) and vision–language models (VLMs) have rapidly entered medicine, demonstrating promising performance in clinical reasoning, radiology report generation, and visual question answering (VQA). However, many current multimodal architectures and pretrained visual backbones remain fundamentally rooted in two-dimensional (2D) image processing, even though major clinical imaging modalities, including computed tomography (CT), magnetic resonance imaging (MRI), optical coherence tomography (OCT), and echocardiography, are inherently volumetric or temporal. This narrative review examines the transition from 2D vision–language systems to volumetric multimodal AI, tracing the evolution from 2D and slice- or projection-based approaches through sequential and video-like methods to three-dimensional (3D) vision foundation models and native 3D VLMs/MLLMs. We examine their representational and computational trade-offs, evaluation gaps, and clinically grounded benchmarks. Approaches differ substantially in how they represent and preserve 3D information. Slice- and projection-based methods offer computational efficiency but may discard spatial context, whereas sequential and native volumetric approaches increasingly model relationships across the full imaging study. Recent 3D foundation models and multimodal systems demonstrate the feasibility of reusable volumetric representations and language-enabled 3D image interpretation, but face barriers in computational cost, training-data scale, evaluation methodology, and clinical reliability. Only 53% of Med-Gemini-3D reports were judged clinically acceptable, and natural language processing (NLP) metrics such as BLEU and ROUGE correlate poorly with diagnostic correctness. True 3D multimodal medical intelligence remains in its early stages. Future progress requires efficient volumetric representation strategies, clinically grounded evaluation frameworks, standardized benchmarks, and robust cross-institution validation.

ComputationVol. 14(9)
Bar-Ilan University (IL), Sheba Medical Center (IL)
Quality Education
Openalex Percentile: Top 13%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.