CT-Bench: A Comprehensive Benchmark for Multimodal AI in Computed Tomography Analysis
Publicly available computed tomography (CT) resources with lesion-level image, spatial, and textual annotations remain limited, restricting reproducible development and evaluation of multimodal medical AI. We present CT-Bench, a CT lesion image–text resource and visual question answering benchmark constructed by extending DeepLesion with curated lesion-level text, standardized lesion-size representations, structured attributes, and QA tasks. CT-Bench contains 20,335 lesion instances from 7,795 CT studies and 3,793 patients, including CT key-slice images, bounding-box annotations, PACS-derived lesion descriptions, lesion size measurements, and associated metadata. The resource also includes a multiple-choice QA benchmark with 2,850 question-answer pairs across seven lesion analysis tasks: image-to-description matching, multi-slice CT-to-description matching, description-to-image retrieval, description-to-bounding-box localization, lesion size estimation, image-based attribute recognition, and multi-slice CT attribute recognition. Hard negative answer choices are included to evaluate fine-grained image-text alignment and lesion-level reasoning. The resource is available at https://kaggle.com/datasets/cd1661d6d6aeab08b8eb99b58885b4489d76fc5ac07d5aac76cae577e6426e2f with documentation, metadata files, and reproducibility code. Validation includes annotation review by trained annotators and medical experts, hard-negative verification, human-reader comparison, and baseline experiments with general and medical vision-language models.
Authors
- Kenneth C. Wang (ORCID: https://orcid.org/0000-0002-6475-814X)
- Benjamin Hou (ORCID: https://orcid.org/0000-0003-3968-1707)
- Zhizheng Wang (ORCID: https://orcid.org/0000-0003-2584-136X)
- Maame Sarfo-Gyamfi
- Praveen T. S. Balamuralikrishna (ORCID: https://orcid.org/0009-0008-8046-4417)
- Tejas S. Mathai
- Yin Fang
- Qiao Jin
- Ronald M. Summers
- Yifan Yang
- Ran Gu
- Zhiyong Lu
- Qingqing Zhu
Institutions
- National Institutes of Health (US)
- Howard University (US)
- Department of Veterans Affairs (AU)
Publication Details
- Journal
- The Journal of Machine Learning for Biomedical Imaging
- Published
- 2026-09-21
- DOI
- https://doi.org/10.59275/j.melba.2026-b141
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00