AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Model Capabilities in Astronomy

Astronomical image interpretation requires not only visual perception, but also familiarity with domain-specific representations, observational conventions, and astrophysical reasoning. Although recent multimodal large language models (MLLMs) have shown strong general visual–language capabilities, their performance on specialized astronomical figures remains insufficiently evaluated. To address this gap, we introduce AstroMMBench, an astronomy-specific benchmark for evaluating MLLMs on figure-associated astronomical multiple-choice questions. AstroMMBench contains 592 expert-screened multiple-choice questions covering six major astrophysical subfields: Astrophysics of Galaxies, Cosmology and Extragalactic Astrophysics, Earth and Planetary Astrophysics, High Energy Astrophysical Phenomena, Instrumentation and Methods for Astrophysics, and Solar and Stellar Astrophysics. The questions were generated from astronomical figures and associated textual context through an automated pipeline, followed by multi-stage filtering and expert review to assess scientific correctness, image–question alignment, and answer uniqueness. Using AstroMMBench, we evaluated 25 selected MLLMs, including 22 open-source and 3 closed-source models. The results show substantial performance differences across models and astrophysical subfields. Ovis2-34B achieved the highest point-estimate overall accuracy of 71.3%, with performance comparable to strong closed-source models such as ChatGPT-4o and Doubao-1.5-vision-pro. Subfield-level analysis shows the lowest mean point-estimate performance in cosmology, whereas high-energy astrophysics exhibits substantial between-model variability rather than uniformly low performance. Question-only results on the 592 benchmark questions yielded 27.53% for Qwen2.5-VL-7B and 24.83% for InternVL3-38B, compared with image-present accuracies of 58.11% and 68.24%, respectively. These descriptive comparisons suggest a substantial contribution from visual input in the two models. AstroMMBench provides a structured and extensible resource for comparing model behavior on figure-associated astronomical questions.

Authors

Institutions

Publication Details

Journal
Universe
Published
2026-09-30
DOI
https://doi.org/10.3390/universe12100296
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Model Capabilities in Astronomy

Jinghang Shi, Xiao Kong, Yanxia Zhang, Xiaoyu Tang et al.
Universe
Multimodal Machine Learning Applications
article

AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Model Capabilities in Astronomy

Jinghang Shi, Xiao Kong, Yanxia Zhang, Xiaoyu Tang, Yang Huang, Caizhan Yue, A-Li Luo, Yuyang Li
article en

Abstract

Astronomical image interpretation requires not only visual perception, but also familiarity with domain-specific representations, observational conventions, and astrophysical reasoning. Although recent multimodal large language models (MLLMs) have shown strong general visual–language capabilities, their performance on specialized astronomical figures remains insufficiently evaluated. To address this gap, we introduce AstroMMBench, an astronomy-specific benchmark for evaluating MLLMs on figure-associated astronomical multiple-choice questions. AstroMMBench contains 592 expert-screened multiple-choice questions covering six major astrophysical subfields: Astrophysics of Galaxies, Cosmology and Extragalactic Astrophysics, Earth and Planetary Astrophysics, High Energy Astrophysical Phenomena, Instrumentation and Methods for Astrophysics, and Solar and Stellar Astrophysics. The questions were generated from astronomical figures and associated textual context through an automated pipeline, followed by multi-stage filtering and expert review to assess scientific correctness, image–question alignment, and answer uniqueness. Using AstroMMBench, we evaluated 25 selected MLLMs, including 22 open-source and 3 closed-source models. The results show substantial performance differences across models and astrophysical subfields. Ovis2-34B achieved the highest point-estimate overall accuracy of 71.3%, with performance comparable to strong closed-source models such as ChatGPT-4o and Doubao-1.5-vision-pro. Subfield-level analysis shows the lowest mean point-estimate performance in cosmology, whereas high-energy astrophysics exhibits substantial between-model variability rather than uniformly low performance. Question-only results on the 592 benchmark questions yielded 27.53% for Qwen2.5-VL-7B and 24.83% for InternVL3-38B, compared with image-present accuracies of 58.11% and 68.24%, respectively. These descriptive comparisons suggest a substantial contribution from visual input in the two models. AstroMMBench provides a structured and extensible resource for comparing model behavior on figure-associated astronomical questions.

UniverseVol. 12(10)
Tianjin University (CN), Chinese Academy of Sciences (CN), Zhejiang Lab (CN), National Astronomical Observatories (CN), University of Chinese Academy of Sciences (CN)
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.