Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research
In this study, presented a comprehensive evaluation of abstractive and extractive summarization performance across three prominent large language models (LLMs): ChatGPT, DeepSeek, and Gemini. A total of 8,000 cardiovascular-related research abstracts were collected from PubMed and summarized using two distinct prompting strategies: abstractive and extractive. This process yielded a dataset of 48,000 summaries. To assess summarization quality, applied a multi-metric evaluation framework including semantic similarity (SBERT cosine), BLEU, GLEU, ROUGE-1 F1, ROUGE-2 F1, ROUGE-L F1, and METEOR. The results indicate that extractive summaries, particularly those generated by ChatGPT, consistently achieve higher scores across most metrics, suggesting stronger lexical fidelity and sequence retention. While Gemini shows balanced performance between abstraction and extraction, DeepSeek yields lower scores in both approaches. This work highlights critical differences in LLM behavior depending on the summarization method and offers a benchmark dataset and evaluation pipeline for future research on AI-assisted biomedical summarization.
Authors
- Aytuğ Onan (ORCID: https://orcid.org/0000-0002-9434-5880)
- Burcu Baştürk (ORCID: https://orcid.org/0009-0005-4781-353X)
Institutions
- Izmir Kâtip Çelebi University (TR)
Publication Details
- Journal
- Sakarya University Journal of Computer and Information Sciences
- Published
- 2026-09-30
- DOI
- https://doi.org/10.35377/saucis...1780353
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00