Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research

In this study, presented a comprehensive evaluation of abstractive and extractive summarization performance across three prominent large language models (LLMs): ChatGPT, DeepSeek, and Gemini. A total of 8,000 cardiovascular-related research abstracts were collected from PubMed and summarized using two distinct prompting strategies: abstractive and extractive. This process yielded a dataset of 48,000 summaries. To assess summarization quality, applied a multi-metric evaluation framework including semantic similarity (SBERT cosine), BLEU, GLEU, ROUGE-1 F1, ROUGE-2 F1, ROUGE-L F1, and METEOR. The results indicate that extractive summaries, particularly those generated by ChatGPT, consistently achieve higher scores across most metrics, suggesting stronger lexical fidelity and sequence retention. While Gemini shows balanced performance between abstraction and extraction, DeepSeek yields lower scores in both approaches. This work highlights critical differences in LLM behavior depending on the summarization method and offers a benchmark dataset and evaluation pipeline for future research on AI-assisted biomedical summarization.

Authors

Institutions

Publication Details

Journal
Sakarya University Journal of Computer and Information Sciences
Published
2026-09-30
DOI
https://doi.org/10.35377/saucis...1780353
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research

Aytuğ Onan, Burcu Baştürk
Sakarya University Journal of Computer and Information Sciences
Topic Modeling
article

Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research

Aytuğ Onan, Burcu Baştürk
article en

Abstract

In this study, presented a comprehensive evaluation of abstractive and extractive summarization performance across three prominent large language models (LLMs): ChatGPT, DeepSeek, and Gemini. A total of 8,000 cardiovascular-related research abstracts were collected from PubMed and summarized using two distinct prompting strategies: abstractive and extractive. This process yielded a dataset of 48,000 summaries. To assess summarization quality, applied a multi-metric evaluation framework including semantic similarity (SBERT cosine), BLEU, GLEU, ROUGE-1 F1, ROUGE-2 F1, ROUGE-L F1, and METEOR. The results indicate that extractive summaries, particularly those generated by ChatGPT, consistently achieve higher scores across most metrics, suggesting stronger lexical fidelity and sequence retention. While Gemini shows balanced performance between abstraction and extraction, DeepSeek yields lower scores in both approaches. This work highlights critical differences in LLM behavior depending on the summarization method and offers a benchmark dataset and evaluation pipeline for future research on AI-assisted biomedical summarization.

Sakarya University Journal of Computer and Information SciencesVol. 9(4)
Izmir Kâtip Çelebi University (TR)
Quality Education
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.