Scientific text summarization by attention-augmented deep clustering and semantic-aware sentence rating

One of the biggest challenges facing researchers is the rising amount of scientific literature. Automatic summarization offers an effective solution by generating concise summaries that aid comprehension and accelerate analysis. For lengthy documents, extractive summarization is a common method; however, selecting sentences that accurately convey the document’s meaning remains challenging. This study proposes an improved approach that leverages context comprehension to rank sentences. To better capture complex semantic structures in scientific text, deep clustering is enhanced using an Attention-Augmented Deep Embedding Clustering (ATT-DEC) mechanism built upon a Scientific BERT (SciBERT) encoding backbone, with sentence-level relevance scores assigned by a Multi-Layer Perceptron (MLP). Furthermore, a non-linear mathematical model is used to balance selection. Recall-Oriented Understudy for Gisting Evaluation (ROUGE) F1 metrics were used to assess the method on the ArXiv and PubMed datasets. On ArXiv (R1 = 0.4337, R2 = 0.2402, RL = 0.3527) and PubMed (R1 = 0.4508, R2 = 0.2585, RL = 0.3724), the results demonstrate improvements over state-of-the-art baselines across most evaluation conditions, highlighting the model’s potential to enhance researchers’ effectiveness and the accessibility of scientific documents.

Authors

Institutions

Publication Details

Journal
Ain Shams Engineering Journal
Published
2026-09-28
DOI
https://doi.org/10.1016/j.asej.2026.104446
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Scientific text summarization by attention-augmented deep clustering and semantic-aware sentence rating

Mohammad‐Reza Feizi‐Derakhshi, Sammer Sami Abdulkareem
Ain Shams Engineering Journal
Topic Modeling
article

Scientific text summarization by attention-augmented deep clustering and semantic-aware sentence rating

Mohammad‐Reza Feizi‐Derakhshi, Sammer Sami Abdulkareem
article en

Abstract

One of the biggest challenges facing researchers is the rising amount of scientific literature. Automatic summarization offers an effective solution by generating concise summaries that aid comprehension and accelerate analysis. For lengthy documents, extractive summarization is a common method; however, selecting sentences that accurately convey the document’s meaning remains challenging. This study proposes an improved approach that leverages context comprehension to rank sentences. To better capture complex semantic structures in scientific text, deep clustering is enhanced using an Attention-Augmented Deep Embedding Clustering (ATT-DEC) mechanism built upon a Scientific BERT (SciBERT) encoding backbone, with sentence-level relevance scores assigned by a Multi-Layer Perceptron (MLP). Furthermore, a non-linear mathematical model is used to balance selection. Recall-Oriented Understudy for Gisting Evaluation (ROUGE) F1 metrics were used to assess the method on the ArXiv and PubMed datasets. On ArXiv (R1 = 0.4337, R2 = 0.2402, RL = 0.3527) and PubMed (R1 = 0.4508, R2 = 0.2585, RL = 0.3724), the results demonstrate improvements over state-of-the-art baselines across most evaluation conditions, highlighting the model’s potential to enhance researchers’ effectiveness and the accessibility of scientific documents.

Ain Shams Engineering JournalVol. 17(12)
University of Tabriz (IR)
Quality Education
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.