Uncovering Latent Scientific Structures in Albanian Texts Using Transformer-Based Clustering

This study analyzes the semantic structure of an Albanian corpus comprising scientific texts with 11.5 million words and approximately 74,500 paragraphs. It aims to cluster the documents according to semantic similarity and uncover thematic structures. The proposed methodology integrates Sentence-BERT to construct semantic representations, UMAP for dimensionality reduction, and HDBSCAN for clustering. Large-scale analysis was conducted without predefined labels, allowing the clusters to emerge directly from the textual content. The paraphrase-multilingual-MiniLM-L12-v2 model achieved the best clustering performance and produced a stable seven-cluster structure. Two representation strategies were evaluated: full-text representation and section-based representation using abstracts, introductions, and conclusions. External validation metrics for both representations indicated moderate alignment with disciplinary classifications, reflecting the interdisciplinary nature of scientific texts. For the section-based representation, the Silhouette score was 0.7154, the Davies-Bouldin index was 0.5025, and the Calinski-Harabasz index was 651.99. The relationship between the resulting clusters and academic fields was statistically significant, with χ^2=748.443 and p=4.90×〖10〗^(-134), while Cramer’s V = 0.746 indicated a strong association. These results suggest that the section-based representation produces clearer and more semantically coherent clusters by focusing on the main thematic information and reducing non-discriminative content. The proposed framework offers a practical and scalable approach for analyzing large scientific corpora and may be adapted to other languages and low-resource research domains.

Authors

Institutions

Publication Details

Journal
WSEAS TRANSACTIONS ON COMPUTER RESEARCH
Published
2026-09-18
DOI
https://doi.org/10.37394/232018.2026.14.48
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Uncovering Latent Scientific Structures in Albanian Texts Using Transformer-Based Clustering

Luela Prifti, Leona Rexhaj
WSEAS TRANSACTIONS ON COMPUTER RESEARCH
Topic Modeling
article

Uncovering Latent Scientific Structures in Albanian Texts Using Transformer-Based Clustering

Luela Prifti, Leona Rexhaj
article en

Abstract

This study analyzes the semantic structure of an Albanian corpus comprising scientific texts with 11.5 million words and approximately 74,500 paragraphs. It aims to cluster the documents according to semantic similarity and uncover thematic structures. The proposed methodology integrates Sentence-BERT to construct semantic representations, UMAP for dimensionality reduction, and HDBSCAN for clustering. Large-scale analysis was conducted without predefined labels, allowing the clusters to emerge directly from the textual content. The paraphrase-multilingual-MiniLM-L12-v2 model achieved the best clustering performance and produced a stable seven-cluster structure. Two representation strategies were evaluated: full-text representation and section-based representation using abstracts, introductions, and conclusions. External validation metrics for both representations indicated moderate alignment with disciplinary classifications, reflecting the interdisciplinary nature of scientific texts. For the section-based representation, the Silhouette score was 0.7154, the Davies-Bouldin index was 0.5025, and the Calinski-Harabasz index was 651.99. The relationship between the resulting clusters and academic fields was statistically significant, with χ^2=748.443 and p=4.90×〖10〗^(-134), while Cramer’s V = 0.746 indicated a strong association. These results suggest that the section-based representation produces clearer and more semantically coherent clusters by focusing on the main thematic information and reducing non-discriminative content. The proposed framework offers a practical and scalable approach for analyzing large scientific corpora and may be adapted to other languages and low-resource research domains.

WSEAS TRANSACTIONS ON COMPUTER RESEARCHVol. 14
Polytechnic University of Tirana (AL)
Reduced inequalities
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Uncovering Latent Scientific Structures in Albanian Texts Using Transformer-Based Clustering — Luela Prifti, Leona Rexhaj · WSEAS TRANSACTIONS ON COMPUTER RESEARCH (2026) | TGRS Research Map | TGRS