Leona Rexhaj, Luela Prifti · WSEAS TRANSACTIONS ON COMPUTER RESEARCH 2026 · 2026
DOI: 10.37394/232018.2026.14.48
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This study analyzes the semantic structure of an Albanian corpus comprising scientific texts with 11.5 million words and approximately 74,500 paragraphs. It aims to cluster the documents according to semantic similarity and uncover thematic structures. The proposed methodology integrates Sentence-BERT to construct semantic representations, UMAP for dimensionality reduction, and HDBSCAN for clustering. Large-scale analysis was conducted without predefined labels, allowing the clusters to emerge directly from the textual content. The paraphrase-multilingual-MiniLM-L12-v2 model achieved the best clustering performance and produced a stable seven-cluster structure. Two representation strategies were evaluated: full-text representation and section-based representation using abstracts, introductions, and conclusions. External validation metrics for both representations indicated moderate alignment with disciplinary classifications, reflecting the interdisciplinary nature of scientific texts. For the section-based representation, the Silhouette score was 0.7154, the Davies-Bouldin index was 0.5025, and the Calinski-Harabasz index was 651.99. The relationship between the resulting clusters and academic fields was statistically significant, with χ^2=748.443 and p=4.90×〖10〗^(-134), while Cramer’s V = 0.746 indicated a strong association. These results suggest that the section-based representation produces clearer and more semantically coherent clusters by focusing on the main thematic information and reducing non-discriminative content. The proposed framework offers a practical and scalable approach for analyzing large scientific corpora and may be adapted to other languages and low-resource research domains.
No comments yet — start the discussion below.