
Mohammed Arif Hasan Chowdhury, Shusmoy Chowdhury, Fariha Tabassum, Nisha Baul, Muhammad Minhazul Haque Bhuiyan, Shamima Islam · Engineering Technology & Applied Science Research 2026 · 2026
DOI: 10.48084/etasr.20584
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The problem of semantic text similarity is a basic and practical application of Natural Language Processing (NLP), used widely to detect plagiarism, perform information retrieval, distinguish duplicate questions, assess machine translation, and support educational systems. Although transformer-based sentence similarity models have succeeded for languages with abundant linguistic resources, Bangla sentence similarity is difficult due to the poor variety of linguistic resources, complicated morphology, differences in dialects, and the high processing power requirements of deep learning sentence similarity models. This paper proposes a lightweight and interpretable feature engineering approach for Bangla sentence similarity detection based on classical machine learning approaches. Five normalized character-based similarity measures—Damerau-Levenshtein, Levenshtein, Jaro, Jaro-Winkler, and Longest Common Substring (LCS)—are obtained as complementary lexical features and subsequently used to train Support Vector Machine (SVM), Random Forest (RM), Decision Tree (DT), and Logistic Regression classifiers. The framework is tested on the BanglaParaphrase dataset with more than 40,000 manually annotated sentence pairs, using an 80:20 train/test split and 10-fold Cross-Validation (CV). The experimental results indicate that all classifiers have accuracy and F1-scores better than 97%, while LR and SVM present the most uniform performance on the validation folds. Correlation analysis shows that the proposed similarity measures are correlated and represent complementary information. The proposed framework gives an interpretable, effective, and efficient baseline method for Bangla sentence similarity detection that works well in both low-resource and resource-affluent NLP applications.
No comments yet — start the discussion below.