
Muhammad Zidny Naf’an, Edi Winarko, Sigit Priyanta · Engineering Technology & Applied Science Research 2026 · 2026
DOI: 10.48084/etasr.19578
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
In this paper, we introduce MAQASID, a closed-domain dataset for exploring the optimal semantic representation of Indonesian Islamic texts. It is designed to solve challenges in low-resource languages and closed-domain religious datasets, such as Arabic-to-Indonesian transliteration inconsistencies and semantic shifts. Using this dataset, we evaluate the trade-offs between non-contextual embeddings (Continuous Bag-of-Words (CBOW), Skip-Gram, and fastText) and the contextual model (fine-tuned IndoBERT) using both intrinsic and extrinsic evaluation approaches. Our intrinsic evaluation showed that CBOW with simple preprocessing obtained the highest proportion of 46.83% related words on top-10 similarities, and fastText achieved the best accuracy of 51.35% on analogy tasks due to its reliance on subword information. Extrinsic evaluation measures the downstream classification task result with a CNN-static and a BERT classifier. In this extrinsic task, Skip-Gram combined with stemming and stopword removal preprocessing, yielded the best test F1-score of 0.5746 when used as CNN-static features. Fine-tuned IndoBERT generalizes well on stemming preprocessed texts (F1 = 0.6411) but shows degraded performance (F1 = 0.6029) when the aggressive preprocessing paired with stemming and stopword removal disrupts the syntactic context required by self-attention. The findings emphasize that model selection and preprocessing strategies must be carefully aligned to effectively capture the detailed theological semantics of a closed-domain Islamic corpus.
No comments yet — start the discussion below.