Iza Guspian, Chien-Chang Lin, Stephen J.H. Yang · KSII Transactions on Internet and Information Systems 2026 · 2026
DOI: 10.3837/tiis.2026.09.008
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This study examined whether adding linguistic features to a pre-trained multilingual BERT model could reduce L1-L2 bias in identifying AI-generated text in Indonesian, thereby enabling the model to generalize better to non-native writers (L2).To isolate the effect of linguistic features, we compared a baseline (mBERT + LR) against a hybrid model (mBERT + Ling36D + LR).The results show that linguistic features are effective since they reduce the L1-L2 false-positive gap (FPR gap) from 0.0233 to 0.0150.This is a relative reduction in residual bias of ≈ 35.6% with no loss of global performance.We established the initial representational bias by benchmarking against independent OpenAI and Gemini centroid probes, and further validated the stability of our hybrid model's fairness improvements using bootstrap resampling.Finally, to support potential use, we also propose a component-based explanation layer to explain detection scores based on more explicit semantic and linguistic evidence.Our findings suggest that such a hybrid system could provide a potential route towards fairer and more transparent AI assessment in multilingual education.
No comments yet — start the discussion below.