Moh. Aminudin, Riza Arifudin · Recursive Journal of Informatics 2026 · 2026
DOI: 10.15294/rji.v4i2.62732
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Email spam filters may classify unmodified messages accurately yet remain vulnerable to deliberate feature manipulation. Hyperparameter optimization can improve predictive performance, but clean-data gains do not by themselves characterize behavior under test-time perturbations. A controlled evaluation is therefore needed to assess both properties within the same leakage-controlled protocol. Purpose: This study compared the clean classification performance and feature-space robustness of a library-default Random Forest (RF-default) and a Bayesian-optimized Random Forest (BO-RF) on Spambase. Methods/Study design/approach: Exact duplicates were removed and predictor-identical records were grouped to prevent cross-fold leakage. The models were evaluated using two repetitions of five-fold nested stratified group-aware cross-validation with shared outer folds. Bayesian Optimization maximized inner-validation spam-class F1. Frozen outer-test folds were assessed on clean inputs, class-center-directed perturbations at α = 0.05, 0.10, 0.20, and 0.40, and 30 magnitude-matched random directions. Retained inference covered clean-versus-directed and directed-versus-random contrasts in F1, Recall, and model-specific attack success rate (ASR). Result/Findings: BO-RF achieved 95.01 ± 0.66% clean Accuracy compared with 94.92 ± 0.74% for RF-default; its mean clean Precision, Recall, and F1 were also slightly higher. Directed perturbations progressively reduced performance. At α = 0.40, Accuracy was 89.24% for RF-default and 89.32% for BO-RF, while Recall declined to 77.94% and 78.03%, respectively, indicating increased false negatives. Directed ASR reached 15.470% and 15.484%, and relative F1 degradation reached 8.915% for both models. All 24 clean-to-directed comparisons and all 24 directed-versus-random comparisons for F1, Recall, and ASR met the Holm-adjusted threshold. Novelty/Originality/Value: The study combines duplicate-aware nested evaluation, shared folds, and identical perturbation settings to examine clean optimization and directed feature-space degradation in one paired protocol. Findings are restricted to the evaluated feature-space setting.
No comments yet — start the discussion below.