Ifeoluwa D Oseni, Ilobekemen P. Oladoja, Olumide Adewale · Cureus Journal of Computer Science. 2026 · 2026
DOI: 10.7759/s44389-026-00300-x
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Phishing is a significant form of cybercrime.It exploits unsuspecting users through the use of deceptive websites designed to harvest sensitive information.This study proposes a practical hybrid phishing detection framework that integrates machine learning, explainable artificial intelligence, and a rule-based component specifically for handling inactive URLs using HTTP status heuristics.Random Forest (RF) and Extreme Gradient Boosting served as the baseline classifiers, with weighted soft voting and stacking ensembles applied to integrate their predictions and improve overall performance.The classifiers were trained on the 2024 Mendeley and UCI PhiUSIIL phishing datasets.XGBoost achieved better accuracy, recall, and F1-score among the two baseline classifiers, while RF achieved better precision.Both ensemble methods surpassed the individual classifiers in accuracy, recall, and F1-score.The stacking ensemble, with a logistic regression meta-learner, slightly outperformed the weighted voting ensemble strategy, achieving an accuracy of 98.88%, recall of 99.01%, and an F1-score of 98.73%.To manage non-responsive links, an operational filter leverages HTTP status heuristics to assign a final classification.For URLs outside the 200-299 status range, this rule-based label takes precedence, while the underlying machine learning models concurrently generate a supplementary probability estimate using retrievable and imputed attributes.SHapley Additive exPlanations is used to interpret active URL model prediction and help highlight the most influential features that distinguish legitimate websites from phishing websites.This enhances transparency and also provides insights into phishing behaviors, such as suspicious domain structures and DNS anomalies.This proposed interpretable phishing detection approach is easy to use and deploy across diverse user groups, regardless of their level of technical expertise.
No comments yet — start the discussion below.