Hamza Shahbaz · Advances in Artificial Intelligence Research 2026 · 2026
DOI: 10.54569/aair.1975838
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Background: The exponential growth of scientific literature creates a practical bottleneck: researchers must manually navigate domain taxonomies, identify research gaps, formulate titles, and bootstrap a methodology before writing begins. Methods: This study presents an end-to-end Intelligent Research Paper Assistant built over 499,999 scholarly abstracts spanning 10 domains and 25 subdomains, integrating six modules: title suggestion, abstract retrieval, methodology generation, dual-branch domain classification, post-hoc explainability, and research-gap ranking. A deterministic six-stage NLP pipeline feeds a TF-IDF encoder (12,000 features); an incremental SGD linear baseline and a fine-tuned DistilBERT transformer (66.96M parameters) are trained in parallel, with SHAP and LIME providing global and local explainability. Results: Domain classification achieves Accuracy = 1.00 (test-set per-domain accuracy 99.78%) for both models. SHAP and LIME correctly recover domain-discriminative vocabulary, and a composite novelty score identifies Social History and Digital Humanities as the most under-researched subdomains. Limitations: Results derive from a generated corpus with probable label co-occurrence inflating domain separability; real deployment requires retraining on verified bibliographic data.
No comments yet — start the discussion below.