An Intelligent Research Paper Assistant: Integrated Retrieval, Explainable Domain Classification, and Research-Gap Ranking
Background: The exponential growth of scientific literature creates a practical bottleneck: researchers must manually navigate domain taxonomies, identify research gaps, formulate titles, and bootstrap a methodology before writing begins. Methods: This study presents an end-to-end Intelligent Research Paper Assistant built over 499,999 scholarly abstracts spanning 10 domains and 25 subdomains, integrating six modules: title suggestion, abstract retrieval, methodology generation, dual-branch domain classification, post-hoc explainability, and research-gap ranking. A deterministic six-stage NLP pipeline feeds a TF-IDF encoder (12,000 features); an incremental SGD linear baseline and a fine-tuned DistilBERT transformer (66.96M parameters) are trained in parallel, with SHAP and LIME providing global and local explainability. Results: Domain classification achieves Accuracy = 1.00 (test-set per-domain accuracy 99.78%) for both models. SHAP and LIME correctly recover domain-discriminative vocabulary, and a composite novelty score identifies Social History and Digital Humanities as the most under-researched subdomains. Limitations: Results derive from a generated corpus with probable label co-occurrence inflating domain separability; real deployment requires retraining on verified bibliographic data.
Authors
- Hamza Shahbaz (ORCID: https://orcid.org/0009-0005-7552-1585)
Institutions
- Government College University, Faisalabad (PK)
Publication Details
- Journal
- Advances in Artificial Intelligence Research
- Published
- 2026-09-17
- DOI
- https://doi.org/10.54569/aair.1975838
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Government College University, Lahore