
P. Senthilkumar, K. Nandhini · Engineering Technology & Applied Science Research 2026 · 2026
DOI: 10.48084/etasr.21127
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Mock tests play an important role in preparing for academic and competitive examinations, and most mock tests have been designed based on the Question Answering (QA) technique, which is one of the areas in Natural Language Processing (NLP). There are two main tasks in QA: Automatic Question Generation (AQG), which entails generating questions from a given context based on available keywords, and Automatic Answer Generation (AAG), which entails generating an answer based on the given question and context. To build a QA system, a proper dataset is required, and a domain-specific dataset can provide advantages over a generic-domain dataset in terms of processing time, hardware requirements, and accuracy. However, fewer domain-specific datasets are available for subjective computer science examinations. To fulfill this need, we develop a question generation model for building a domain-specific dataset. The proposed approach consists of four phases. In the first phase, a hybrid training dataset is constructed by curating subjective questions from two established benchmarks, the Stanford Question Answering Dataset (SQuAD) and the ReAding Comprehension Dataset from Examinations (RACE). In the second phase, computer science-related data are collected from Wikipedia, comprising 406 computer science topics with 17,520 sentences. In the third phase, the Text-to-Text Transfer Transformer (T5) model is utilized with the hybrid dataset to generate subjective questions. In the fourth phase, the generated questions are evaluated using Bilingual Evaluation Understudy (BLEU), Recall-Oriented Understudy for Gisting Evaluation (ROUGE-L), and Metric for Evaluation of Translation with Explicit ORdering (METEOR), as well as human judgment based on a five-point Likert scale to assess the quality of the generated questions and the consistency of the human evaluations. The proposed model generates subjective questions for the computer science dataset and outperforms other baseline models according to BLEU, ROUGE-L, and METEOR scores, as well as human annotator scores.
No comments yet — start the discussion below.