Chanyi Lee, Sang Ho Oh · Electronics 2026 · 2026
DOI: 10.3390/electronics15194518
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This study proposes a multi-stage framework for generating, evaluating, refining, and assigning cognitive domains to story-based candidate questions for potential cognitive training applications. Using 20 FairytaleQA stories, 1668 questions were initially generated, of which 1351 remained after quality evaluation, refinement, and Solver–Supervisor review. The framework combines criterion-based LLM evaluation, iterative self-refinement, supervisory review, evidence-grounded answerability assessment, and task-based multi-label cognitive-domain assignment. In the question-quality ablation study, Supervisor + Refinement achieved an 83.9% branching pass rate with 3.59 mean refinement iterations, compared with 68.7% and 4.63 iterations without the Evaluation Supervisor. Relative to Zero-shot generation, coherence, consistency, fluency, and relevance increased by 1.28, 1.21, 0.13, and 1.96 points, respectively. Independent evaluation with Granite and Phi-4-mini consistently showed improvements in consistency and relevance. For domain assignment, the Full Framework achieved inter-seed Jaccard, MASI, Exact Match, and Fleiss’ κ values of 0.917, 0.904, 0.883, and 0.874, respectively. These results support the framework’s utility for improving LLM-assessed question quality and assignment stability, while expert validation remains necessary to establish cognitive validity and training suitability.
No comments yet — start the discussion below.