Huy Tran, Linh Luong · DS Journal of Digital Science and Technology 2026 · 2026
DOI: 10.59232/dst-v5i3p102
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Keeping track of how students behave in class is hard to do consistently, and right now it mostly comes down to a teacher's own judgment, formed by watching the room. Visual Question Answering (VQA) datasets rarely help here: most are English, built around everyday photos, or simply not designed with a classroom in mind. This paper presents Classroom VQA, a Vietnamese dataset built specifically for classroom learning-behavior understanding, comprising 3,000 classroom images drawn from Places 365 and 9,000 question–answer pairs produced and checked through a six-stage annotation pipeline. The questions span four categories – confirmation, object and attribute, counting, and action and state – chosen to cover different kinds of visual reasoning a classroom-monitoring system might need. A PhoBERT–EVA02 baseline is also proposed, using bidirectional Co-Cross-Attention so that Text-to-Image and Image-to-Text signals can inform one another. Tested across four EVA02 backbone sizes, this baseline reaches 70.78–71.44% accuracy, with under one percentage point separating the smallest and largest models. Classroom VQA is compared against six existing VQA datasets, and its question distribution, category-level error patterns, ethical considerations, and limitations are examined in detail. Together, the dataset and baseline give Vietnamese classroom VQA a concrete starting point and, more broadly, a foundation for education-oriented multimodal research.
No comments yet — start the discussion below.