Zerun Lin, Junming Chen, Ke Jin, Yanli Chen · Applied Sciences 2026 · 2026
DOI: 10.3390/app16199754
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Detecting student classroom behaviors from ordinary classroom videos is complicated by weak features for small targets such as rear-row mobile phones, feature loss caused by inter-student occlusion, imbalance between frequent learning behaviors and rarer off-task behaviors, and accuracy degradation when a detector trained in one classroom setting is transferred to another. Existing YOLO-series detectors are typically limited by shallow detail extraction and insufficient multi-scale fusion, making it difficult to reconcile accuracy with real-time inference. To address these limitations, we propose YOLO11-GELSK, an optimized detection architecture built upon YOLO11s that integrates a gated Global Edge Information Transfer (GEIT) module with a star-shaped Large-Kernel Separable Attention (LSKA) fusion neck. The gated GEIT module extends the shallow, Sobel-based edge-transfer design of SCB-YOLO with an image-level channel gate and cross-stage residual transfer, strengthening contour extraction for fine-grained actions such as bowing the head, leaning over the table, or holding a phone, while the star-fusion neck combined with LSKA attention captures global spatial context, mitigating occlusion and missed detections of distant targets. A scale-adaptive Focal-CIoU composite loss balances training across behavior categories to suppress false positives and negatives for underrepresented classes, and a Maximum Mean Discrepancy (MMD) unsupervised domain-adaptation term aligns neck-level features across classroom subsets; L2 weight decay and stochastic depth further curb overfitting. Six public SCB-Dataset3 behavior categories—hand-raising, reading, writing, using a phone, bowing the head, and leaning over the table—are modeled as an object-detection task, and a multi-target probability fusion mechanism aggregates per-student outputs into a continuous, frame-level off-task score. On a recording-grouped validation split of SCB-Dataset3, YOLO11-GELSK attains 91.7% mAP@0.5 and 90.2% macro-F1, exceeding YOLO11s, ResNet50-based Faster R-CNN, and EfficientNet-B3-based RetinaNet baselines by 5.2–11.9 mAP points and the closest classroom-specific detector, an s-scale re-implementation of SCB-YOLO, by 2.5 points; recording-grouped five-fold stratified cross-validation further yields 91.5 ± 0.4% mAP@0.5 and 89.8 ± 0.4% macro-F1 (the paired per-fold gain over YOLO11s is positive on every fold, 5.0–5.5 mAP points), while the complete video pipeline sustains 33.6 FPS, providing a lightweight and accurate solution for intelligent classroom behavior monitoring.
No comments yet — start the discussion below.