Duc-Hung Nguyen, Cong-Huu Hoang, Dinh-Nhat Loi, Duc-Anh Nguyen, Phạm Ngọc Hùng · International Journal of Pattern Recognition and Artificial Intelligence 2026 · 2026
DOI: 10.1142/s0218001426550189
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Real-Time Detection Transformer (RT-DETR) is the first detection transformer-based model to achieve real-time performance. However, the concatenation operation treats all scale features equally without adaptive weighting, leading to insufficient multi-scale feature fusion. In addition, Generalized Intersection over Union (GIoU) does not explicitly consider the center distance and aspect ratio between the predicted bounding boxes and the ground-truth boxes, leading to slow convergence. To address the above problems, this paper proposes Tri-Focal DEtection TRansformer (TriDETR) by replacing the GIoU loss function with the Tri-Focal Complete Intersection over Union (CIoU) loss to optimize overlap and assign weights, allowing the model to focus on difficult patterns. Additionally, TriDETR uses Selective Boundary Aggregation (SBA) to better fuse multi-scale features. Experiments were conducted on PASCAL VOC 2007, Foggy Cityscapes, and ACDC datasets to demonstrate the effectiveness of the proposed method. Specifically, in terms of mAP 50:95 , TriDETR outperforms RT-DETR with the same backbone by 0.80-0.98 on PASCAL VOC, 0.20-1.23 on Foggy Cityscapes, and 0.30-0.95 on ACDC (snow, rain, and nighttime). Additionally, TriDETR outperforms the best-performing YOLO26 variant while maintaining approximately the same frames per second.
No comments yet — start the discussion below.