Yifei Wang · Applied and Computational Engineering 2026 · 2026
DOI: 10.54254/2755-2721/2026.36995
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This paper compares the cross-domain robustness of six established object detector configurations under a unified data protocol, a fixed training budget, and a single evaluation pipeline. All configurations are trained on the union of PASCAL Visual Object Classes (VOC) 2007 and VOC 2012 trainval, comprising 16551 images, under a fixed budget of 18 epochs of source-data exposure, with horizontal flipping at probability 0.5 as the only augmentation and Common Objects in Context (COCO)-pretrained complete detector initialisation. Each configuration is trained under three random seeds. Evaluation covers the VOC 2007 test set, Clipart1k, and Watercolor2k, with all predictions routed through the same pycocotools evaluator. Watercolor degradation is measured against a matched six-category VOC baseline. The Spearman rank correlation between in-domain and Clipart1k average precision at an intersection-over-union threshold of 0.50 is 0.257 across the six configurations. Single Shot Detector (SSD300) ranks fifth in-domain at 0.726 and first on both target domains, whereas Real-Time Detection Transformer (RT-DETR-l) ranks first in-domain at 0.831 and second on both. Observed cross-domain seed variability exceeds in-domain variability for every configuration by factors between 3.4 and 16.2.
No comments yet — start the discussion below.