Chao Ding, Kun Zou, Yu Zhong, Yong Liu, Yanan Zhang · Eng—Advances in Engineering 2026 · 2026
DOI: 10.3390/eng7090492
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Industrial surface inspection must localize defects when annotated data are limited. We present Asymmetric CNN–ViT Residual Fusion (ACVRF), in which a convolutional neural network (CNN) generates proposals and primary region-of-interest features, while a Vision Transformer (ViT) supplies a residual with a bounded mixing coefficient on the same regions. Frozen expert losses determine training-only sample weights. Across three fusion-stage seeds with fixed experts, ACVRF obtained 0.78442 ± 0.00222 mean average precision at an intersection-over-union threshold of 0.50 (mAP50) and 0.37404 ± 0.00089 mAP50:95 on NEU-DET-pro, compared with 0.75955 and 0.35440 for CNN. These results use validation-selected checkpoints. Matched controls support a contribution from the trained ViT/SFP branch, while ACVRF exceeded the strongest tested concatenation baseline by only 0.00458 mAP50. On a supplementary GC10 split with separate validation and test subsets and three independently trained expert pairs, mean test mAP50 was 0.39751 for ACVRF and 0.36804 for CNN. This historically used data pool does not provide new external validation, and the GC10 evidence for a Transformer-specific contribution remains limited. Teacher weighting slightly increased mean mAP50 but reduced mean mAP50:95 on both protocols. Model-only inference required 43.787 ms per image versus 17.936 ms for CNN. These results support CNN-anchored residual fusion within the evaluated protocols, with modest gains over the strongest tested fusion control and a substantial computational cost.
No comments yet — start the discussion below.