DeokHyun You, Seongbok Baik, Yong-Geun Hong · Applied Sciences 2026 · 2026
DOI: 10.3390/app16188977
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Camera-only BEV 3D object detectors are trained under highly imbalanced category distributions, and their matched object queries can exhibit directional semantic errors toward frequent classes. We investigate this behavior as a diagnostic problem: given fixed geometric predictions and fixed query–ground-truth assignments, how much class information remains accessible in frozen decoder features, which errors can be recovered, and where does recovery fail? We establish a scene-disjoint protocol in which recovery fitting and model selection use separate subsets of the official nuScenes training set, while all 150 validation scenes (6019 samples) remain final-only until all model and post-processing choices are fixed. Geometry-only Hungarian matching produces 158,253 fixed positive pairs on the full validation set. The frozen detector obtains a macro accuracy of 0.6136 on these pairs, while a lightweight factorized head trained on frozen features from decoder layer 4 reaches 0.6959 ± 0.0012 across three seeds. A linear probe achieves a macro accuracy of 0.8628 on the internal tuning split, whereas a shuffled-label control remains at chance (0.1000), indicating that substantial class information remains decodable from the frozen features. Tail-focused analysis further shows that recovered errors are more separable in frozen feature space than unrecovered errors across all 15 class-by-seed comparisons. However, recovery is not consistently observed across the controlled ResNet-18 and ResNet-50 configurations, and locked end-to-end evaluation decreases mAP from 0.2565 to 0.1620 and NDS from 0.3582 to 0.2796. These results support a geometry-controlled diagnosis of partial and class-dependent semantic recoverability, rather than improved localization, architecture-independent recovery, or deployable detection performance.
No comments yet — start the discussion below.