Ulugbek Hudayberdiev, Abdimumin Alikulov, Adkham Israilov, Muhiddin Xidirov, Javokhir Musaev · Journal of Imaging 2026 · 2026
DOI: 10.3390/jimaging12080397
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Landmark recognition for smart tourism is usually validated on curated benchmark images. In deployment, however, the classifier must handle user-generated photographs whose viewpoint, lighting, resolution, occlusion, and compression differ sharply from curated data. This paper evaluates a previously published multi-threshold enhancement and selective YOLO11n-cls ensemble under this shift, and provides a preliminary zero-shot comparison of three general-purpose multimodal large language models (MLLMs) on the same task. To measure the shift, we build Samarkand v2-SNS, a 300-image out-of-distribution test set of social-media photographs of 12 Samarkand landmarks, disjoint from the training and validation data. Under the shift, four supervised baselines fall by 12.73-22.08 percentage points to 73-80% accuracy, and their in-distribution ranking does not hold. The selective ensemble degrades least (99.24% to 93.00%, -6.24 points) and outperforms the strongest baseline by 13 points. A capacity-matched ablation shows that most of this robustness comes from enhancement diversity, not from generic ensembling. In a preliminary comparison, zero-shot MLLMs (GPT-5, Claude Sonnet 4.5, Gemini 2.5) reach only 24.81-54.26%, far below deployment needs. The results argue for reporting out-of-distribution accuracy alongside curated benchmarks, and for hybrid systems that pair compact specialised recognisers with MLLM-based interpretation.
No comments yet — start the discussion below.