
Nguyen Bao Tran, Huu Quang Hoa Nguyen, Hoang Loc Tran · Engineering Technology & Applied Science Research 2026 · 2026
DOI: 10.48084/etasr.20469
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This paper studies ways to integrate pretrained visual features into the Neural Collaborative Filtering (NeuMF) architecture for item cold-start recommendation under implicit feedback. We propose Vis-NeuMF for warm-start and ColdStartVisNeuMF for item cold-start, both extending NeuMF with frozen ResNet-50 visual features through branch-specific projection layers and Layer Normalization (LayerNorm) fusion. The cold-start variant removes item identifier (ID) embeddings and represents cold items purely through visual features, avoiding the uninformative item-embedding noise that affects ID-based collaborative filtering when test items have no positive training interactions. An initial comparison against the strong visual baseline Visual Bayesian Personalized Ranking (VBPR) appears to favor it, but this is confounded by objective mismatch: VBPR uses a pairwise ranking loss (BPR) while the NeuMF variants use the pointwise Binary Cross-Entropy (BCE). To separate the loss effect from the architecture effect, we adopt a controlled comparison matrix (BCE/BPR by keep/remove item ID embedding). On four categories and five paired random seeds from the Amazon Reviews 2023 dataset, switching ColdStartVisNeuMF from BCE to BPR increases mean full-catalog cold-start Hit Ratio at 10 (cold_all HR@10) from 0.126 to 0.198. At matched BPR loss, the average cold_all HR@10 difference between ColdStartVisNeuMF and VBPR is +0.017; paired tests find no significant difference in any category, whereas ColdStartVisNeuMF is significantly stronger on cold_only HR@10 in three of the four categories. The main contribution of this paper is a controlled empirical characterization showing that, in visual cold-start, the training objective accounts for most of the apparent performance difference between architectures, and that a simple visual NeuMF without item ID embeddings is competitive with the baseline once the objective is matched.
No comments yet — start the discussion below.