谭国豪, Jiahui Wan, T. Wen, Dongshuai Li, Xingxian Liu, Daoning Jiang, Xiaoyao Qiu, Yu-Liang Yan, Dan Ou, Haihong Tang, B. Zheng · ACM Conference on Recommender Systems (RecSys) 2026 · 2026
DOI: 10.1145/3773078.3831933
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Semantic retrieval in e-commerce search aims to identify a compact candidate set from billion-scale product catalogs with both high recall and low latency. Dual-Encoders dominate this stage due to their efficient dot-product similarity, but this formulation limits model expressiveness and fails to capture fine-grained relationships between queries and items. While prior work has explored interaction-based similarity, its additional cost often prevents deployment at industrial scale. We present TSSR-Beta (Taobao Search Semantic Retrieval Model - Beta), which improves the expressiveness of our production Dual-Encoder TSSR through a plug-in similarity module, termed the Hybrid Interaction Head. This module introduces fine-grained matching in the representation space through two complementary pathways: InteractMLP, which captures explicit matching patterns with residual MLP blocks, and InteractTrans, which models implicit cross-dimensional interactions with a Transformer layer. Their outputs are combined by a Fusion Head to produce the final similarity score. TSSR-Beta introduces only a small parameter overhead to the Dual-Encoder without changing its architecture, enabling it to be (1) pluggable, readily adapting to diverse Dual-Encoder backbones; (2) efficiently trainable, supporting large-batch contrastive learning with massive negative sampling; and (3) industrially deployable, preserving offline item pre-encoding and supporting low-latency online retrieval with Neighborhood-Aware Approximate Nearest Neighbor (NANN). Offline experiments on the Taobao Search show a +3.90pp Hitrate@500 improvement over TSSR. On public Natural Questions and WebQA, our module further improves Recall@1 by +4.60pp and +0.90pp over public Dual-Encoder, respectively. Deployed in Taobao Search, TSSR-Beta delivers low-latency billion-scale retrieval and achieves +0.63% transaction count and +2.69% GMV gains in online A/B tests.
No comments yet — start the discussion below.