Shubham Goel, Andrew Treadway, Yundi Jiang, Yiding Wen, Jianing Fu, Alan S. Li, Angli Liu, Antoine Simoulin, Himanshu Thakur, Haibo Zhang, Selahattin Akkas, Calvin Ma, Zhen Zeng, Yunyu He, Qin Huang, Benjamin Au, Guy Lebanon, Sagar Chordia · ACM Conference on Recommender Systems (RecSys) 2026 · 2026
DOI: 10.1145/3773078.3831887
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Video advertisements contain rich textual signal overlaid on the creative itself — promotional copy, prices, calls-to-action, brand and product mentions — that current production ranking systems largely under-exploit. We argue this is a high-yield, low-cost ranking modality and demonstrate a production deployment that brings it into Instagram’s ads ranking model at an Internet scale. Making this practical requires three components, each motivated by the modality’s deployment economics: (1) a specialized OCR backbone delivering near-VLM word-level quality at ∼ 80 × higher production throughput, which makes per-frame perception economical at an Internet scale; (2) a smart frame-extraction stage tailored to short-form video, combining a CTR-trained quality scorer with windowless-SSIM de-duplication, that recovers more OCR-relevant frames than uniform sampling at half the per-video frame budget; and (3) a feature stack — TF–IDF tokens and discrete RQ-VAE semantic IDs over a domain-tuned sentence-embedding backbone — that integrates into the ranker’s existing sparse embedding tables without a dedicated dense tower. We validate the stack with a four-lens ablation (LLM-as-judge with 200, 000-ad coverage, deterministic lexical-recall and reconstruction anchors, and a 5-rater human-vs-LLM agreement check with Spearman ρ = 0.74). The launched bundle delivers 0.055–0.06 offline relative NE gain aggregated across all surfaces served by the ranker — concentrated at ≈ 0.12% on short-form video (Reels-class) surfaces where promotional overlay text is densest — and a 0.06–0.08% NE gain in a multi-week 80%-traffic online A/B; per-feature attribution shows the 3-byte discrete codes (looked up through the ranker’s existing embedding tables) carry the majority of the lift at a 500 × smaller per-ad storage footprint than the dense backbone they were trained from, with the TF–IDF tokens contributing complementary high-recall lexical signal that the codes blur (brand names, numeric promotions, tail vocabulary).
No comments yet — start the discussion below.