Loading…
Benchmarking Pre-Trained Vision-Language Models for Bidirectional Image-Text Retrieval on MS-COCO: BERT+ResNet, CLIP, and BLIP · Researchar