Benchmark Granularity and Model Robustness for Image-Text Retrieval

Mariya Hendriksen; Shuo Zhang; Ridho Reinanda; Mohamed Yahya; Edgar Meij; Maarten de Rijke

doi:https://doi.org/10.1145/3726302.3730290

Benchmark Granularity and Model Robustness for Image-Text Retrieval A Reproducibility Study

Authors	Mariya Hendriksen Shuo Zhang Ridho Reinanda Mohamed Yahya Edgar Meij Maarten de Rijke
Publication date	2025
Book title	SIGIR '25
Book subtitle	Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval : July 13-18, 2025, Padua, Italy
ISBN (electronic)	9798400715921
Event	48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2025
Pages (from-to)	3183-3193
Number of pages	11
Publisher	New York, NY: Association for Computing Machinery
Organisations	Faculty of Science (FNWI) - Informatics Institute (IVI)
Abstract	Image-Text Retrieval (ITR) systems are central to multimodal information access, with Vision-Language Models (VLMs) showing strong performance on standard benchmarks. However, these benchmarks predominantly rely on coarse-grained annotations, limiting their ability to reveal how models would perform under real-world conditions, where query granularity varies. Motivated by this gap, we examine how dataset granularity and query perturbations affect retrieval performance and robustness across four architecturally diverse VLMs (ALIGN, AltCLIP, CLIP, and GroupViT). Using both standard benchmarks (MS-COCO, Flickr30k) and their fine-grained variants, we show that richer captions consistently enhance retrieval, especially in text-to-image tasks, where we observe an average improvement of 16.23%, compared to 6.44% in image-to-text. To assess robustness, we introduce a taxonomy of perturbations and conduct extensive experiments, revealing that while perturbations typically degrade performance, they can also unexpectedly improve retrieval, exposing nuanced model behaviors. Notably, word order emerges as a critical factor - contradicting prior assumptions of model insensitivity to it. Our results highlight variation in model robustness and a dataset-dependent relationship between caption granularity and perturbation sensitivity and emphasize the necessity of evaluating models on datasets of varying granularity.
Document type	Conference contribution
Language	English
Published at	https://doi.org/10.1145/3726302.3730290 (Final published version)
Other links	https://www.scopus.com/pages/publications/105011818669
Downloads	3726302.3730290 (Final published version)
Permalink to this page

Back

UvA-DARE

Digital Academic Repository

Benchmark Granularity and Model Robustness for Image-Text Retrieval A Reproducibility Study