On the Challenges in Evaluating Visually Grounded Stories

Authors
Publication date 2025
Host editors
  • R. Campos
  • A.M. Jorge
  • A. Jatowt
  • S. Bhatia
  • M. Litvak
Book title Proceedings of Text2Story — Eighth Workshop on Narrative Extraction From Texts
Book subtitle held in conjunction with the 47th European Conference on Information Retrieval (ECIR 2025) : Lucca, Italy, April 10, 2025
Series CEUR Workshop Proceedings
Event 8th Workshop on Narrative Extraction From Texts
Article number 19
Pages (from-to) 215-223
Number of pages 9
Publisher Aachen: CEUR-WS
Organisations
  • Interfacultary Research - Institute for Logic, Language and Computation (ILLC)
Abstract
Producing stories grounded in visual content is an inherent trait of human intelligence and an integral aspect of interpersonal communication. With the surge of advanced vision-to-language models, there has been increased interest in developing and understanding the capabilities of models to generate visually grounded narratives. However, recent research has highlighted the challenges in evaluating model-generated stories. In this work, we study these evaluation limitations in the visually grounded story generation task by focusing on the recently released Visual Writing Prompts dataset and shared task. Through this study, we also explore the capabilities of several general-purpose vision-to-language foundation models for generating stories grounded in sequences of images. We observe that some recent models, such as Qwen2.5-VL, can generate stories that are coherent, consistent, and well-grounded in the visual data. Nevertheless, in line with the recent studies in this area, we find that the existing automatic evaluation metrics and methods are insufficient in fully capturing all the aspects essential for assessing model-generated stories. We believe our findings reinforce the evidence and arguments emphasizing the need for improvements to automatic approaches that can comprehensively evaluate and understand models for visual storytelling.
Document type Conference contribution
Language English
Published at
https://ceur-ws.org/Vol-3964/paper19.pdf (Final published version)
Other links
Permalink to this page
Back