The Curious Case of Visual Grounding: Different Effects for Speech-and Text-Based Language Encoders

Open Access
Authors
Publication date 2026
Book title ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Book subtitle proceedings : 4-8 May 2026, Barcelona, Spain
ISBN
  • 9798331567026
ISBN (electronic)
  • 9798331567019
Event 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing
Pages (from-to) 17917-17921
Publisher Piscataway, NJ: IEEE
Organisations
  • Interfacultary Research - Institute for Logic, Language and Computation (ILLC)
Abstract
How does visual information included in training affect language processing in audio- and text-based deep learning models? We explore how such visual grounding affects model-internal representations of words, and find substantially different effects in speech-vs. text-based language encoders. Firstly, global representational comparisons reveal that visual grounding increases alignment between representations of spoken and written language, but this effect seems mainly driven by enhanced encoding of word identity rather than meaning. We then apply targeted clustering analyses to probe for phonetic vs. semantic discriminability in model representations. Speech-based representations remain phonetically dominated with visual grounding, but in contrast to text-based representations, visual grounding does not improve semantic discriminability. Our findings could usefully inform the development of more efficient methods to enrich speech-based models with visually-informed semantics.
Document type Conference contribution
Language English
Published at
Downloads
2509.15837v1 (Submitted manuscript)
The_Curious_Case_of_Visual_Grounding_Different_Effects_for_Speech-and_Text-Based_Language_Encoders (Embargo up to 2026-10-21) (Final published version)
Permalink to this page
Back