CTRL-O: Language-Controllable Object-Centric Visual Representation Learning

Open Access
Authors
Publication date 2025
Book title 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition : CVPR 2025
Book subtitle Nashville, Tennessee, USA, 11-15 June 2025 : proceedings
ISBN
  • 9798331543655
ISBN (electronic)
  • 9798331543648
Event 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025
Pages (from-to) 29523-29533
Publisher Los Alamitos, California: IEEE Computer Society
Organisations
  • Faculty of Science (FNWI) - Informatics Institute (IVI)
Abstract
Object-centric representation learning aims to decompose visual scenes into fixed-size vectors called "slots" or "object files", where each slot captures a distinct object. Current state-of-the-art object-centric models have shown remarkable success in object discovery in diverse domains, including complex real-world scenes. However, these models suffer from a key limitation: they lack controllability. Specifically, current object-centric models learn representations based on their preconceived understanding of objects, without allowing user input to guide which objects are represented. Introducing controllability into object-centric models could unlock a range of useful capabilities, such as the ability to extract instance-specific representations from a scene. In this work, we propose a novel approach for user-directed control over slot representations by conditioning slots on language descriptions. The proposed CONTROLLABLE OBJECT-CENTRIC REPRESENTATION LEARNING approach, which we term CTRL- O, achieves targeted object-language binding in complex real-world scenes without requiring mask supervision. Next, we apply these controllable slot representations on two downstream vision language tasks: text-to-image generation and visual question answering. The proposed approach enables instance-specific text-to-image generation and also achieves strong performance on visual question answering.
Document type Conference contribution
Note With supplementary material
Language English
Published at
Published at
Other links
Downloads
Supplementary materials
Permalink to this page
Back