CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
| Authors |
|
|---|---|
| Publication date | 2025 |
| Book title | 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition : CVPR 2025 |
| Book subtitle | Nashville, Tennessee, USA, 11-15 June 2025 : proceedings |
| ISBN |
|
| ISBN (electronic) |
|
| Event | 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 |
| Pages (from-to) | 29523-29533 |
| Publisher | Los Alamitos, California: IEEE Computer Society |
| Organisations |
|
| Abstract |
Object-centric representation learning aims to decompose visual scenes into fixed-size vectors called "slots" or "object files", where each slot captures a distinct object. Current state-of-the-art object-centric models have shown remarkable success in object discovery in diverse domains, including complex real-world scenes. However, these models suffer from a key limitation: they lack controllability. Specifically, current object-centric models learn representations based on their preconceived understanding of objects, without allowing user input to guide which objects are represented. Introducing controllability into object-centric models could unlock a range of useful capabilities, such as the ability to extract instance-specific representations from a scene. In this work, we propose a novel approach for user-directed control over slot representations by conditioning slots on language descriptions. The proposed CONTROLLABLE OBJECT-CENTRIC REPRESENTATION LEARNING approach, which we term CTRL- O, achieves targeted object-language binding in complex real-world scenes without requiring mask supervision. Next, we apply these controllable slot representations on two downstream vision language tasks: text-to-image generation and visual question answering. The proposed approach enables instance-specific text-to-image generation and also achieves strong performance on visual question answering.
|
| Document type | Conference contribution |
| Note | With supplementary material |
| Language | English |
| Published at |
https://doi.org/10.1109/CVPR52734.2025.02749
(Final published version)
|
| Published at | |
| Other links | |
| Downloads |
Didolkar_CTRL-O_Language-Controllable_Object-Centric_Visual_Representation_Learning_CVPR_2025_paper
(Accepted author manuscript)
CTRL-O_Language-Controllable_Object-Centric_Visual_Representation_Learning
(Final published version)
|
| Supplementary materials | |
| Permalink to this page | |