External and internal semantics for visual understanding

Open Access
Authors
Supervisors
Cosupervisors
Award date 29-09-2026
ISBN
  • 9789465346007
Series SIKS Dissertation Series, 2026-34
Number of pages 174
Organisations
  • Faculty of Science (FNWI) - Informatics Institute (IVI)
Abstract
Visual understanding lies at the core of computer vision, aiming to move beyond recognition toward representations that capture semantic structure. Achieving such understanding requires learning from labelled data and prior knowledge, but also leveraging knowledge and relationships present in raw data. Understanding visual data thus requires learning representations that incorporate externally defined semantics, such as human-annotated hierarchies, and internally available semantics, embedded in the data. This thesis explores these complementary perspectives, external and internal semantics for visual understanding.
The first part, External Semantics, focuses on semantics defined outside the data, such as hierarchical relations between concepts. It investigates how such structures can guide learning in hyperbolic spaces, where hierarchical organization is captured. We study how to design hierarchies for hyperbolic embedding tasks, then introduce a method to restructure hierarchies to reduce geometric distortion, and finally show how hierarchical hyperbolic learning can be extended to continual learning scenarios. Together, these works demonstrate how external semantic organization enhances representation learning and increases the likelihood of making hierarchically meaningful mistakes.
The second part, Internal Semantics, examines semantics directly extracted from the data. It studies how semantic relations can be discovered from visual and multi-modal signals without relying on external data sources. The first work explores relative visual composition, learning how image parts relate through self-supervised objectives. The second work extends this idea to multi-modal composition, discovering visual named entities by aligning visual and textual cues from captioned videos. These approaches reveal how intrinsic relationships within data can be leveraged to form meaningful representations.
Document type PhD thesis
Language English
Downloads
Permalink to this page
cover
Back