Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-shot Semantic Segmentation
| Authors | |
|---|---|
| Publication date | 2025 |
| Book title | 2025 IEEE/CVF International Conference on Computer Vision |
| Book subtitle | ICCV 2025 : Honolulu, Hawaii, USA, 19-23 October 2025 : proceedings |
| ISBN |
|
| ISBN (electronic) |
|
| Event | 2025 IEEE/CVF International Conference on Computer Vision |
| Pages (from-to) | 21155-21165 |
| Publisher | Los Alamitos, California: IEEE Computer Society |
| Organisations |
|
| Abstract |
Generalized Few-Shot Semantic Segmentation (GFSS) aims to extend a segmentation model to novel classes with only a few annotated examples while maintaining performance on base classes. Recently, pretrained vision-language models (VLMs) such as CLIP have been leveraged in GFSS to improve generalization on novel classes through multi-modal prototypes learning. However, existing prototype-based methods are inherently deterministic, limiting the adaptability of learned prototypes to diverse samples, particularly for novel classes with scarce annotations. To address this, we propose FewCLIP, a probabilistic prototype calibration framework over multi-modal prototypes from the pretrained CLIP, thus providing more adaptive prototype learning for GFSS. Specifically, FewCLIP first introduces a prototype calibration mechanism, which refines frozen textual prototypes with learnable visual calibration prototypes, leading to a more discriminative and adaptive representation. Furthermore, unlike deterministic prototype learning techniques, FewCLIP introduces distribution regularization over these calibration prototypes. This probabilistic formulation ensures structured and uncertainty-aware prototype learning, effectively mitigating overfitting to limited novel class data while enhancing generalization. Extensive experimental results on PASCAL- 5i and COCO-20i datasets demonstrate that our proposed FewCLIP significantly outperforms state-of-the-art approaches across both GFSS and class-incremental setting. The code is available at https://github.com/jliu4ai/FewCLIP.
|
| Document type | Conference contribution |
| Note | With supplementary file |
| Language | English |
| Published at |
https://doi.org/10.48550/arXiv.2506.22979
(Submitted manuscript)
https://doi.org/10.1109/ICCV51701.2025.01965
(Final published version)
|
| Published at | |
| Downloads |
Liu_Probabilistic_Prototype_Calibration_of_Vision-language_Models_for_Generalized_Few-shot_Semantic_ICCV_2025_paper
(Accepted author manuscript)
Probabilistic_Prototype_Calibration_of_Vision-Language_Models_for_Generalized_Few-Shot_Semantic_Segmentation
(Embargo up to 2026-10-29)
(Final published version)
|
| Supplementary materials | |
| Permalink to this page | |