Elastic ViTs from Pretrained Models without Retraining

Open Access
Authors
Publication date 2025
Host editors
  • D. Belgrave
  • C. Zhang
  • H. Lin
  • R. Pascanu
  • P. Koniusz
  • M. Ghassemi
  • N. Chen
Book title 39th Annual Conference on Neural Information Processing Systems (NeurIPS 2025)
Book subtitle 2-7 December 2025, San Diego, California, USA and 30 November-5 December 2025, Mexico City, Mexico
ISBN (electronic)
  • 9798331338275
Series Advances in Neural Information Processing Systems
Event 39th Annual Conference on Neural Information Processing Systems
Pages (from-to) 27652-27682
Publisher Neural Information Processing Systems Foundation
Organisations
  • Faculty of Science (FNWI) - Informatics Institute (IVI)
Abstract
Vision foundation models achieve remarkable performance but are only available in a limited set of pre-determined sizes, forcing sub-optimal deployment choices under real-world constraints. We introduce SnapViT: single-shot network approximation for pruned Vision Transformers, a new post-pretraining structured pruning method that enables elastic inference across a continuum of compute budgets. Our approach efficiently combines gradient information with cross-network structure correlations, approximated via an evolutionary algorithm, does not require labeled data, generalizes to models without a classification head, and is retraining-free. Experiments on DINO, SigLIPv2, DeIT, and AugReg models demonstrate superior performance over state-of-the-art methods across various sparsities, requiring less than five minutes on a single A100 GPU to generate elastic models that can be adjusted to any computational budget. Our key contributions include an efficient pruning strategy for pretrained Vision Transformers, a novel evolutionary approximation of Hessian off-diagonal structures, and a self-supervised importance scoring mechanism that maintains strong performance without requiring retraining or labels. Code and pruned models are available at: https://elastic.ashita.nl/
Document type Conference contribution
Language English
Published at
https://doi.org/10.48550/arXiv.2510.17700 (Accepted author manuscript)
https://doi.org/10.52202/085713-0824 (Final published version)
Published at
Other links
Downloads
085713-0824open (Final published version)
Permalink to this page
Back