National University of Singapore
DYNAMIC NEURAL NETWORKS FOR EFFICIENT VISION MODEL INFERENCE
Abstract
dc:description.abstractRecent progress in dynamic neural networks is hindered by high training overhead from full-parameter fine-tuning and limited application beyond visual perception. This thesis addresses these challenges through three key contributions: Parameter-Efficient Fine-Tuning (PEFT): We demonstrate that dynamic networks can be developed with negligible trainable parameters, reducing FLOPs by 30% while significantly lowering training costs. Diffusion Transformers (DiT) Acceleration: We introduce a dynamic architecture for visual generation, achieving a 50% FLOPs reduction and 1.7× speedup with improved quality. Vision-Language Model (VLM) Acceleration: By implementing a dynamic small-large model cooperation mechanism, we achieve up to 3× faster inference while maintaining competitive performance. In conclusion, this research provides a comprehensive solution to reduce training costs and extends dynamic neural networks to visual generation and multi-modal domains, unlocking new possibilities for efficient real-world applications.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- ZHAO WANGBO