Skip to main content

Frame Skipping Architecture for Video-Language Model Acceleration

Conferences
Shan, H; Wei, C; Guo, C; Fu, Y; Liang, T; Li, H; Chen, Y
Published in: Glsvlsi 2026 Proceedings of the Great Lakes Symposium on VLSI 2026
June 22, 2026

Video-Language Models (VLMs) have achieved strong performance in video understanding, enabling a wide range of applications. However, the high computational and memory costs of Multimodal Large Language Models (MLLMs), combined with the rapid growth in sequence length introduced by video inputs, hinder efficient deployment, particularly on resource-constrained edge platforms. To address these challenges, we propose a modular extension for VLM accelerators that incorporates a frame-skipping unit prior to MLLM processing. This module dynamically selects informative frames and suppresses redundant ones, effectively reducing unnecessary computation while preserving task-relevant information. The proposed design operates in conjunction with existing VLM accelerator backbones, providing an efficient and edge-friendly solution for video-language tasks. Experimental results show that our method achieves up to 2.42 × speedup and 2.46 × improvement in energy efficiency compared to the baseline, while incurring only a small accuracy degradation, demonstrating a favorable trade-off between efficiency and performance.

Duke Scholars

Altmetric Attention Stats
Dimensions Citation Stats

Published In

Glsvlsi 2026 Proceedings of the Great Lakes Symposium on VLSI 2026

DOI

Publication Date

June 22, 2026

Start / End Page

824 / 829
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Shan, H., Wei, C., Guo, C., Fu, Y., Liang, T., Li, H., & Chen, Y. (2026). Frame Skipping Architecture for Video-Language Model Acceleration. In Glsvlsi 2026 Proceedings of the Great Lakes Symposium on VLSI 2026 (pp. 824–829). https://doi.org/10.1145/3787109.3816385
Shan, H., C. Wei, C. Guo, Y. Fu, T. Liang, H. Li, and Y. Chen. “Frame Skipping Architecture for Video-Language Model Acceleration.” In Glsvlsi 2026 Proceedings of the Great Lakes Symposium on VLSI 2026, 824–29, 2026. https://doi.org/10.1145/3787109.3816385.
Shan H, Wei C, Guo C, Fu Y, Liang T, Li H, et al. Frame Skipping Architecture for Video-Language Model Acceleration. In: Glsvlsi 2026 Proceedings of the Great Lakes Symposium on VLSI 2026. 2026. p. 824–9.
Shan, H., et al. “Frame Skipping Architecture for Video-Language Model Acceleration.” Glsvlsi 2026 Proceedings of the Great Lakes Symposium on VLSI 2026, 2026, pp. 824–29. Scopus, doi:10.1145/3787109.3816385.
Shan H, Wei C, Guo C, Fu Y, Liang T, Li H, Chen Y. Frame Skipping Architecture for Video-Language Model Acceleration. Glsvlsi 2026 Proceedings of the Great Lakes Symposium on VLSI 2026. 2026. p. 824–829.

Published In

Glsvlsi 2026 Proceedings of the Great Lakes Symposium on VLSI 2026

DOI

Publication Date

June 22, 2026

Start / End Page

824 / 829