Skip to main content

Keyframe-Oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-form Video Processing

Conferences
Liu, Y; Sun, J; Lin, Y; Zhang, J; Yin, M; Wang, Q; Li, H; Chen, Y
Published in: Proceedings of the IEEE International Conference on Computer Vision
January 1, 2025

Vision language models (VLMs) demonstrate strong capabilities in jointly processing visual and textual data. However, they often incur substantial computational overhead due to redundant visual information, particularly in longform video scenarios. Existing approaches predominantly focus on either vision token pruning, which may overlook spatio-temporal dependencies, or keyframe selection, which identifies informative frames but discards others, thus disrupting contextual continuity. In this work, we propose KVTP (Keyframe-oriented Vision Token Pruning), a novel framework that overcomes the drawbacks of token pruning and keyframe selection. By adaptively assigning pruning rates based on frame relevance to the query, KVTP effectively retains essential contextual information while significantly reducing redundant computation. To thoroughly evaluate the long-form video understanding capacities of VLMs, we curated and reorganized subsets from seven datasets into a unified benchmark that highlights real-world scenarios with sparse but crucial events. Our experiments with VLMs of various scales show that KVTP can reduce token usage by 80% without compromising spatiotemporal and contextual consistency, significantly cutting computation while maintaining the performance. These results demonstrate our approach's effectiveness in efficient long-video processing, facilitating more scalable VLM deployment. The code for this paper is available at GitHub.

Duke Scholars

Altmetric Attention Stats
Dimensions Citation Stats

Published In

Proceedings of the IEEE International Conference on Computer Vision

DOI

EISSN

2380-7504

ISSN

1550-5499

Publication Date

January 1, 2025

Start / End Page

20802 / 20811
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Liu, Y., Sun, J., Lin, Y., Zhang, J., Yin, M., Wang, Q., … Chen, Y. (2025). Keyframe-Oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-form Video Processing. In Proceedings of the IEEE International Conference on Computer Vision (pp. 20802–20811). https://doi.org/10.1109/ICCV51701.2025.01934
Liu, Y., J. Sun, Y. Lin, J. Zhang, M. Yin, Q. Wang, H. Li, and Y. Chen. “Keyframe-Oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-form Video Processing.” In Proceedings of the IEEE International Conference on Computer Vision, 20802–11, 2025. https://doi.org/10.1109/ICCV51701.2025.01934.
Liu Y, Sun J, Lin Y, Zhang J, Yin M, Wang Q, et al. Keyframe-Oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-form Video Processing. In: Proceedings of the IEEE International Conference on Computer Vision. 2025. p. 20802–11.
Liu, Y., et al. “Keyframe-Oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-form Video Processing.” Proceedings of the IEEE International Conference on Computer Vision, 2025, pp. 20802–11. Scopus, doi:10.1109/ICCV51701.2025.01934.
Liu Y, Sun J, Lin Y, Zhang J, Yin M, Wang Q, Li H, Chen Y. Keyframe-Oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-form Video Processing. Proceedings of the IEEE International Conference on Computer Vision. 2025. p. 20802–20811.

Published In

Proceedings of the IEEE International Conference on Computer Vision

DOI

EISSN

2380-7504

ISSN

1550-5499

Publication Date

January 1, 2025

Start / End Page

20802 / 20811