🔗 aidailypost.com/news/avllms-...
🔗 aidailypost.com/news/avllms-...
(1) Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
(2) OLA-VLM: Elevating Visual Perception in Multimodal LLMs with Auxiliary Embedding Distillation
🔍 More at researchtrend.ai/communities/MLLM
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
https://arxiv.org/abs/2412.09530
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
https://arxiv.org/abs/2412.09530
Abstract: Video large language models (VideoLLM) excel at video understanding, but face efficiency challenges due to the quadratic complexity of abundant visual tokens. Our systematic analysis of token [1/6 of https://arxiv.org/abs/2505.14454v1]
Abstract: Video large language models (VideoLLM) excel at video understanding, but face efficiency challenges due to the quadratic complexity of abundant visual tokens. Our systematic analysis of token [1/6 of https://arxiv.org/abs/2505.14454v1]
arxiv.org/html/2505.03...
arxiv.org/html/2505.03...
Read more: https://arxiv.org/html/2505.03829v1
Read more: https://arxiv.org/html/2505.03829v1
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
https://arxiv.org/abs/2511.08003
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
https://arxiv.org/abs/2511.08003
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
https://arxiv.org/abs/2508.17686
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
https://arxiv.org/abs/2508.17686
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
https://arxiv.org/abs/2508.15641
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
https://arxiv.org/abs/2508.15641
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
https://arxiv.org/abs/2506.21116
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
https://arxiv.org/abs/2506.21116