Multimodal KV Cache Quantization for Long-Video Understanding
Developing modality-aware KV cache quantization for long-video MLLMs, exploiting temporal redundancy across video tokens and distinct characteristics of visual and text KV caches. The goal is to enable memory-efficient long-video understanding while preserving model performance.