arXiv 2606.10651 introduces Keye-VL-2.0, a 30B multimodal model from Kuaishou. The technical highlight: it's the first to adapt DSA (Dynamic Sparse Attention) to the GQA (Grouped-Query Attention) multimodal setting, achieving 4× inference speedup on long-video understanding with no quality loss.
The technical path: DSA is a recently proposed sparse-attention method that dynamically selects the most relevant KV cache entries per query. Keye-VL-2.0's contribution: extending DSA to GQA-Multimodal — i.e., applying DSA to both the visual encoder's attention and the LLM's cross-attention. The challenge is that the visual encoder has a different KV layout than the LLM, and the "relevance" criterion needs to be modality-aware.
The result: on the VideoMME long-video benchmark, Keye-VL-2.0-30B matches the quality of Qwen2.5-VL-72B (the previous SOTA) while running 4× faster. The speedup comes from two sources: 3× from DSA itself (sparse attention reduces compute), and 1.3× from GQA-friendly memory layout (the GQA structure means fewer KV entries need to be stored).
The bigger takeaway: "sparse attention × multimodal × GQA" is the right combination for long-video understanding. The current SOTA models are all in the 70B+ range, and they're slow. Keye-VL-2.0 proves that with the right sparse-attention design, a 30B model can match a 72B model at 4× the speed. This is a significant result for the "video understanding at the edge" use case.
For the industry, the takeaway is that sparse attention is moving from "research curiosity" to "production must-have." Long-context workloads (video, code, book-length text) cannot be served with dense attention at scale, and DSA-style methods are the most promising direction.