Every token stuffs its K/V cache into HBM, making the "long CoT + high concurrency" combination for inference LLM the real testing ground for KV cache compression research. Han Shen and Yuyang Wu of Carnegie Mellon, in their Kara (arXiv:2607.01237), take on the two old problems of "window boundary" and "retention granularity" and give the cleanest set of solutions currently available. Kara's core is to compress only the most recently generated context window — this avoids the repeated compression overhead of SnapKV/AdaKV-class "threshold-trigger + full-window re-scoring"; more importantly, Kara uses bidirectional attention rather than the single-direction backward-looking query to score KV pairs, letting retention candidates span forward and backward positions, no longer dominated by prefix position. Then the Token2Chunk module extends candidate discrete KV pairs into "arbitrary-length continuous chunks", both retaining the pointer property of discrete key tokens and preserving the semantic continuity of chunks — this neatly fills the "rigid boundary" short board of ChunkKV. In the KvLLM framework landed on PagedAttention, the design uses a periodic-trigger strategy instead of threshold-trigger, directly sidestepping the "compression overhead actually lowers throughput" concurrency-throughput inversion problem. Experiments on Qwen3-4B/14B and DeepSeek-R1-Distill-Llama-8B show that Kara at 30% retention nearly maintains uncompressed accuracy on MATH-500, AIME24, AMC23, and the retrieval performance on NIAH also significantly outperforms ChunkKV and AdaKV. Opinion: Kara's bidirectional scoring + flexible chunk combination essentially upgrades "KV retention" from a one-dimensional sorting problem to a two-dimensional layout problem. This upgrade lets 7B/14B-tier inference models running high-concurrency workloads on 8×H100/H200 for the first time have the possibility of "compression without losing accuracy, throughput can still rise". For deployers, KvLLM's periodic-trigger strategy is more suitable for long-CoT inference services than SnapKV-class threshold-trigger — this path is worth following, but engineering landing still depends on whether the synchronization overhead of PagedAttention across nodes is masked by the periodic trigger.