36Kr published a deep dive into Xiaomi's MiMo-V2.5 inference optimization techniques. The five techniques — speculative decoding, KV cache compression, MoE expert scheduling, dynamic batching, and quantization — combine to reduce inference cost by 60% while maintaining quality, allowing Xiaomi to offer aggressive pricing without margin loss.