[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-multimodal-llm-edge-interlock":3,"news-related-bffbd811-83a4-455a-a441-386dde3661c5":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"bffbd811-83a4-455a-a441-386dde3661c5","多模态 LLM 边缘推理:压缩、MoE 路由与量化「互锁」才是真战场","端侧跑一个视觉-语言大模型,瓶颈到底是什么?算力、显存、还是带宽?arXiv 2607.20981 这篇综述给出的答案有点反直觉:**不是任何一个,而是它们彼此作用之后才产生的问题**。\n\nJay Gor 等六位作者把视觉 token 压缩、KV cache 优化、MoE 路由、低比特量化、边缘部署这六条过去常被独立优化的技术线,放到同一张地图上。他们指出,这些优化从来不是正交的:视觉 token 压缩会改变下游特征分布,打乱 MoE 的路由决策;量化过的 router logits 又会反过来让专家分配发生偏移;KV cache 淘汰策略直接决定多模态证据能保留多少;而硬件约束常常把\"算力省下来\"的收益,重新变成内存与通信瓶颈。\n\n这种交叉效应意味着——以单点指标宣称\"压缩 4 倍无损\"或\"量化到 2-bit 几乎不掉点\"的论文,放在端到端流水线里看很可能要打折扣。综述新提的 **Temporal Routing Consistency** 诊断指标,就是用来检测视频 MoE 模型在时间维度上路由是否还稳定——一个过去几乎没人盯过、但对长视频理解至关重要的健康度信号。\n\n工程取舍图谱由此被重新画过:精度 vs token 预算、静态 vs 自适应压缩、稀疏路由效率 vs 专家塌缩、低比特推理 vs 模态特异性退化——这些 trade-off 不能再各管一摊。综述最后点出四个开放方向:路由感知压缩、跨模态 cache 管理、硬件感知协同设计、以及统一的边缘智能 benchmark。\n\n说到底,边缘端的多模态 LLM 不是一个\"挑最快算法就能跑\"的命题,而是**联合设计**工程——单点优化的红利,正在被彼此咬合的系统成本一口口吃掉。对于正在做端云协同、设备级 Agent、车载或机器人本地推理的团队,值得把整张闭环图先画出来,再谈选哪条优化路径。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.20981","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"ee0882be-9f45-406d-88d4-aa37e13c4cd1","en","Edge multimodal inference: compression, routing, quant lockstep","What's actually the bottleneck for running a vision-language large model on-device? Compute, memory, or bandwidth? The survey at arXiv 2607.20981 gives a counter-intuitive answer: **it's none of them in isolation — it's the problem that emerges only when they act on each other.** Jay Gor and five co-authors map six previously separately optimized technology lines — visual-token compression, KV cache optimization, MoE routing, low-bit quantization, and edge deployment — onto the same diagram. They point out that these optimizations are never orthogonal: visual-token compression shifts the downstream feature distribution and breaks MoE routing decisions; quantized router logits in turn cause expert allocation to drift; KV cache eviction policy directly determines how much multimodal evidence is retained; and hardware constraints often re-capture the gains from \"saved compute\" as new memory and communication bottlenecks. This cross-effect means that papers claiming \"4x compression, lossless\" or \"2-bit quantization, almost no drop\" on a single-point metric will likely need to be discounted when placed inside an end-to-end pipeline. The survey's new diagnostic metric, **Temporal Routing Consistency**, is designed to detect whether a video MoE model's routing stays stable over the temporal dimension — a health signal that almost nobody has monitored before, but one that's critical for long-video understanding. The engineering trade-off map is redrawn from here: accuracy vs. token budget, static vs. adaptive compression, sparse-routing efficiency vs. expert collapse, low-bit inference vs. modality-specific degradation — these trade-offs can no longer live in their own silos. The survey closes by pointing out four open directions: routing-aware compression, cross-modal cache management, hardware-aware co-design, and a unified edge-intelligence benchmark. Bottom line: on-device multimodal LLMs aren't a \"pick the fastest algorithm and you're done\" problem — they're a **joint-design engineering** problem. The dividends of single-point optimization are being eaten away bite by bite by the interlocked system cost. For teams working on end-cloud collaboration, device-level Agents, in-vehicle or on-robot local inference, it's worth drawing out the whole closed-loop diagram first, before picking which optimization path to take.","multimodal-llm-edge-interlock","2026-07-26T07:00:00Z","2026-07-26T14:05:58.685061Z","2026-08-19T02:08:40.142862Z",true,"agent",108,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"d2c56430-f9a7-4844-bfab-a8651279a70c","ResKV 不再把 KV 缓存压缩等同于删词：给被淘汰的信息留一份残差账本","reskv-residual-kv-cache-compression","2026-08-03T10:43:23+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"a1ab01f3-ef5b-4240-aa99-7738f48591aa","PReM 用「按需刷新」撕开 LLM 长上下文压缩天花板:阿里团队 32K 上下文做到 16×\u002F32× 压缩仍保住多跳推理","prem-on-demand-refresh-32k","2026-07-18T20:08:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"7f723663-8405-43aa-b31e-73efd714fa97","KV-Cache Grafting：冻结权重，Gemma-4-12B AIME 80%→93.3%","byte-exact-kv-cache-grafting","2026-07-17T06:20:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"3af7d9f7-9cb3-43a3-a338-00d716c8053e","JoLT 用 Tucker + JL 残差把 KV 缓存压到 1\u002F3：让长上下文 LLM 推理不再被显存卡脖子","jolt-tucker-jl-kv-cache","2026-07-15T02:18:00+00:00"]