[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen-3-5-deepstack-vit-multilayer":3,"news-related-009ab609-27bf-4c34-92a6-2ce53eb25b69":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"009ab609-27bf-4c34-92a6-2ce53eb25b69","Qwen 3.5 原生多模态新思路：DeepStack Vision Transformer 多层特征融合解析","阿里 Qwen 3.5 的视觉编码器设计走了一条不同于主流的路线。传统 Vision Transformer 如 LLaVA、MiniGPT 等，主要依赖单层输出特征，再通过投影层与语言模型对齐。Qwen 3.5 的 DeepStack Vision Transformer 则采用了多层级特征融合——将编码器多个中间层的特征进行整合，而非只看最后一层输出。同时，它用 Conv3D 将视频作为第三维度处理，实现原生时序建模，而非事后拼接帧序列。这种设计的核心收益在于：细粒度纹理与全局语义不再对立，可以同时保留。对于视频问答、时序推理等任务，提升效果显著。更重要的是，这套视觉编码器不是独立外挂的模块，而是直接融入了语言模型的多模态链路，体现了「原生多模态」的设计取向——从架构层面而非后训练对齐来解决融合问题。","https:\u002F\u002Fqwen.ai\u002Fblog?id=qwen3.5","c36a21ac-2a77-421b-9519-1e150695732a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"e676a5cf-1f24-472f-a765-86fa21a1bc3c","ai-model",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"b6aa453d-b37d-4b1d-8af5-adeebc972c59","en","DeepStack ViT: multi-layer fusion for Qwen 3.5 multimodality","Alibaba's Qwen 3.5 vision encoder design takes a different path from the mainstream. Traditional Vision Transformers like LLaVA, MiniGPT, etc., mainly rely on single-layer output features, then align with the language model via a projection layer. Qwen 3.5's DeepStack Vision Transformer uses multi-layer feature fusion — integrating features from multiple intermediate layers of the encoder, rather than just looking at the final layer's output. Meanwhile, it uses Conv3D to treat video as a third dimension for processing, achieving native temporal modeling rather than post-hoc frame-stitching. The core benefit of this design: fine-grained texture and global semantics are no longer opposed, both can be preserved. For video Q&A, temporal reasoning, and other tasks, the improvement is significant. More importantly, this vision encoder isn't an independent external module, but is directly integrated into the language model's multimodal chain, embodying a \"native multimodal\" design orientation — solving the fusion problem at the architecture level rather than via post-training alignment.","qwen-3-5-deepstack-vit-multilayer","2026-05-09T13:10:00Z","2026-05-09T13:13:11.215255Z","2026-08-19T02:08:40.142862Z",true,"agent",165,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f5c0b227-faf9-47e3-863f-3c102365cd41","LongCat-Next 开源：把文字、图像和声音统一成离散 Token","longcat-next-discrete-native-multimodal","2026-08-09T08:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"40095b51-97b0-4fd4-9b1d-f636c970572e","阿里 Qwen 团队发布 Qwen3.8-Max:2.4 万亿参数 MoE 模型首度开放权重","qwen3-8-max-2-4t-moe-open-weights","2026-08-07T02:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"804b44fb-66c6-4355-83a4-b3a03a776d2a","Inkling-Small 开放权重：12B 激活参数换来更高 Agent 效率，也暴露事实性短板","inkling-small-multimodal-moe-efficiency","2026-08-05T16:32:13+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"7cc1b87c-fe06-495a-9c01-9516d0c16354","腾讯混元 HunyuanImage-3.0 全面开源：80B 总参 \u002F 13B 激活的自回归 MoE，把多模态理解和生图拉到同一框架","tencent-hunyuanimage-3-moe-autoregressive","2026-08-05T01:00:00+00:00"]