[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-wemm-embedding-wechat-multimodal":3,"news-related-4c7f5330-3aff-458a-9ef5-f04cc5585703":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","微信视觉团队开源 WeMM-Embedding 全模态嵌入家族(2B\u002F4B\u002F9B,Apache 2.0),文本、图像、视频、视觉文档对齐进同一向量空间:2B 在 MMEB-v2 反超此前 8B 开源基线,9B 官方报告 80.6 为表内最高,已通过 14 个线上 A\u002FB 验证。","检索和推荐的地基是嵌入模型——你刷到的每一条视频号、每一篇公众号文章,背后都是向量在算距离。大模型时代大家盯着对话和生成,嵌入这一层往往被忽视,但它恰恰是把多模态内容接进大系统的转接头。微信视觉团队(WeChat Vision)这次把自家生产环境里的全模态嵌入家族 WeMM-Embedding 整个开源了:2B、4B、9B 三档,Apache 2.0 许可,权重上了 Hugging Face,代码和评测全在 GitHub(arXiv: 2608.24053)。\n\n## 一张表里的跨度\n\n按官方技术报告,WeMM-Embedding 支持文本、图像、视频、视觉文档和任意交错多模态输入,统一映射到同一向量空间,嵌入取自专用 \u003Cembedding> token 最后一层隐状态再做 L2 归一化。训练分两阶段:先大规模多模态对齐,再用精选数据精调,其中包括细粒度相关性监督与跨尺度知识转移。\n\nMMEB-v2(78 个数据集)上的官方结果:2B 版拿到 77.9 的总分,反超此前领先的 8B 开源基线——对照组里 Qwen3-VL-Embedding 8B 是 77.8,VLM2Vec 8B 只有 53.2。9B 版总分 80.6,其中图像 81.9、视频 74.3、视觉文档 83.3,是该表内的最高分(表内两个未公开权重的闭源提交 DME-Small\u002FMedium 也没有超过它)。这些是官方口径的 benchmark 数字,复现代码已随仓库放出。\n\n## 套娃维度与部署细节\n\n工程上有两个值得注意的点。一是 Matryoshka(套娃)表示:三档模型分别支持 64 到 2048\u002F2560\u002F4096 的可变输出维度,官方数据显示 2B 模型截到 256 维仍保留全维度图文性能的 98.7%——意味着向量库存储成本可以压掉一个数量级而几乎不损失检索质量。二是推理栈:官方验证了 vLLM 0.27.0 和 SGLang 0.5.9 两种 serving 方案,推荐 transformers 5.2.0,评测管线里视频按 64 帧采样。\n\n## 更严苛的 MMEB-v3 与自曝短板\n\n团队同时放了 MMEB-v3 的成绩:190 个任务,含 78 个 v2 任务、53 个文本任务、47 个 agent 任务、11 个音频任务和 MCMR,不支持的任务直接计零分。WeMM 2B\u002F4B\u002F9B 分别拿到 56.0、58.2、59.5,都高于表内同档的 Qwen3-VL-Embedding(2B 50.9、8B 53.5),也压过 E5-Omni 和 Omni-Embed-Nemotron。但有意思的是音频列:WeMM 全线 0 分,因为模型目前不支持音频输入——官方把这条短板原样写进表格,而不是藏起来。\n\n## 从跑分到朋友圈\n\n真正的分量在于线上验证。技术报告称,这套模型在 26 项内部基准上取得显著增益,并通过 14 个线上 A\u002FB 测试的持续检验,目前部署在微信的视频号、公众号、朋友圈和电商等推荐搜索场景。也就是说,这不是一个只在榜单上好看的学术模型,而是在微信核心场景里跑过全链路的工业件。\n\n## 所以呢\n\n嵌入模型是 RAG 和 Agent 时代的水电煤:多模态理解做得再好,检索层接不住就到不了用户手里。把一套生产级全模态嵌入连权重带评测全开源,等于给多模态检索生态补了一块地基——而 MMEB-v3 那一列显眼的 0 分也提醒我们:全模态的\"全\",永远差一块。\n\n参考:arXiv 2608.24053(huggingface.co\u002Fpapers\u002F2608.24053)与 github.com\u002FTencent\u002FWeMM-Embedding。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.24053","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"655a1a2c-b8d8-4304-9012-47a40167d762","en","WeChat Open-Sources WeMM-Embedding: 2B Beats 8B, 9B Hits 80.6","WeChat open-sources WeMM-Embedding (2B\u002F4B\u002F9B, Apache 2.0) for text, image, video and docs; 2B beats the old 8B baseline, 9B hits 80.6 on MMEB-v2.","Retrieval and recommendation run on embedding models — every video and article your feed serves you is a vector distance computation under the hood. While the spotlight sits on chat and generation, embeddings are the adapter that wires multimodal content into large systems. Tencent's WeChat Vision team has now open-sourced its production multimodal embedding family, WeMM-Embedding: three sizes (2B, 4B, 9B) under Apache 2.0, with weights on Hugging Face and code plus evaluation on GitHub (arXiv: 2608.24053).\n\n## The Span in One Table\n\nPer the official technical report, WeMM-Embedding supports text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs, mapped into one shared vector space; embeddings are taken from the last-layer hidden state at a dedicated \u003Cembedding> token and L2-normalized. Training runs in two stages: large-scale multimodal alignment first, then refinement with curated data, fine-grained relevance supervision, and cross-scale knowledge transfer.\n\nOfficial MMEB-v2 results (78 datasets): the 2B model scores 77.9 overall, surpassing the previously leading 8B open-source baseline — Qwen3-VL-Embedding 8B sits at 77.8, VLM2Vec 8B at 53.2. The 9B model reaches 80.6 overall (image 81.9, video 74.3, visual document 83.3), the highest score in the table; the two closed-source submissions without public weights (DME-Small\u002FMedium) do not exceed it either. These are vendor-reported benchmark numbers, and the reproduction code ships with the repository.\n\n## Matryoshka Dimensions and Serving\n\nTwo engineering details stand out. First, Matryoshka representations: the three sizes support variable output dimensions from 64 up to 2048\u002F2560\u002F4096, and the 2B model retains 98.7% of its full-dimensional image and video performance when truncated to 256 dimensions — vector-store storage costs can drop by an order of magnitude with almost no retrieval-quality loss. Second, the inference stack: the team validated vLLM 0.27.0 and SGLang 0.5.9 serving, recommends transformers 5.2.0, and samples video at 64 frames in its evaluation pipeline.\n\n## The Harsher MMEB-v3, and a Self-Reported Weakness\n\nThe team also published MMEB-v3 results: 190 tasks, including the 78 v2 tasks, 53 text tasks, 47 agent tasks, 11 audio tasks, and MCMR, with unsupported tasks scored as zero. WeMM 2B\u002F4B\u002F9B score 56.0, 58.2, and 59.5 respectively, above Qwen3-VL-Embedding in the same table (2B 50.9, 8B 53.5), and above E5-Omni and Omni-Embed-Nemotron as well. Notably, the audio column reads all zeros for WeMM — audio input is not currently supported — and the team left that weakness in the table rather than hiding it.\n\n## From Benchmarks to Moments\n\nThe real weight is online validation. The report claims substantial gains on a 26-task in-house benchmark and consistent improvements across 14 online A\u002FB tests, with deployment across WeChat Channels, Official Accounts, Moments, and e-commerce recommendation and search. This is not a leaderboard-only academic model; it is an industrial component that has run the full pipeline inside WeChat's core surfaces.\n\n## So What\n\nEmbedding models are the plumbing of the RAG and agent era: multimodal understanding is worthless if the retrieval layer cannot carry it to users. Open-sourcing a production-grade universal multimodal embedder — weights and evaluation included — lays a foundation for multimodal retrieval ecosystems. And that glaring zero in the MMEB-v3 audio column is a reminder: \"universal\" always has one modality missing.\n\nReferences: arXiv 2608.24053 (huggingface.co\u002Fpapers\u002F2608.24053) and github.com\u002FTencent\u002FWeMM-Embedding.","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30Z","2026-08-26T21:08:56.549075Z","2026-08-26T21:08:56.549082Z",true,"agent",13,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"652b1484-31eb-4a83-a932-21fcf97b3a50","Boson AI Higgs Audio v3 TTS：4B 参数原生可控百语种语音生成","higgs-audio-v3-boson-4b-100-language","2026-06-04T18:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"12a5f49d-8c83-4c40-82af-1c0b7f1c8b3e","DeepSeek 给 V4-Flash 装上眼睛:Vision-Exp 实验模型两项基准反超 Opus 4.8","deepseek-v4-flash-vision-exp-multimodal","2026-08-21T23:05:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00"]