[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ovis-omni-embedding-3b":3,"topics-all":38,"news-related-d78ea8a8-bebb-4b94-b340-127eb2874a73":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"d78ea8a8-bebb-4b94-b340-127eb2874a73","Ovis 全模态嵌入 3B:综合分领先 5.19","阿里 ATH-MaaS 团队发布 Ovis-Omni-Embedding-3B,基于 Qwen2.5-Omni-3B 改造的全模态嵌入模型,覆盖文本、图像、视频、音频与视觉文档检索。官方自报 MMEB-v3 聚合分 58.46,领先最强对比基线 5.19 分,音频组优势最大达 6.91;权重暂未开源。","做检索的人一直面对同一个尴尬:文本用一套模型,图像换一套,音频再换一套,跨模态检索要么拼接多个嵌入向量,要么干脆放弃。9 月 21 日,阿里 ATH-MaaS 团队在 arXiv 公开[技术报告](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25165)(编号 2609.25165),放出一个试图终结这种碎片化的方案——Ovis-Omni-Embedding-3B,一个覆盖文本、图像、视频、音频、视觉文档乃至交错多模态输入的统一嵌入模型。\n\n## 不拼塔,直接用全模态底座\n\n大多数多模态嵌入模型的思路是\"组装\":文本塔加视觉塔,各自编码再对齐。Ovis 团队反其道而行,直接拿预训练好的 Qwen2.5-Omni-3B 全模态模型当底座,砍掉语音生成用的 Talker 模块和语言建模头,保留原生文本 tokenizer、视觉编码器、音频编码器和共享的 Thinker 主干——最后一个非填充 token 的末层隐藏状态,直接拿来当检索向量。\n\n训练侧的三板斧也都围绕\"统一\"展开:对比学习加低秩初始化做适配;构造覆盖文本、图像、视频、音频和交错多模态数据的高质量语料,用同源采样保证批内负样本够\"难\";损失函数用 focal loss 强调困难样本,再加一个基于相似度的嵌入蒸馏,从互补专家模型转移细粒度相似结构。推理时的低秩特征分解则允许弹性降维——2048 维可以压到 1024、512、256 甚至 128,性能损失有限。\n\n## 分数说话:音频组领先最多\n\n官方自报的成绩单([GitHub README](https:\u002F\u002Fgithub.com\u002FATH-MaaS\u002FOvis-Omni-Embedding)):在覆盖 190 个数据集的 MMEB-v3 上,Ovis-Omni-Embedding-3B 综合 58.46 分,比最强对比基线的 53.27 高出 5.19 分,六个模态组全部第一。分组看,音频组优势最大(50.08 vs 43.17,领先 6.91 分),Agent 检索组次之(45.52 vs 39.42,领先 6.10),图像组反而只领先 3.72 分。在完整对比表的 31 个条目里,它 22 个第一、8 个第二,唯一掉出前二的是 MultiConIR。\n\n三个补充基准同样值得看:音频嵌入基准 MAEB 57.29(对比 LCO-Embedding-Omni-7B 的 53.54),视频基准 MVEB 61.77(对比 57.58)。但在纯文本检索的 RTEB 上,它对 Qwen3-Embedding-4B 的优势只有 67.35 比 67.27——0.08 分的差距,基本等于噪声。\n\n## 冷水与看点\n\n两盆冷水先泼。第一,这些分数全部来自团队自报,评测是本地跑分后插入对应榜单快照,尚无第三方复现。第二,也是更关键的:README 明确写着模型权重\"暂未开源,将于近期放出\"。一个声称统一全模态检索的模型,如果最终权重不放或放得抠门,行业影响力会大打折扣——3B 参数、2048 维嵌入的方案,只有在别人真能部署时才有意义。\n\n看点同样有两个。其一,3B 参数在嵌入赛道属于轻量级,却能同时吃下六种模态,如果权重如期放出,对做 RAG、多模态知识库、Agent 记忆检索的团队是现成的地基。其二,弹性降维设计(2048→128)直接对标 Matryoshka 类方案,给了存储敏感场景一个官方退路。\n\n所以呢?嵌入模型正在从\"每个模态一个专家\"走向\"一个底座通吃\",Ovis 证明了全模态底座改造成检索模型的路走得通,但 0.08 分的文本险胜和未放出的权重都在提醒:统一是趋势,护城河还没建起来。权重放出来那天,这篇报告才真正值得再看一遍。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25165","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"a6d9881c-1b58-48e2-b35a-6e8883acfc58","en","Ovis-Omni-Embedding-3B: One Backbone, Every Modality","Alibaba ATH-MaaS released Ovis-Omni-Embedding-3B, a universal embedding model on Qwen2.5-Omni-3B covering text, image, video and audio. MMEB-v3: 58.46, +5.19.","Anyone building retrieval systems lives with the same awkwardness: one model for text, another for images, a third for audio. Cross-modal retrieval means either stitching together multiple embedding vectors or giving up entirely. On September 21, Alibaba's ATH-MaaS team published a [technical report on arXiv](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25165) (2609.25165) proposing a way out of this fragmentation: Ovis-Omni-Embedding-3B, a unified embedding model covering text, images, video, audio, visual documents, and interleaved multimodal inputs.\n\n## No Tower Assembly, Just an Omni Backbone\n\nMost multimodal embedding models follow an \"assembly\" approach: a text tower plus a vision tower, each encoding separately before alignment. The Ovis team went the opposite way. They took a pretrained Qwen2.5-Omni-3B omni-modal model as the foundation, removed the speech-generation Talker module and the language-modeling head, and kept the native text tokenizer, vision encoder, audio encoder, and the shared Thinker backbone. The final-layer hidden state at the last non-padding token is used directly as the retrieval embedding.\n\nThe training recipe has three planks, all centered on \"unified\": contrastive training with low-rank initialization for adaptation; a broad, high-quality corpus spanning text, images, video, audio, and interleaved multimodal data, with homogeneous-source sampling ensuring informative in-batch negatives; and focal loss to emphasize hard examples plus similarity-based Embedding Distillation transferring fine-grained similarity structure from complementary experts. At inference time, low-rank feature decomposition allows flexible dimensionality — 2048 dimensions can compress to 1024, 512, 256, or even 128 with limited performance loss.\n\n## The Numbers: Audio Leads the Pack\n\nThe self-reported scorecard ([GitHub README](https:\u002F\u002Fgithub.com\u002FATH-MaaS\u002FOvis-Omni-Embedding)): on MMEB-v3, which spans 190 datasets, Ovis-Omni-Embedding-3B scores 58.46 overall, 5.19 points above the strongest compared baseline's 53.27, ranking first on every modality group. By group, audio shows the largest margin (50.08 vs 43.17, +6.91), agent retrieval comes second (45.52 vs 39.42, +6.10), while the image group leads by only 3.72. Across the 31 aggregate and sub-task entries in the full comparison, it ranks first on 22 and second on 8. MultiConIR is the only entry where it falls outside the top two.\n\nThree supplementary benchmarks are worth noting: 57.29 on MAEB (vs LCO-Embedding-Omni-7B's 53.54) for audio embedding, and 61.77 on MVEB (vs 57.58) for video. But on RTEB, a text-only retrieval benchmark, its edge over Qwen3-Embedding-4B is just 67.35 vs 67.27 — a 0.08-point gap that is essentially noise.\n\n## Cold Water and Silver Linings\n\nTwo buckets of cold water first. All scores are self-reported by the team; evaluations were run locally and inserted into the corresponding leaderboard snapshots, with no third-party replication yet. Second, and more critical: the README explicitly states model weights are \"not open-sourced yet\" and will be released in the near future. A model claiming to unify omni-modal retrieval loses much of its industry impact if the weights arrive late or stingy — a 3B-parameter, 2048-dimension embedding only matters if others can actually deploy it.\n\nThere are two silver linings. First, 3B parameters is lightweight for the embedding track, yet it swallows six modality groups at once; if weights land as promised, it is a ready-made foundation for RAG, multimodal knowledge bases, and agent memory retrieval. Second, the elastic dimensionality design (2048 down to 128) directly parallels Matryoshka-style approaches, offering an official escape hatch for storage-sensitive deployments.\n\nSo what? Embedding models are moving from \"one expert per modality\" toward \"one backbone for everything.\" Ovis shows the omni-backbone-to-retrieval path works, but the 0.08-point text squeaker and the unreleased weights both remind us: unification is the trend, the moat is not built yet. The day the weights actually drop, this report becomes worth a second read.","ovis-omni-embedding-3b","2026-09-23T21:08:45Z","2026-09-23T21:09:46.958739Z","2026-09-23T21:09:46.958748Z",true,"agent",923,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"12c67d52-17a2-4df5-8386-35d18ffd221a","JEPA-Anything:一套预测框架打通七个领域,湿实验也给了背书","jepa-anything-orthogonal-predictive-factorization","2026-09-19T23:10:37+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"d8bc7b5e-9eb0-475e-91b7-5a3390d2c6a6","2026年开源LLM爆发：Meta、阿里、Google竞相发布新一代模型","open-source-llm-boom-2026-q1-meta-alibaba-google","2026-04-24T04:06:08+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"33bd7c4c-27f8-458c-8404-265134fc6ce8","视频生成缺的不是算力,是记忆:282 篇论文拼出一张全景地图","ar-video-generation-memory-survey","2026-09-24T21:09:28+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"5d3c50e8-5087-43e9-a8f1-c64973f712c1","Qwen 拆掉 ASR 管道:音视频原生对话靠合成数据练成","qwen-omnivchat-native-audio-visual-dialogue","2026-09-21T15:14:15+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"c99751d5-418e-49d5-99d3-e43b84c80ec7","IBM与NASA开源月球基础模型:Lunar Foundation Model","nasa-ibm-lunar-foundation-model-sombench","2026-09-19T09:30:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"28c41f06-d20f-481c-b133-cd109af3aed1","答对之后停不下来:微软团队揪出在线蒸馏的 EOS 错配元凶","eos-mismatch-opd-length-inflation","2026-09-18T21:09:06+00:00"]