[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen3-7-text-embedding-launch":3,"news-related-40210e0d-84e3-460b-bde2-295b77573ab8":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"40210e0d-84e3-460b-bde2-295b77573ab8","Qwen3.7-Text-Embedding 上线:20% 检索增益、256-2560 可变维度,阿里把 RAG 的地基悄悄重浇了一遍","8 月 11 日,阿里通义实验室在 QwenCloud 上线基于 Qwen3.7 的多语言文本向量模型 qwen3.7-text-embedding。相比 text-embedding-v4,它在 MTEB 多语言、中英与代码检索任务上提升约 20%,支持 256-2560 自定义向量维度、131K 输入,定价 $0.07\u002F百万 token。这篇聊聊为什么 embedding 更新比新旗舰更值得开发者马上动手。","# Qwen3.7-Text-Embedding:没有发布会的更新,RAG 开发者却该马上动手\n\n8 月 11 日,阿里通义实验室在 QwenCloud 平台上线了新一代文本向量模型 **qwen3.7-text-embedding**。没有发布会,没有铺天盖地的通稿,只是 changelog 里短短一段——但对所有做检索增强生成(RAG)、语义搜索和企业知识库的团队来说,这类\"配角级\"更新的实际价值,往往比又一个新旗舰 LLM 更直接:它决定你的应用检索回来的上下文质量,而检索质量几乎等于 RAG 的天花板。\n\n## 官方口径的核心数字\n\n根据 QwenCloud 官方模型页与 changelog,这次发布的要点可以浓缩为几条:\n\n- **基于 Qwen3.7 训练的多语言统一文本向量模型**,由通义实验室出品;\n- 对比上一代 **text-embedding-v4**,在**文本检索、聚类、分类**三类任务上显著提升;\n- 在 **MTEB 多语言、中英、代码检索**等评测任务上取得 **20% 的提升**;\n- 支持 **256 到 2560 的自定义向量维度**;\n- 最大输入与上下文均为 **131K token**;\n- 定价 **$0.07 \u002F 百万 token**,速率上限为 1M TPM、2K RPM,并支持微调。\n\n## 可变维度:一个被低估的工程杠杆\n\n向量维度是 embedding 应用里最典型的权衡题:维度越高,表达精度越好,但存储成本和检索计算量也随之上涨;维度低则便宜、快,却容易丢信息。\n\nqwen3.7-text-embedding 把维度开放到 256-2560 由用户自定义,意味着同一个模型可以在一套系统里扮演多个角色——粗排召回用低维向量省存储、抢延迟,精排或高价值语料用高维向量保精度。对存量向量库,这也是一条低摩擦的升级路径:不必推翻重来,按场景逐步切换维度策略即可。\n\n## 131K 输入 + 代码检索增益\n\n另一个值得注意的是 **131K 的最大输入**。传统 embedding 模型的输入窗口普遍偏短,长文档必须切成小块再分别向量化,而分块策略的粗暴程度直接决定检索质量——一段被拦腰截断的上下文,检索回来也答不对题。长输入窗口让整篇长文档、完整代码文件直接进入向量化流程成为可能,分块这一 RAG 中最容易失真的环节有了被绕开的余地。\n\n对代码场景,官方明确列出了代码检索任务的提升。代码检索是 embedding 最难啃的场景之一——变量名、API 签名、跨文件的调用关系都不是自然语言模型天然擅长的。官方给出的 20% 评测提升覆盖了这一项,对做代码搜索、代码库问答的团队是个明确信号。\n\n## 冷静的补充\n\n两点需要保持克制的解读。第一,20% 是官方在自选评测任务上的口径,不同业务语料上的实际增益会有差异,上线前值得用自己的数据集做一轮回归评测。第二,官方页面并未给出与第三方开源 embedding 模型的横向对比数据,竞争格局如何,还需要社区榜单和实际项目来回答。\n\n## 所以呢\n\n大模型的聚光灯永远打在聊天旗舰上,但真正决定 AI 应用成败的,常常是检索这一层没人鼓掌的基础设施。$0.07\u002F百万 token 的定价加上 20% 的检索增益,qwen3.7-text-embedding 是那种\"升级成本一个下午、收益贯穿整个应用\"的改动。如果你手上有 RAG 系统在跑,这可能是本月性价比最高的一次迁移评估。\n\n参考:[QwenCloud 官方模型页](https:\u002F\u002Fwww.qwencloud.com\u002Fmodels\u002Fqwen3.7-text-embedding) 与 [QwenCloud Model Changelog](https:\u002F\u002Fdocs.qwencloud.com\u002Fchangelog\u002Fmodels)。","https:\u002F\u002Fwww.qwencloud.com\u002Fmodels\u002Fqwen3.7-text-embedding","c36a21ac-2a77-421b-9519-1e150695732a",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"e676a5cf-1f24-472f-a765-86fa21a1bc3c","ai-model",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"bf8b61ca-28ad-4d98-a5c8-72166021612c","en","Qwen3.7-Text-Embedding: +20% retrieval, variable dimensions","On August 11, Alibaba Tongyi Lab released qwen3.7-text-embedding, a multilingual text vector model built on Qwen3.7, on the QwenCloud platform. Compared to text-embedding-v4, it achieves roughly 20% improvements on MTEB multilingual, Chinese-English, and code retrieval tasks, supports customizable vector dimensions from 256 to 2560, accepts 131K input, and is priced at $0.07 per million tokens. This piece explains why an embedding update deserves developers' immediate attention more than yet another flagship LLM.","# Qwen3.7-Text-Embedding: An Update With No Keynote, Yet One RAG Developers Should Act On Immediately\n\nOn August 11, Alibaba's Tongyi Lab launched its new-generation text vector model, **qwen3.7-text-embedding**, on the QwenCloud platform. No launch event, no wall-to-wall press coverage — just a short entry in the changelog. But for every team building retrieval-augmented generation (RAG), semantic search, or enterprise knowledge bases, this kind of \"supporting-cast\" update often delivers more direct value than yet another flagship LLM: it determines the quality of the context your application retrieves, and retrieval quality is effectively the ceiling of RAG.\n\n## The Core Numbers, per the Official Announcement\n\nAccording to the official QwenCloud model page and changelog, the release boils down to a few key points:\n\n- A **multilingual unified text vector model trained on Qwen3.7**, produced by Tongyi Lab;\n- Significant improvements over the previous generation **text-embedding-v4** in **text retrieval, clustering, and classification**;\n- A **20% improvement** on evaluation tasks including **MTEB multilingual, Chinese-English, and code retrieval**;\n- Support for **customizable vector dimensions from 256 to 2560**;\n- A maximum input and context of **131K tokens**;\n- Pricing at **$0.07 per million tokens**, with rate limits of 1M TPM and 2K RPM, plus fine-tuning support.\n\n## Custom Dimensions: An Underrated Engineering Lever\n\nVector dimensionality is the classic trade-off in embedding applications: higher dimensions mean better representational precision, but storage costs and retrieval compute climb accordingly; lower dimensions are cheap and fast but risk losing information.\n\nBy opening up dimensions from 256 to 2560 for user customization, qwen3.7-text-embedding lets a single model play multiple roles within one system — low-dimensional vectors for coarse recall to save storage and latency, high-dimensional vectors for fine ranking or high-value corpora to preserve accuracy. For existing vector databases, this also offers a low-friction upgrade path: no need to tear everything down; you can shift dimension strategies per use case, step by step.\n\n## 131K Input + Code Retrieval Gains\n\nAnother notable point is the **131K maximum input**. Traditional embedding models generally have short input windows, forcing long documents to be split into chunks before vectorization — and how crudely you chunk directly determines retrieval quality. A passage cut in half mid-context won't answer the question correctly even when retrieved. A long input window makes it feasible to feed entire long documents or complete code files directly into vectorization, giving practitioners room to sidestep chunking — the most distortion-prone stage of RAG.\n\nFor code scenarios, the official announcement explicitly lists gains on code retrieval tasks. Code retrieval is one of the hardest arenas for embedding models — variable names, API signatures, and cross-file call relationships are not things natural-language models handle natively. The officially reported 20% evaluation improvement covers this area, a clear signal for teams building code search and codebase Q&A.\n\n## A Sober Caveat\n\nTwo points deserve restrained interpretation. First, the 20% figure is the official result on self-selected evaluation tasks; actual gains on different business corpora will vary, so a regression evaluation on your own dataset is worthwhile before switching over. Second, the official page does not provide head-to-head comparisons against third-party open-source embedding models, so where this lands in the competitive landscape remains a question for community leaderboards and real-world projects to answer.\n\n## So What\n\nThe spotlight in large language models always shines on chat flagships, but what truly determines the fate of AI applications is often the unglamorous infrastructure layer: retrieval. At $0.07 per million tokens with a 20% retrieval gain, qwen3.7-text-embedding is the kind of change where \"the upgrade costs an afternoon and the benefit runs through your entire application.\" If you have a RAG system in production, this may be the highest-ROI migration review you do this month.\n\nReferences: [QwenCloud official model page](https:\u002F\u002Fwww.qwencloud.com\u002Fmodels\u002Fqwen3.7-text-embedding) and [QwenCloud Model Changelog](https:\u002F\u002Fdocs.qwencloud.com\u002Fchangelog\u002Fmodels).","qwen3-7-text-embedding-launch","2026-08-14T13:10:00Z","2026-08-14T13:06:59.698104Z","2026-08-14T13:06:59.698113Z",true,"agent",139,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"37eb94e4-2155-416a-a8d9-bc5075541a27","Qwen-Image-3.0 发布:把文生图从「好看」推向「好用」","qwen-image-3-0","2026-07-21T08:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"7258978b-dfcd-4cb4-91c4-3b8569cd5deb","Qwen-Audio-3.0-TTS双版本发布:Plus登顶Artificial Analysis,Flash压到300ms首包延时","qwen-audio-3-tts","2026-07-20T10:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"12a5f49d-8c83-4c40-82af-1c0b7f1c8b3e","DeepSeek 给 V4-Flash 装上眼睛:Vision-Exp 实验模型两项基准反超 Opus 4.8","deepseek-v4-flash-vision-exp-multimodal","2026-08-21T23:05:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"d74a088d-e7f5-41cc-8c55-aedb10b8101d","Gemini-3-Pro 也只拿 66.4 分:南京大学开源全模态视频助手基准 OmniAssistBench","omniassistbench-omni-llm-video-assistant","2026-08-21T17:59:52+00:00"]