[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ml-embed-3d-matryoshka-low-resource":3,"news-related-5e4ce9dc-9454-4e4c-997d-467617d00fee":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"5e4ce9dc-9454-4e4c-997d-467617d00fee","打破高质量嵌入的「不可能三角」：ML-Embed 三维 Matryoshka 框架直击低资源语言痛点","文本嵌入模型已广泛用于 RAG、语义搜索等场景，但高质量嵌入长期面临「不可能三角」：计算成本、覆盖语种、模型透明度三者难以兼得。ICML 2026 录用论文 ML-Embed 试图打破这一困局。\n\nML-Embed 提出了三维 Matryoshka 学习框架（3D-ML），在模型全生命周期三个维度同时优化：MRL（Matryoshka 表征学习）减少存储开销，MLL（Matryoshka 层学习）支持推理时按需调整深度，MEL（Matryoshka 嵌入学习）提升参数效率。模型参数量从 1.4 亿到 80 亿，在 430 个任务上完成了评估，在 MTEB 基准的 17 个子集中刷新了 9 项纪录，尤其在低资源语言上的表现超出预期。\n\n更值得注意的，是团队选择了全面开源模型、数据和代码。在当前 embedding 服务普遍依赖闭源 API 的背景下，这为学术研究和中小开发者提供了一条低成本的入场路径。\n\n从工程视角看，ML-Embed 的三层解耦设计值得借鉴——存储、推理、参数效率分别优化，最终在端侧部署场景的可行性显著提升。如何在保持多语言覆盖的同时控制推理延迟，仍是后续研究的关键课题。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.15081","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"26c9ae34-4f2e-4ce5-95e2-d415757b5eb2","en","ML-Embed: 3D Matryoshka embeddings for low-resource languages","Text embedding models are widely used in RAG, semantic search, and other scenarios, but high-quality embeddings have long faced an \"impossible triangle\": compute cost, language coverage, and model transparency are hard to combine. The ICML 2026 paper ML-Embed attempts to break this deadlock.\n\nML-Embed proposes a three-dimensional Matryoshka learning framework (3D-ML), optimizing across three dimensions of the model lifecycle: MRL (Matryoshka Representation Learning) reduces storage overhead, MLL (Matryoshka Layer Learning) supports on-demand depth adjustment at inference, and MEL (Matryoshka Embedding Learning) improves parameter efficiency. The model has parameter counts from 140M to 8B, evaluated across 430 tasks, setting 9 new records across 17 MTEB benchmark subsets — with particularly impressive performance on low-resource languages.\n\nWhat's even more noteworthy is the team's choice to fully open-source the model, data, and code. Against the backdrop of embedding services widely relying on closed-source APIs, this provides a low-cost onramp for academic research and small-to-mid developers.\n\nFrom an engineering perspective, ML-Embed's three-layer decoupling design is worth learning from — storage, inference, and parameter efficiency are optimized separately, and the feasibility of edge deployment scenarios is significantly improved. How to maintain multilingual coverage while controlling inference latency remains a key issue for follow-up research.","ml-embed-3d-matryoshka-low-resource","2026-05-17T04:10:00Z","2026-05-17T04:08:48.545993Z","2026-08-19T02:08:40.142862Z",true,"agent",109,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"c84f5d0b-65d1-411c-8f76-75c301a748b2","多模态AI的token成本困局：Image Prompt Packaging带来推理降本新思路","image-prompt-packaging-multimodal-35-91pct","2026-05-26T04:08:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"ea444bd9-4683-486b-b606-c222d98f1ba7","标注即 rollout:南开 OraRL 把视频多模态 RL 训练成本砍半,9B 空间智能超 GPT-5","orarl-annotations-as-rollouts-video-rl","2026-08-26T17:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"c94bdf86-5de9-49fe-8c98-0f5c47611bfe","SGLang v0.5.18 发布:大模型冷启动提速 2.38 倍,710 个 PR 都改了什么","sglang-v0-5-18-cold-start-2-38x","2026-08-24T23:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"92433e6b-113a-4ada-af77-fbb8995a9850","LFM2.5-DSpark 开源:300M 草稿模型让端侧推理快 2.87 倍,输出零损耗","lfm2-5-dspark-draft-models","2026-08-21T21:10:00+00:00"]