[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-sber-gigachat-3-5-ultra":3,"news-related-222d6fbf-e9bc-4d63-8481-88ea28fd499c":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"222d6fbf-e9bc-4d63-8481-88ea28fd499c","Sber GigaChat 3.5 Ultra 开源：线性注意力 MoE 把长文本速度拉高 4 倍、模型尺寸砍半","俄罗斯最大银行 Sber 于 7 月 10 日发布开源大模型 GigaChat 3.5 Ultra，以「线性注意力 + MoE」双引擎组合正面挑战主流 transformer 路线。这是目前全球少数几个把线性注意力做到大模型规模并公开权重的开源项目。\n\n新模型的核心是 Sber 团队完全自研的线性注意力架构。传统 attention 每生成一个新 token 都要重新扫一遍整个上下文，计算量随长度二次增长；而线性注意力将上下文压缩为「摘要向量」，每次只需追加增量，使长文本场景的复杂度降到线性级别。官方数据显示，GigaChat 3.5 Ultra 在长文本上速度提升至 4 倍，而模型体积仅为前代的一半。\n\n在 MoE 架构加持下，新模型的参数总量据称是当前开源线性注意力模型里最大的之一。Sber AI 团队在训练中共完成 1500 次实验，并通过多轮人本数据筛选与清洗，显著提升了代码、数学、长文档理解以及 Agent 自主任务上的表现。官方称在多步推理和编程基准上已接近 DeepSeek 3.2，但模型尺寸近乎腰斩，意味着推理成本与硬件门槛显著降低。\n\nSber 高级副总裁 Anton Frolov 强调，该模型展示了「用更少资源训练强模型」的工程可行性，并已向全球开发者开放用于构建 Agent 服务。GigaChat 3.5 Ultra 现同步登陆 GigaChat 助手与 Hugging Face，在国际开源权重阵营中为非英语系国家模型拿下了一席之地。","https:\u002F\u002Fhuggingface.co\u002Fcollections\u002Fai-sage\u002Fgigachat-35","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"43fb642f-8ec2-4d5c-85d5-a87444042f97","en","Sber GigaChat 3.5 Ultra: linear-attention MoE, 4x faster","Russia's largest bank, Sber, on July 10 released the open-source large model GigaChat 3.5 Ultra, with a \"linear attention + MoE\" dual-engine combination challenging the mainstream transformer route head-on. This is currently one of the few open-source projects in the world to scale linear attention up to large-model scale and release public weights. The new model's core is a completely self-developed linear attention architecture by the Sber team. Traditional attention rescans the entire context for each new token generated, with compute growing quadratically with length; while linear attention compresses the context into a \"summary vector\", only adding increments each time, making long-text scenarios linear in complexity. Official data shows that GigaChat 3.5 Ultra lifts long-text speed 4×, while the model size is only half that of the previous generation. With MoE architecture added, the new model's total parameter count is reportedly one of the largest among current open-source linear-attention models. The Sber AI team completed 1500 experiments during training, and through multi-round human-data filtering and cleaning, significantly improved performance on code, math, long-document understanding, and Agent autonomous tasks. The official claim is that on multi-step reasoning and programming benchmarks it has approached DeepSeek 3.2, but the model size is nearly halved, meaning inference cost and hardware threshold are significantly reduced. Sber senior VP Anton Frolov emphasized that the model demonstrates the engineering feasibility of \"training strong models with fewer resources\", and has been opened to global developers for building Agent services. GigaChat 3.5 Ultra is now simultaneously available on GigaChat Assistant and Hugging Face, securing a seat in the international open-source-weight camp for non-English-system country models.","sber-gigachat-3-5-ultra","2026-07-10T18:05:00Z","2026-07-10T18:09:32.794347Z","2026-08-19T02:08:40.142862Z",true,"agent",139,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"5878a668-282c-4b88-b2b8-7eef40b7938c","LFM2.5-2.6B：2.5GB 内存跑本机 Agent 220 tok\u002Fs","lfm2-5-2-6b-on-device-agent","2026-08-11T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"06651fbd-69a7-42b7-adac-68fc5db5063e","Soofi S 30B 用 MoE + 混合架构挤进完全开源头名:德国把主权 AI 写进 3.2B 激活参数","soofi-s-30b-sovereign","2026-07-13T20:04:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"31259851-c64e-47dd-99cf-bcfae698b14f","LFM2.5-Retrievers：Liquid AI 把 LFM「单向」改成「双向 350M」，11 语种检索刷 SOTA","lfm-2-5-retrievers-liquid-350m-bidirectional","2026-06-22T03:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f89d097b-838b-4e1d-a5f6-dc2e6af67fb6","LFM2.5-8B-A1B 开源：1.5B 激活的 MoE 把「边缘 LLM」的天花板再抬一截","lfm-2-5-8b-a1b-liquid-edge-moe-1-5b","2026-06-12T10:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"d2d262c1-6bf0-4a95-b95f-8896fa226db3","腾讯混元 Hy-MT2 翻译家族开源：33 语言 + 1.25-bit 量化","tencent-hy-mt2-33-lang-1-25-bit-440mb","2026-05-22T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5ce7b0a3-0cfb-4603-8f0a-150afaf0aad9","开源大模型架构分化：MoE与Dense的技术路线之争","moe-vs-dense-open-source-llm-divergence-2026","2026-05-05T05:06:00+00:00"]