[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gemini-3-1-ultra-2m-context-rag-pushed":3,"news-related-a4201ba3-a84a-4711-a6b6-6436d121a122":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"a4201ba3-a84a-4711-a6b6-6436d121a122","Gemini 3.1 Ultra 发布：200万 token 上下文将 RAG 推下神坛","Gemini 3.1 Ultra 将上下文窗口扩展至 200万 token，是 Google 再次刷新自己保持的纪录。这一数字意味着它可一次性吞下约1500页文本或30000行代码，绝大多数真实场景下，根本不再需要 RAG 管道。\n\nRAG（检索增强生成）长期以来是处理长文档的标配方案：切分、Embedding、召回、拼接，每一步都有信息丢失风险。200万 token 上下文改变了这个逻辑——当模型能直接消化所有 token 时，召回这一步就成了多余。\n\nGoogle 在 Gemini 3.1 Ultra 中采用了稀疏 MoE（混合专家）架构，这是头部厂商在长上下文赛道的一致选择：不是让所有参数参与每次推理，而是只激活与当前任务相关的专家模块，在保持高质量生成的同时控制了推理成本。Gemini 3.1 Ultra 原生多模态能力（文本、图像、音频、视频统一处理）进一步拓展了长上下文的应用边界。\n\n对行业而言，RAG 并未消亡，但它的必要性正在被重新评估。当模型能直接读完一整本技术手册、整个代码仓库时，开发者需要重新思考：哪些场景真的需要检索，哪些只是因为模型太慢、上下文太短而习以为常的权宜之计。当然，在知识频繁更新的场景下，RAG 仍有不可替代的价值——但那种上下文不够长的焦虑，确实可以缓解了。","https:\u002F\u002Fdeepmind.google\u002Fmodels\u002Fmodel-cards\u002Fgemini-3-1-ultra\u002F","35ce748f-48b7-4638-88ef-effa57a7e749",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a9524a82-a7c5-4daa-bb4b-a7ee77bb0b94","gemini",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"8cf7490f-2449-4ba7-be19-61befa0d92b4","google",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"792b6f20-5ae8-4891-bc8b-4bea0fa5acac","en","Gemini 3.1 Ultra released: 2M-token context pushes RAG off its throne","Google DeepMind released Gemini 3.1 Ultra on June 2, with a 2-million-token context window. At this scale, the traditional RAG (retrieval-augmented generation) pattern becomes less necessary — the model can fit the entire knowledge base in its context, querying directly. This may mark a major shift in how AI applications handle large document collections.","gemini-3-1-ultra-2m-context-rag-pushed","2026-06-02T06:01:00Z","2026-06-02T10:04:33.390200Z","2026-08-19T02:08:40.142862Z",true,"agent",89,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"34edaffc-6b5c-4df1-9e2f-d864cada6063","Gemini 走进 K-12 课堂：Google 把「上下文」塞进每个作业","gemini-classroom-k12-contextualized-prompts","2026-08-07T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"6f033005-3b0b-41f4-8e0b-125cb9cc0c0d","Gemini 3.1与Google压缩算法：AI效率革命的双重突破","gemini-3-1-google-compression-april-2026-roundup","2026-04-22T01:03:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7b9cdf6e-5ef0-4ece-ab6c-e8cec1b02397","Google 重组 DeepMind 领导层,Gemini 研发提速应对 Anthropic 与 OpenAI 竞争","google-deepmind-reshuffle-gemini-speed","2026-08-25T07:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"bcedeb8e-e5eb-4bbc-98b8-ea12f869055f","Google 收编 DeepMind：25 年最大 AI 重组","google-deepmind-centralization-gemini-catchup","2026-08-14T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"4bd8e8bd-7066-4ab7-bd97-e24ea3921395","Gemini 因编程落后推迟两月:Brin 4 月督促背后,Google 把研发「收回到一个人」手里的组织账本","google-gemini-coding-behind-deepmind-reshuffle","2026-08-14T03:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"16856034-439d-4915-aed4-80b42ae09c68","Gemini 3.7 Flash：FrontierCode 43.6%，价格腰斩","gemini-3-7-flash-coding-agent-fast","2026-08-13T09:00:00+00:00"]