[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gemini-3-1-flash-lite-ga-multimodal-1-50":3,"news-related-386ce7fe-6fde-4d4a-8438-8b90f16bb963":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"386ce7fe-6fde-4d4a-8438-8b90f16bb963","Gemini 3.1 Flash-Lite 正式版发布：Google 最快最便宜的 Gemini 3 模型来了","**Google 于 5 月 7 日正式推出 Gemini 3.1 Flash-Lite 通用版本**，这是 Gemini 3 系列中速度最快、成本最低的模型，标志着 Google 在高效推理赛道上的最新落子。\n\n## 定位：速度与成本的极致平衡\n\nFlash-Lite 专为对延迟敏感、并发量大的企业场景打造，涵盖软件工程、客服、创意工具和金融等高实时性领域。Google 披露，该模型在分类任务上实现亚秒级响应，在高并发压力下 p95 延迟约为 1.8 秒，相较前代产品有显著提升。\n\n## 多模态能力落地\n\n值得注意的是，Flash-Lite 是 Gemini 3 系列中首款支持多模态（文本 + 图像）的 Lite 级别模型，支持工具调用（tool calling）和编排（orchestration）等 Agent 能力，标志着轻量级模型也能承载复杂 Agent 工作流。\n\n## 定价：再次拉低大模型使用门槛\n\nFlash-Lite 的定价为每百万输入 tokens 0.25 美元、每百万输出 tokens 1.50 美元，延续了 Google 近年来在高效率模型上持续压缩成本的策略。这也是 Google 面向大规模企业部署给出的最低单价方案。\n\n## 行业影响\n\nJetBrains、Gladly、Ramp 等企业已率先在生产环境中采用。Google 此番将 Flash-Lite 推至 GA（正式发布），既是对 Preview 阶段用户反馈的回应，也预示着今年 I\u002FO 大会上 Gemini 3.2 Flash 等更高端型号即将面世——Flash 系列正在成为 Google 覆盖企业需求的主力价格锚点。","https:\u002F\u002Fcloud.google.com\u002Fblog\u002Fproducts\u002Fai-machine-learning\u002Fgemini-3-1-flash-lite-is-now-generally-available","93669252-6081-4192-bf76-3a6814fc60cf",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a9524a82-a7c5-4daa-bb4b-a7ee77bb0b94","gemini",{"id":18,"name":19,"slug":19,"description":13,"color":13},"8cf7490f-2449-4ba7-be19-61befa0d92b4","google",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"c2214db0-a434-4d34-a7ac-0d7e2a0336b1","en","Gemini 3.1 Flash-Lite GA: the fastest, cheapest Gemini 3","**Google officially launched Gemini 3.1 Flash-Lite GA on May 7**, the fastest and lowest-cost model in the Gemini 3 series, marking Google's latest move in the high-efficiency inference race.\n\n## Positioning: The Ultimate Speed-Cost Balance\n\nFlash-Lite is purpose-built for latency-sensitive, high-concurrency enterprise scenarios, covering software engineering, customer service, creative tools, and finance. Google disclosed that the model achieves sub-second response on classification tasks, with p95 latency of about 1.8 seconds under high concurrency, a significant improvement over the previous generation.\n\n## Multimodal Capability Landing\n\nNotably, Flash-Lite is the first Lite-tier model in the Gemini 3 series to support multimodality (text + image), along with tool calling and orchestration Agent capabilities, marking that lightweight models can also carry complex Agent workflows.\n\n## Pricing: Lowering the LLM Bar Again\n\nFlash-Lite is priced at $0.25 per million input tokens and $1.50 per million output tokens, continuing Google's strategy of continuously compressing cost on high-efficiency models. This is also Google's lowest single-price solution for large-scale enterprise deployment.\n\n## Industry Impact\n\nCompanies including JetBrains, Gladly, and Ramp have been early adopters in production environments. Google's move to push Flash-Lite to GA is both a response to Preview-stage user feedback and a sign that more high-end models like Gemini 3.2 Flash are about to debut at this year's I\u002FO — the Flash series is becoming Google's anchor point for covering enterprise demand at various price points.","gemini-3-1-flash-lite-ga-multimodal-1-50","2026-05-08T11:04:00Z","2026-05-08T19:04:26.725264Z","2026-08-19T02:08:40.142862Z",true,"agent",105,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"6f033005-3b0b-41f4-8e0b-125cb9cc0c0d","Gemini 3.1与Google压缩算法：AI效率革命的双重突破","gemini-3-1-google-compression-april-2026-roundup","2026-04-22T01:03:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"7b9cdf6e-5ef0-4ece-ab6c-e8cec1b02397","Google 重组 DeepMind 领导层,Gemini 研发提速应对 Anthropic 与 OpenAI 竞争","google-deepmind-reshuffle-gemini-speed","2026-08-25T07:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"bcedeb8e-e5eb-4bbc-98b8-ea12f869055f","Google 收编 DeepMind：25 年最大 AI 重组","google-deepmind-centralization-gemini-catchup","2026-08-14T08:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"4bd8e8bd-7066-4ab7-bd97-e24ea3921395","Gemini 因编程落后推迟两月:Brin 4 月督促背后,Google 把研发「收回到一个人」手里的组织账本","google-gemini-coding-behind-deepmind-reshuffle","2026-08-14T03:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"16856034-439d-4915-aed4-80b42ae09c68","Gemini 3.7 Flash：FrontierCode 43.6%，价格腰斩","gemini-3-7-flash-coding-agent-fast","2026-08-13T09:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"8b6c20ec-7222-48cf-af2c-ac97466a2b0a","Gemini 月活破 10 亿:Google 第一次把 AI 助手做成自家「最快十亿用户产品」","gemini-app-1b-monthly-users","2026-08-12T03:00:00+00:00"]