[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gemini-3-8-live-voice-s2s-number-one":3,"topics-all":41,"news-related-039ff515-68e7-4f11-866a-1da97e26eb45":60},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"039ff515-68e7-4f11-866a-1da97e26eb45","Gemini 3.8 Live 拿下 S2S 实时语音榜第一","Google 9 月 15 日上线 Gemini 3.8 Live 与 Extended Thinking 两款实时语音模型,实时语音质量指数 82.6 排第一,Big Bench Audio 97.7%,支持 97 种语言无缝切换并能边对话边后台执行工具。","9 月 15 日,Google 在博客里把两款新的实时语音模型一次性摆到了台面上——Gemini 3.8 Live 和 Gemini 3.8 Live Extended Thinking,定位是「最会说话的 Gemini」,从今天起在 Gemini API、Google AI Studio、Gemini Enterprise 和 Search Live 同步开放。\n\n## 拿数据说话:两个版本怎么分工\n\n两款模型都围绕「让对话更像和人聊天」这一件事展开。普通版 Gemini 3.8 Live 强调成本与可扩展性,在 Speech Agent Arena 上拿到了用户偏好第二名。Extended Thinking 版本则在复杂任务上发力——它在 Artificial Analysis 的 Speech to Speech Quality Index 上拿到 82.6 分排第一,在 Sierra 的 τ-Voice-banking 智能体基准上做到了 35.1%,在 Big Bench Audio 推理基准上拿到了 97.7%。ServiceNow 的 EVA-Bench 上,两个版本联手把语音智能体的准确率-体验帕累托前沿推到了新位置。\n\n## 模型能力:听、想、说、读、切换、后台工具,一次到位\n\n模型设计上,Google 这次给出了几个关键细节:实时视觉输入、近乎实时的语音响应,以及会话过程中自动检测并切换 97 种语言的能力。Extended Thinking 版本最值得注意的特性是「边想边说」——模型在后台跑多步推理的同时,用\"Let me check that...\"这类语言化提示保持对话流不被掐断,等到后台任务完成再继续推进。这意味着开发者在 Gemini Live API 上能直接搭出既能多步规划、又能在用户等待时持续聊天的语音智能体。工具调用也被彻底重写:两个版本都允许模型在后台跑工具调用和 API 请求而不打断语音对话,前端只用一句\"我先帮你查一下\"自然过渡。SynthID 水印自动嵌入所有音频输出,Google 同时披露了 Agora、LiveKit、Pipecat、Vercel、LangChain、Fishjam、Vision Agents 8 家已经基于 Gemini Live API 推出集成的生态合作伙伴。\n\n## 定价、行业意义与落地节奏\n\n值得拆开看的细节是定价。Google 在博客里没有公布具体数字,但 Artificial Analysis 的 cost-per-hour-of-input-audio 横轴和 Big Bench Audio 的性价比定位显示 Extended Thinking 处于「前沿但便宜」的象限。从行业角度看,这次发布把实时语音对话从「TTS + ASR + LLM 串联」的拼装时代推进到了端到端统一模型时代——Gemini 3.8 Live 同时处理听、想、说、读视觉、切换语言、后台跑工具,这些过去需要几个独立系统协作的环节现在都收进了一个模型权重里。对做客服、教育、陪伴类语音产品的团队来说,基础设施层面的拼装工作量会被显著压缩,真正的差异化只能往产品逻辑和用户体验上去做。\n\n模型卡同步在 Google DeepMind 的 model-cards 仓库公开,Safety & Responsibility 报告与 SynthID 水印策略一并发布,Gemini Enterprise 的 private preview 也在今天同步开启——对企业 IT 来说,这是首个不需要额外接 NLU\u002FASR\u002FTTS 三件套就能直接上线生产环境的语音模型。Salesforce、Genspark、Lumeris 这些已经把 3.8 Live 接进产品的合作伙伴,接下来会把这些能力推进到真实场景里去试。","https:\u002F\u002Fblog.google\u002Finnovation-and-ai\u002Fmodels-and-research\u002Fgemini-models\u002Fgemini-3-8-live-gemini-3-8-live-extended-thinking\u002F","d884df39-706a-45db-b7ed-371f12e54f1f",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"a9524a82-a7c5-4daa-bb4b-a7ee77bb0b94","gemini",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"8cf7490f-2449-4ba7-be19-61befa0d92b4","google",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":25,"name":26,"slug":26,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"2a681042-a8ac-4522-ab56-07aec49e3eea","en","Gemini 3.8 Live takes the top spot on the S2S real-time voice index","Google launched Gemini 3.8 Live and the Extended Thinking variant on Sept 15, with a 82.6 Speech-to-Speech Quality Index score, 97.7% on Big Bench Audio, 97-language mid-conversation switching, and background tool execution while the user keeps talking.","On September 15, Google put two new real-time voice models on the table at once — Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking — billed as \"the most conversational Gemini yet\", rolling out today across the Gemini API, Google AI Studio, Gemini Enterprise, and Search Live.\n\n## What the numbers say: how the two variants split the work\n\nBoth models orbit a single goal: making conversation feel like talking to a person. The standard Gemini 3.8 Live emphasizes cost and scalability, taking second place on user preference in the Speech Agent Arena. The Extended Thinking variant pushes into complex tasks — it tops Artificial Analysis's Speech to Speech Quality Index at 82.6, hits 35.1% on Sierra's τ-Voice-banking agentic benchmark, and lands 97.7% on the Big Bench Audio reasoning benchmark. On ServiceNow's EVA-Bench, both versions jointly push the accuracy-vs-experience Pareto frontier for voice agents to a new position.\n\n## Capabilities: listen, think, speak, see, switch, background-tool — all in one shot\n\nGoogle surfaced several concrete capabilities this round: real-time visual input, near-real-time voice response, and automatic detection and switching among 97 supported languages mid-conversation. The Extended Thinking variant's standout feature is \"thinking while talking\" — the model runs multi-step reasoning in the background while emitting language cues like \"Let me check that...\" to keep the conversation flowing, only resuming the full response when the background task completes. That means developers on the Gemini Live API can ship voice agents that do multi-step planning without ever freezing the live conversation. Tool calling was rewritten as well: both versions run tool calls and API requests in the background without interrupting the voice thread, with the front-end transitioning naturally on a single \"let me check that for you\". SynthID watermarking is baked into every audio output. Google also named eight ecosystem partners already shipping integrations on top of Gemini Live API — Agora, LiveKit, Pipecat, Vercel, LangChain, Fishjam, Vision Agents.\n\n## Pricing, industry meaning, and rollout cadence\n\nWorth pulling apart: pricing. Google didn't publish raw numbers in the blog, but Artificial Analysis's cost-per-hour-of-input-audio axis and the Big Bench Audio value-positioning both place Extended Thinking in the \"frontier but cheap\" quadrant. From an industry angle, this release moves real-time voice out of the \"TTS + ASR + LLM stitched together\" era into an end-to-end unified-model era — Gemini 3.8 Live simultaneously handles listening, reasoning, speaking, visual reading, language switching, and background tool execution, collapsing what used to require several independent systems into a single set of weights. For teams building customer service, education, and companion voice products, the infrastructure assembly work gets compressed significantly; differentiation now has to live in product logic and user experience.\n\nThe model card went live on Google DeepMind's model-cards repo alongside the Safety & Responsibility report and the SynthID policy. Gemini Enterprise's private preview also opened today — making this the first voice model enterprises can drop into production without bolting on NLU\u002FASR\u002FTTS separately. Salesforce, Genspark, and Lumeris, already shipping Gemini 3.8 Live integrations, will be the testbed for whether these capabilities survive real scenarios.","gemini-3-8-live-voice-s2s-number-one","2026-09-15T17:00:00Z","2026-09-17T05:04:47.672770Z","2026-09-17T05:04:47.672785Z",true,"agent",54,[42,51],{"slug":43,"tag_slug":43,"title_zh":44,"title_en":45,"intro_zh":46,"intro_en":47,"id":48,"is_active":38,"created_at":49,"modified_at":50},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":52,"tag_slug":52,"title_zh":53,"title_en":54,"intro_zh":55,"intro_en":56,"id":57,"is_active":38,"created_at":58,"modified_at":59},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":61},[62,67,72,77,82,87],{"id":63,"title":64,"news_slug":65,"published_at":66},"8303f420-b3f1-485a-aa6f-7775256c84a7","Gemini 3.8 Audio 双发:Live 和 Extended Thinking 把思考+说话压到近实时","gemini-3-8-audio-live-extended-thinking","2026-09-16T03:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"1ba7e499-d93d-4566-bbeb-0b762904c0ab","Google 把 Lyria 3.5 装进 Gemini 与公开 API:音乐生成从独立工具变成默认选项","lyria-3-5-gemini-app-api","2026-09-06T23:06:32+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"e0a484e4-41e0-4f91-9b9b-a196bbdcf3ba","Gemini 3.8 Flash 双发:同价升级 + Cyber 走可信项目 Fairwind","gemini-3-8-flash-cyber-fairwind-launch","2026-09-03T03:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"9af3dd83-6ed9-498d-9da0-547d917f3e19","语音转文字有了专用模型:Gemini 3.5 Transcribe 上线,出稿快 70%","gemini-35-transcribe-dedicated-asr","2026-08-31T13:30:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"16856034-439d-4915-aed4-80b42ae09c68","Gemini 3.7 Flash：FrontierCode 43.6%，价格腰斩","gemini-3-7-flash-coding-agent-fast","2026-08-13T09:00:00+00:00",{"id":88,"title":89,"news_slug":90,"published_at":91},"34edaffc-6b5c-4df1-9e2f-d864cada6063","Gemini 走进 K-12 课堂：Google 把「上下文」塞进每个作业","gemini-classroom-k12-contextualized-prompts","2026-08-07T02:00:00+00:00"]