[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gpt-realtime-2-128k-context-parallel-tool":3,"topics-all":36,"news-related-89fc3a4f-8176-48a8-95cd-e9728573a436":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"89fc3a4f-8176-48a8-95cd-e9728573a436","OpenAI 推出 GPT‑Realtime‑2：语音交互从「命令执行」迈向「真正对话」","5月7日，OpenAI 在 API 中上线三款音频模型：GPT‑Realtime‑2（集成 GPT‑5 级推理的语音模型）、GPT‑Realtime‑Translate（覆盖 70+ 输入语言、13 种输出语言的实时翻译）以及 GPT‑Realtime‑Whisper（流式语音转文字）。\n\n**这次不同在哪里？**\n\n之前大多数语音 AI 本质上是「语音化的命令执行器」——听清一句话、执行单一指令、结束。GPT‑Realtime‑2 的核心升级在于将大模型推理直接嵌入语音交互链路。几个值得注意的技术细节：\n\n- **上下文窗口从 32K 扩展至 128K**：足以支撑多轮复杂任务，例如连贯的旅行规划会话。\n- **并行工具调用 + 过程透明化**：模型可同时执行多个工具，并用「正在查询您的日历」等语音反馈告知用户状态，而不是干等最终答案。\n- **更强容错与恢复能力**：工具调用失败时，模型会生成自然的补救话术，而非沉默或崩溃。\n\n**实时翻译的落地价值**\n\nGPT‑Realtime‑Translate 将翻译从「说完一段再翻」推进到「边说边翻」。Deutsche Telekom 已宣布将其用于多语言客户支持，Priceline 计划用其帮助旅客完成全程语音行程管理。这对跨语言客服、医疗咨询等场景有直接价值。\n\n**评论：语音正在成为真正的 UI**\n\n过去语音助手稍复杂的任务就露馅，GPT‑Realtime‑2 代表了一次质变——将强推理模型直接暴露在用户面前，而非藏在文字输入框后面。对企业而言，下一步的挑战更多是响应延迟和 SLA 保证，而非模型能力本身。2026 年，或许是企业市场真正检验这条路线可行性的元年。","https:\u002F\u002Fopenai.com\u002Findex\u002Fadvancing-voice-intelligence-with-new-models-in-the-api\u002F","15975962-b5fe-49e5-ae68-687ba6cb7015",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"baf131c1-687a-49f4-87f6-4dd87c1c692f","gpt",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"9fc3ae61-6ebb-41e9-a4ed-ade9e5eb501d","en","GPT-Realtime-2: voice moves from commands to conversation","On May 7, OpenAI launched three audio models in its API: GPT-Realtime-2 (a voice model integrated with GPT-5-class reasoning), GPT-Realtime-Translate (covering 70+ input languages and 13 output languages for real-time translation), and GPT-Realtime-Whisper (streaming speech-to-text).\n\n**What's different this time?**\n\nMost previous voice AIs were essentially \"voice-ified command executors\" — hear one sentence, execute one command, end. GPT-Realtime-2's core upgrade is embedding large-model reasoning directly into the voice interaction chain. Several technical details worth attention:\n\n- **Context window expanded from 32K to 128K**: sufficient to support multi-turn complex tasks, such as coherent travel-planning conversations.\n- **Parallel tool calling + process transparency**: the model can execute multiple tools simultaneously, and inform users of status with voice feedback like \"looking up your calendar\" rather than waiting silently for the final answer.\n- **Stronger error tolerance and recovery**: when a tool call fails, the model generates natural remediation language rather than silence or crash.\n\n**Real-world value of real-time translation**\n\nGPT-Realtime-Translate advances translation from \"wait until finished, then translate\" to \"translate as you speak.\" Deutsche Telekom has announced using it for multilingual customer support, while Priceline plans to use it to help travelers with full voice itinerary management. This has direct value for cross-language customer service, medical consultation, and similar scenarios.\n\n**Commentary: Voice is becoming a real UI**\n\nPast voice assistants were exposed on tasks of even modest complexity. GPT-Realtime-2 represents a qualitative change — exposing strong reasoning models directly to users, not hiding them behind text input boxes. For enterprises, the next challenges are more about response latency and SLA guarantees than model capability itself. 2026 may be the first year enterprise markets truly test the viability of this approach.","gpt-realtime-2-128k-context-parallel-tool","2026-05-08T01:10:00Z","2026-05-08T01:07:31.536966Z","2026-08-19T02:08:40.142862Z",true,"agent",135,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"0f466258-290e-4c6b-b471-2169ba6a393f","GPT-Live-1 进 API:全双工语音层 0.05 美元一分钟,推理外包给 GPT-6 Astra","gpt-live-1-api-launch","2026-09-12T17:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"e12d2e7d-35b7-42d2-b02f-bdcc0a547878","机械臂实测 GPT-6 Astra:19\u002F20 对 8\u002F20 完胜 Fable 5.1,精细插入却全员卡壳","gpt-6-astra-robot-arm-benchmark","2026-09-07T19:13:54+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"390c2437-4e4f-45ec-8270-67c5bfa4fa47","ChatGPT、Claude、Grok、Gemini 罕见同时下线,周四早晨全球 AI 集体失声","chatgpt-claude-grok-gemini-thursday-outage","2026-09-05T06:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"b5be4ce8-4a41-461c-9202-148e64fab329","GPT-6 Astra 系统卡:零日自用、对齐升 53%,CoT 可监控性反向下滑","gpt-6-astra-system-card-2026-monitorability","2026-09-04T03:30:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"5bf8fa2d-258e-41d3-bfb6-5c2053433cfd","GPT-6 Astra 正式上线:8 月因安全被暂停的旗舰回来了","gpt-6-astra-launch","2026-09-04T03:12:38+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"9e58d587-3c1b-44c5-ad36-daf23aeb42a2","微软叫停 tokenmaxxing:GitHub Copilot 默认切回 GPT-5.6 Sol,Parikh 设 token 预算","microsoft-token-budget-gpt-5-6-default","2026-09-03T00:30:00+00:00"]