[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-phonellm-alpha-1-voice-agent-open-model":3,"news-related-61de017b-bdd6-44b3-9f45-d4fb233bd24d":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"61de017b-bdd6-44b3-9f45-d4fb233bd24d","PhoneLLM 开源:30B MoE 电话客服模型,自称比 GPT-5.6 Terra 便宜 94%","Daily(Pipecat)开源 PhoneLLM Alpha 1:基于 Nemotron 3 Nano 30B-A3B 微调、3.5B 激活的电话语音智能体模型。厂商口径:电话任务与 GPT-5.6 Terra 相当、成本降 94%、P95 首 token 快 1.3 秒;自托管约 0.00025 美元\u002F分钟。","打电话给客服,人类的耐心大约是 1.5 秒——Daily 团队(Pipecat 开源框架背后的公司)在发布 PhoneLLM Alpha 1 时给了个硬数字:voice-to-voice 延迟要压在 1,500ms 左右,通话体验才自然。而他们测得 GPT 5.6 Terra 快速模式的 P95 首 token 延迟约 1,900ms——也就是说,光 LLM 这一环就把预算撑爆了,还没算 STT、TTS 和网络。这就是 PhoneLLM 存在的理由:一个专为电话场景训练的开源权重模型。\n\n## 30B MoE,3.5B 激活,只为接电话\n\nPhoneLLM Alpha 1 的底细写在模型卡里:基于 NVIDIA Nemotron 3 Nano 30B-A3B 做全参数微调,训练用 NVIDIA NeMo 框架;架构是 Hybrid Mamba-Transformer MoE,总参数 30B,激活只有 3.5B;上下文 262,144 tokens;BF16 safetensors;BSD 2-Clause 许可证,无商用限制。推荐的推理设置很干脆:temperature=0、thinking 关闭——因为模型就是按这个口径训的。\n\n场景也收得很窄:金融、医疗、零售、酒店的 inbound 客服 + 常规 outbound 外呼。它要解决的是一个非常具体的痛点:关掉 thinking 之后,大小模型在多轮长对话里的工具调用都不可靠——模型嘴上说「好的,桌已订好」,实际上根本没调用工具。PhoneLLM 的训练目标就是在不开思考链的前提下,该调工具时调对工具。\n\n## 每分钟 0.00025 美元是怎么算出来的\n\n模型卡给了一笔很工程化的账:单张 B200 塞 44 个并发 agent 进程(双卡节点 88 个);Modal 上 B200 基础价 $6.2496\u002F小时,区域锁定乘 1.5 得 $9.3744,按 70% 利用率折算 $13.392\u002F小时,即 $0.2232\u002F分钟;除以 88 个并发,得到每个 agent 每分钟 $0.00025。单请求 TTFT P95 在 B200 上低于 100ms;配 Modal AutoEndpoints 的定制配置后,在 sub-600ms P95 首 audio token 目标下,最大并发大约是 vLLM 通用 cookbook 配置的两倍。\n\nDaily 的官方口径更大胆:电话任务表现与 GPT 5.6 Terra 相当,成本便宜 94%,P95 首 token 快 1,300ms。联合创始人 Kwindla Hultman Kramer 在 X 上的说法是 1\u002F3 延迟、1\u002F18 成本——两处数字在数学上对得上(94% ≈ 1\u002F18)。\n\n## 自建考场的问题\n\n这些数字全部来自 Daily 自己的 PhoneBench v1:LLM 裁判对人类标注做校准,评的是电话语体、工具调用准确性、言行一致(say\u002Fdo consistency)、事实接地、对话连贯、鉴权与升级纪律、呼叫结果。独立媒体 explainx.ai 的报道把该泼的冷水泼足了:Alpha 1 是开发方自己贴的标签;benchmark 是自建的,没有第三方公开榜单复现;数字对配置敏感(temperature=0 + 关 thinking 才复现得出来);而且是「自托管专用部署 vs 通用托管 API」的不对等比较——Terra 的延迟和成本是 OpenAI 端点零基建直接给的,PhoneLLM 的数字前提是你自己运维 B200。GPT-5.6 Terra 在 7 月底降价后是 $2\u002F$12 每百万 token,这个对比框架下每分钟成本怎么折算,两家口径并不在同一张账本上。\n\n## 所以呢\n\n即便把厂商口径打对折,这个发布仍然值得注意:它不是「又一个小模型」,而是把「任务专用微调 + MoE 低激活 + 自托管并发工程」三件事拧在一根延迟轴上。Daily 的判断是行业正在转向小开源专用模型——用生产 agent 轨迹和私有数据按月迭代权重。电话客服恰好是延迟、成本、工具可靠性三重约束最苛刻的场景之一,PhoneLLM 能不能站住,等第三方复测;但「用 1\u002F18 的成本干一个具体工种」这条路线,已经有人交卷了。\n\n原文与权重:https:\u002F\u002Fhuggingface.co\u002Fpipecat-ai\u002Fphonellm-alpha-1","https:\u002F\u002Fhuggingface.co\u002Fpipecat-ai\u002Fphonellm-alpha-1","3e8d298f-64af-4ae9-bae7-05de4a653ecb",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"d11f0044-8aef-487c-bebe-89ce4683a4a3","moe",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"78e7a693-6c53-4b03-a846-e6c63140418d","en","PhoneLLM Alpha 1: An Open 30B MoE Built to Answer the Phone","Daily, the team behind Pipecat, open-sources PhoneLLM Alpha 1: a full fine-tune of NVIDIA's Nemotron 3 Nano 30B-A3B with 3.5B active parameters, purpose-built for phone voice agents. Vendor-reported: GPT-5.6 Terra parity on phone tasks at 94% lower cost, 1,300ms faster P95 TTFT, $0.00025 per self-hosted agent-minute, BSD licensed.","Call a customer-service line, and human patience lasts about 1.5 seconds. That is the hard number Daily — the company behind the open-source Pipecat voice-agent framework — anchored its PhoneLLM Alpha 1 release on: voice-to-voice latency needs to stay around 1,500ms for a phone conversation to feel natural. Their own measurement put GPT 5.6 Terra's P95 time-to-first-token in fast mode at about 1,900ms, meaning the LLM alone blows the budget before STT, TTS, and network overhead even enter the picture. That is the reason PhoneLLM exists: an open-weights model trained specifically for the phone.\n\n## A 30B MoE with 3.5B active parameters, built to take calls\n\nThe model card lays out the specifics: PhoneLLM Alpha 1 is a full-parameter fine-tune of NVIDIA's Nemotron 3 Nano 30B-A3B, trained with the NVIDIA NeMo framework. The architecture is a hybrid Mamba-Transformer mixture-of-experts — 30B total parameters, only 3.5B active — with a 262,144-token context window, BF16 safetensors, and a BSD 2-Clause license with no commercial restrictions. The recommended inference settings are blunt: temperature=0, thinking disabled, because that is exactly the regime the model was trained for.\n\nThe target scenarios are deliberately narrow: inbound customer service plus routine outbound calling for financial services, healthcare, retail, and hospitality. The problem it attacks is very specific: with thinking disabled, models large and small are unreliable at tool invocation across long, multi-turn conversations — an LLM will cheerfully say \"Yes, I've booked that table for you\" without ever calling the booking tool. PhoneLLM is trained to invoke the right tools at the right time, no chain-of-thought required.\n\n## Where $0.00025 per minute comes from\n\nThe model card walks through a very engineering-grade cost calculation. A single B200 hosts 44 concurrent agent processes (88 per two-GPU node). On Modal, a B200 costs $6.2496\u002Fhour at base; region pinning adds a 1.5x multiplier for $9.3744\u002Fhour; targeting 70% utilization yields an effective $13.392\u002Fhour, or $0.2232\u002Fminute. Divide by 88 concurrent agents and you get $0.00025 per agent-minute. Single-request TTFT P95 lands below 100ms on a B200, and with Modal AutoEndpoints' tuned configuration, maximum concurrency at the sub-600ms P95 time-to-first-audio-token target is roughly double what the generic vLLM cookbook configuration delivers.\n\nDaily's official framing is bolder: performance on par with GPT 5.6 Terra on voice-agent tasks, at 94% lower cost and with a 1,300ms faster P95 time-to-first-token. Co-founder Kwindla Hultman Kramer put it on X as one-third the latency and one-eighteenth the cost — and the two framings are mathematically consistent (94% off is roughly 1\u002F18th the price).\n\n## The self-authored benchmark problem\n\nAll of those numbers come from Daily's own PhoneBench v1: LLM judges calibrated against human labels, grading telephone speaking style, tool-call accuracy, say\u002Fdo consistency, factual grounding, conversation coherence, authentication and escalation discipline, and caller outcome. The independent outlet explainx.ai poured the appropriate cold water: \"Alpha 1\" is a label the developers gave it themselves; the benchmark is self-authored, with no third-party public leaderboard reproduction; the numbers are configuration-sensitive (temperature=0 plus thinking off, or they likely won't replicate); and the comparison is a mismatched one — a specialized self-hosted deployment versus a general-purpose hosted API. Terra's latency and cost are what OpenAI's endpoint hands you with zero infrastructure work; PhoneLLM's numbers assume you operate the B200 yourself. After OpenAI's late-July price cut, GPT-5.6 Terra lists at $2\u002F$12 per million tokens, and the two framings do not sit on the same ledger when it comes to converting that into per-minute cost.\n\n## So what\n\nEven with the vendor's numbers cut in half, the release is worth attention. This is not \"another small model\" — it welds task-specific fine-tuning, low-activation MoE, and self-hosted concurrency engineering onto a single latency axis. Daily's read is that the industry is shifting toward small, open, purpose-built models, with weights iterated monthly using production agent traces and proprietary data. Phone customer service happens to be one of the harshest scenes for the triple constraint of latency, cost, and tool reliability — whether PhoneLLM holds up awaits third-party testing. But the route of \"doing one specific job at 1\u002F18th the cost\" already has its first graded submission.\n\nModel card and weights: https:\u002F\u002Fhuggingface.co\u002Fpipecat-ai\u002Fphonellm-alpha-1","phonellm-alpha-1-voice-agent-open-model","2026-08-29T21:10:00Z","2026-08-29T21:10:48.166659Z","2026-08-29T21:10:48.166674Z",true,"agent",31,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"33f3b08b-c8a2-43ec-81cf-85e2b918f913","腾讯开源 Hy4 preview:770B MoE、1M 上下文,模型首次参与自身训练","tencent-hy4-preview-770b-moe","2026-08-29T15:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"3d36921f-3b84-4663-97a0-fee7d4eff795","汤森路透开源 Thomson-1.0-Small:持续学习改造 Qwen,3B 激活的 35B MoE","thomson-1-0-small-continual-learning","2026-08-28T19:10:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"c4027891-42ca-4517-817a-83a48550b1bb","Qwen3.8-Flash-Next 开源:6B 激活参数跑赢 Opus,训练成本仅 1\u002F9","qwen-flash-next-gdn-qsa-architecture","2026-08-27T15:10:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"68072ee1-fc37-4064-ab18-09550ae72d1b","GLM-5.3-Flash 把 320B MoE 跑在国产芯片上:Flash 价位和 $0.15 API 的混合注意力栈","glm-5-3-flash-chinese-chips-hybrid-attention","2026-08-27T03:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"7958a2f1-028c-4b4e-b134-0d5de9afc1c1","Motif 3 收官:韩国 314B MoE 改用 MIT 许可,从零起步架构首次面向商用","motif-3-mit-license-sovereign-ai","2026-08-24T00:00:00+00:00"]