[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-meta-muse-voice-transcribe-streaming-asr":3,"topics-all":38,"news-related-6062d551-9068-4a9f-ae8e-4e99269cd838":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"6062d551-9068-4a9f-ae8e-4e99269cd838","Muse Voice Transcribe 发布:流式转写、20+ 说话人分离、端点检测,Meta 全塞进一个模型","Meta 超级智能实验室发布首个实时音频感知模型 Muse Voice Transcribe,把流式 ASR、20+ 说话人分离和端点检测合并进单个自回归模型,官方图表演示流式词错率 3.1%、说话人分离错误率 17.5%,目前已上线 Artificial Analysis 流式语音转写榜第一。","做语音产品的团队都习惯三件套架构:一个模型转写、一个模型分离说话人、再加一个检测器判断用户何时说完。每次交接都引入延迟和新的出错点。Meta 超级智能实验室 9 月 1 日发布 Muse Voice Transcribe,把这三件事压进一个自回归模型——官方称其为首个实时音频感知模型。\n\n## 一个模型怎么同时干三件事\n\n技术路线上,Muse Voice Transcribe 出身 Muse Spark 家族,是自回归多模态模型。音频按 80 毫秒(12.5Hz)分块,每块转成一个 soft token;模型在每个音频块上自己决定是继续听(`\u003C|next_audio|>`),还是吐出文本 token。说话人分离和端点检测也不是外挂模块,而是靠额外特殊 token(`\u003C|start_of_turn|>`、`\u003C|speaker_A|>`、`\u003C|speech_endpoint|>` 等)在同一条序列里完成,官方强调三个任务联合训练,无需后处理。\n\n最有意思的设计是「自适应延迟」:等待越久越准、但延迟越高,Meta 用强化学习把词错率奖励和延迟奖励相乘,让模型逐词决定等多久再落笔。\n\n## 数字层面\n\n按官方图表,流式最终转写词错率 3.1%,同图七个系统在 3.4%–4.0% 之间;说话人分离平均错误率(AMI-IHM、AMI-SDM、VoxConverse)17.5%,其余五个系统 21.1%–28.6%。速度-精度散点图上,该模型约 0.16 秒达到约 3.0% 错误率,低于 Soniox、Cartesia、ElevenLabs 系统构成的前一代 Pareto 前沿。Meta 同时声明该模型在 Artificial Analysis 流式语音转写和公开说话人分离基准上排名第一(截至 2026 年 9 月 1 日)——榜单归属是官方口径,数字可回溯至上文图表。\n\n第三方媒体 MarkTechPost 的独立报道给出更多对照:同榜 Cartesia Ink-2(语义端点)3.4% WER \u002F 0.43 秒,ElevenLabs Scribe v2 Realtime 3.6% \u002F 0.14 秒;定价折合每小时音频 0.18 美元,低于 Cartesia 的 0.24 美元和 ElevenLabs 的 0.39 美元,且目前仅通过 API 提供、未开放权重(此三项为 MarkTechPost 单源转述)。\n\n## 语言能力与落地\n\n模型用 70+ 种语言训练,其中 25 种经过充分验证,官方推荐首发使用这 25 种;原生支持句内句间任意 code-switching,演示里一段中英混杂的极客闲聊——本地部署 Muse Glimmer、4bit 量化、350 瓦 TDP——转得明明白白。长音频原生支持超过 1 小时、20+ 说话人,官方博客放了一段 1 小时 52 秒、11 人闲聊的完整转录演示。可用渠道是 Meta Model API、Mac 版 Meta AI 和 Muse Code,Mac 端按 Fn 即全局听写。\n\n## 值得留意的两点\n\n一是只放 API 不放权重,和 Meta 此前开源 Muse Glimmer 30B、Muse Spark 系列的姿态形成反差——「个人超级智能」的叙事先从闭源耳朵开始。二是把 ASR、分离、端点检测统一进 token 空间,和文本模型把工具调用、思考过程 token 化同向:管线越短,实时性越好,失败模式越少。做语音 Agent 的团队值得逐段读一遍官方博客(原文:https:\u002F\u002Fresearch.meta.ai\u002Fblog\u002Fintroducing-muse-voice-transcribe)。\n\n当转写、分离、断句都变成同一个模型的几个特殊 token,语音交互的下一块短板,可能就不再是「听不听得清」了。","https:\u002F\u002Fresearch.meta.ai\u002Fblog\u002Fintroducing-muse-voice-transcribe","245423b3-0e3c-47a8-9370-34e7a3b3988e",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"e676a5cf-1f24-472f-a765-86fa21a1bc3c","ai-model",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"350e5a2d-925e-4411-ad7b-7db0d1160400","en","Meta's Muse Voice Transcribe: ASR + Diarization + Endpointing","Meta's first real-time audio model merges streaming ASR, diarization and endpointing in one: 3.1% streaming WER, No.1 on Artificial Analysis.","Production voice stacks usually run three stitched systems: one model transcribes, a second separates speakers, a detector decides when the user stopped talking. Every hand-off adds latency and a new failure mode. On September 1, Meta Superintelligence Labs released Muse Voice Transcribe, collapsing those three jobs into one autoregressive model — what the lab calls its first real-time audio perception model.\n\n## One Model, Three Jobs\n\nTechnically, Muse Voice Transcribe comes from the Muse Spark family. Audio is processed in 80ms chunks (12.5 Hz), each transformed into a single soft token; at each chunk the model decides whether to keep listening (`\u003C|next_audio|>`) or emit a text token. Diarization and endpointing are not bolted-on modules either: extra special tokens (`\u003C|start_of_turn|>`, `\u003C|speaker_A|>`, `\u003C|speech_endpoint|>`) carry them inside the same sequence. Meta says all three tasks are trained jointly, with no required post-processing.\n\nThe most interesting design choice is \"adaptive delay\": waiting longer before transcribing improves accuracy but raises latency, so Meta used reinforcement learning to combine a word-error-rate reward and a delay reward multiplicatively, letting the model decide per word how much audio context to wait for.\n\n## The Numbers\n\nPer Meta's published charts, final-transcription streaming word error rate is 3.1%, with seven other systems in the same chart ranging from 3.4% to 4.0%; average diarization error rate across AMI-IHM, AMI-SDM and VoxConverse is 17.5%, versus 21.1%–28.6% for five other systems. On the speed-accuracy scatter, the model reaches about 3.0% error at roughly 0.16 seconds, below the previous Pareto frontier formed by Soniox, Cartesia and ElevenLabs systems. Meta also states the model ranks first on Artificial Analysis streaming speech-to-text and on public diarization benchmarks as of September 1, 2026 — a vendor claim, with the numbers traceable to the charts above.\n\nIndependent coverage from MarkTechPost adds context: on the same board, Cartesia Ink-2 (semantic endpoints) sits at 3.4% WER \u002F 0.43s and ElevenLabs Scribe v2 Realtime at 3.6% \u002F 0.14s; pricing works out to $0.18 per audio hour, below Cartesia's $0.24 and ElevenLabs' $0.39, with an API-only release and no open weights so far (these three items are single-sourced to MarkTechPost).\n\n## Languages and Availability\n\nThe model is trained on 70+ languages, 25 of them extensively verified and recommended for the initial release. It natively supports arbitrary code-switching within and between sentences — one demo is a long Mandarin-English monologue about running Muse Glimmer locally on an RTX 3090 with 4-bit quantization that transcribes cleanly. Long audio over one hour with 20+ speakers is natively supported; the official blog posts a full transcript of a 1-hour-52-second, 11-person real conversation. Availability is via Meta Model API, the Meta AI for Mac client, and Muse Code, where holding the Fn key gives system-wide voice dictation.\n\n## Two Things Worth Noting\n\nFirst, this is an API-only release with no open weights — a contrast with Meta's earlier generosity in open-sourcing Muse Glimmer 30B and the Muse Spark line. The \"personal superintelligence\" story starts with closed ears. Second, folding ASR, diarization and endpointing into one token space mirrors what text models did with tool calls and reasoning traces: shorter pipelines, better real-time behavior, fewer failure modes. Teams building voice agents should read the official blog end to end (source: https:\u002F\u002Fresearch.meta.ai\u002Fblog\u002Fintroducing-muse-voice-transcribe).\n\nOnce transcription, speaker separation and turn-taking are just special tokens inside one model, the next bottleneck for voice interaction may no longer be whether the machine can hear clearly.","meta-muse-voice-transcribe-streaming-asr","2026-09-05T13:11:00Z","2026-09-05T13:11:16.764597Z","2026-09-05T13:11:16.764606Z",true,"agent",144,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"fdf05035-e02e-4954-8beb-c697ecc7975a","Cohere Parse 发布:$1.5 每千页的文档解析模型,ParseBench 79.2 超 Mistral OCR 4","cohere-parse-document-parsing-model","2026-08-28T13:30:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"12a5f49d-8c83-4c40-82af-1c0b7f1c8b3e","DeepSeek 给 V4-Flash 装上眼睛:Vision-Exp 实验模型两项基准反超 Opus 4.8","deepseek-v4-flash-vision-exp-multimodal","2026-08-21T23:05:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"40210e0d-84e3-460b-bde2-295b77573ab8","Qwen3.7-Text-Embedding 上线:20% 检索增益、256-2560 可变维度,阿里把 RAG 的地基悄悄重浇了一遍","qwen3-7-text-embedding-launch","2026-08-14T13:10:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"7258978b-dfcd-4cb4-91c4-3b8569cd5deb","Qwen-Audio-3.0-TTS双版本发布:Plus登顶Artificial Analysis,Flash压到300ms首包延时","qwen-audio-3-tts","2026-07-20T10:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"c75518c3-da86-45a3-80fd-004e4650c06b","阿里把「世界模型」搬上百炼:HappyOyster 1.0 用一句话生成可实时交互的开放世界","alibaba-happyoyster-1-0","2026-07-19T22:00:00+00:00"]