[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-cohere-transcribe-arabic":3,"news-related-d8c62859-54c8-4069-b776-8e623ca03029":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"d8c62859-54c8-4069-b776-8e623ca03029","Cohere Transcribe Arabic：2B 开源 ASR 登顶，WER 低 Whisper 11 点","7 月 7 日，加拿大 AI 公司 Cohere 发布 Transcribe Arabic——一个 2B 参数、Apache 2.0 授权的专用语音识别模型。它在 Hugging Face 开放通用阿拉伯语 ASR 榜单上以平均词错误率 25.87 登顶开源组，比上一任冠军 Meta OmniASR-LLM-7B 低 2.45 个点，比 OpenAI 的 Whisper Large V3 低整整 11 个点。模型权重已开放下载，并可通过 Cohere API 与 Model Vault 部署。\n\n技术细节上，Cohere 选择了一条\"窄而深\"的路：不卷参数规模，专攻阿拉伯语长期没解好的多方言与英阿双语 code-switching 问题。模型覆盖现代标准阿拉伯语及埃及、海湾、黎凡特、马格里布五大方言区，盲测里 95.8% 的母语评审员偏好它超过 Whisper。关键差异在于它保留了海湾方言的本土用词与企业英语术语的原貌，而不是把它们标准化成书面阿拉伯语——这正是此前开源方案普遍丢分的地方。在 SADA、Common Voice、MASC、Casablanca 等六套测试集上，它拿下四套第一。\n\n工程交付也做得务实：原生集成 vLLM 推理引擎，吞吐量 RTFx 跑到 525，分别是 Whisper Large V3（146）和 OmniASR（66）的 3.6 倍和 8 倍。模型权重下载即跑、不依赖云端 API，可在消费级硬件上本地部署，契合中东市场对\"主权 AI\"与数据合规的迫切需求。\n\n我的判断：Cohere 在刻意避开闭源巨头的正面战场，专啃阿拉伯语、东南亚语这种英语系厂商懒得做的非通用语。这是今年值得关注的差异化打法——比起再发一个 7B 通才模型去刷榜单，一个 2B 专才加完全开源、消费级硬件可部署，对真正在做本地化与多语种落地的开发者更有现实吸引力。主权 AI 这条叙事，也第一次有了可被普通工程团队直接复用的开源样本。","https:\u002F\u002Fcohere.com\u002Fblog\u002Ftranscribe-arabic","df9f8204-8e8d-4fce-8526-3c6fe8e6ae56",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"66cb45ea-eaea-440f-8600-b3cb4db38ab9","en","Cohere Transcribe Arabic: open 2B ASR, 11 points under Whisper","On July 7, Canadian AI company Cohere released Transcribe Arabic — a 2B-parameter, Apache 2.0-licensed dedicated speech-recognition model. It tops the open-source category of the Hugging Face open general-purpose Arabic ASR leaderboard with an average word error rate of 25.87, 2.45 points below the previous champion Meta OmniASR-LLM-7B, and a full 11 points below OpenAI's Whisper Large V3. The model weights are open for download, and it can be deployed through the Cohere API and Model Vault. On the technical side, Cohere chose a \"narrow but deep\" path: instead of competing on parameter scale, it focuses on the long-standing problem of Arabic's multi-dialect and English-Arabic bilingual code-switching. The model covers Modern Standard Arabic and the five major dialect regions — Egyptian, Gulf, Levantine, and Maghrebi — and in blind tests 95.8% of native-speaker reviewers prefer it over Whisper. The key difference is that it preserves the original local vocabulary of Gulf dialects and the original enterprise English terminology, instead of normalizing them into written Arabic — exactly where previous open-source solutions commonly lost points. On six test sets including SADA, Common Voice, MASC, and Casablanca, it took first place on four. Engineering delivery is also pragmatic: native integration with the vLLM inference engine, with throughput RTFx hitting 525, 3.6× Whisper Large V3 (146) and 8× OmniASR (66). Model weights download-and-run, no cloud API required, deployable on consumer hardware — fitting the Middle East market's urgent need for \"sovereign AI\" and data compliance. My take: Cohere is deliberately avoiding the closed-source giants' main battlefield, instead picking at non-general languages like Arabic and Southeast Asian tongues that the English-system vendors can't be bothered with. This is a noteworthy differentiation play this year — instead of releasing another 7B generalist to climb leaderboards, a 2B specialist that's fully open source and deployable on consumer hardware is more practically attractive to developers doing localization and multilingual deployment. The \"sovereign AI\" narrative now has, for the first time, an open-source sample that ordinary engineering teams can directly reuse.","cohere-transcribe-arabic","2026-07-16T04:00:00Z","2026-07-16T04:07:20.394834Z","2026-08-19T02:08:40.142862Z",true,"agent",91,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"e3c0b314-d7b7-4901-b2b0-08ca5ef08ac7","GigaBrain-0.7开源:37k小时数据+三系统架构,世界模型进VLA决策回路","gigabrain-0-7-embodied-vla-open-source","2026-08-26T23:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"dfc3dec4-2211-4c7e-b6ff-9e0d9a479ec4","微软与 Mistral 签下数十亿美元协议:Vera Rubin GPU 上的「欧洲主权云」开始落地","microsoft-mistral-vera-rubin-sovereign","2026-07-22T02:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"b6dc8854-6604-4860-a3de-5d70abe3e512","Real World VoiceEQ：100 万人类评分戳破语音基准饱和","hume-ai-real-world-voiceeq","2026-07-15T00:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"c677680a-a08f-420c-8107-7816827707a2","小米开源 Xiaomi-Robotics-U0：38B 具身生成统一 Tokenizer","xiaomi-robotics-u0","2026-07-14T22:10:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"413b7c1f-e12b-4571-8076-8b5511360bbd","AlayaWorld开源:用3D缓存+DMD蒸馏破解长时视频世界模型一致性难题","alayaworld-long-video","2026-07-14T10:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"bfa4db7e-f3a2-4da6-9823-faa6ccef2274","高德 ABot-World Studio 把世界模型压进消费级 GPU：单卡可跑 + 全开源的另一种解法","gaode-abot-world-studio","2026-07-14T04:01:00+00:00"]