[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-cohere-parse-document-parsing-model":3,"topics-all":38,"news-related-fdf05035-e02e-4954-8beb-c697ecc7975a":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"fdf05035-e02e-4954-8beb-c697ecc7975a","Cohere Parse 发布:$1.5 每千页的文档解析模型,ParseBench 79.2 超 Mistral OCR 4","Cohere 发布文档解析模型 Parse:PDF 进、Markdown 出，表格结构保留。官方 ParseBench 评测 79.2 分超 Mistral OCR 4,定价 1.5 美元每千页,8 卡 H100 节点每分钟 2160 页。","企业 RAG 的第一公里，长期卡在文档解析上。合同、保单、发票、论文——这些真正躺着企业知识的地方，往往是最难喂给模型的：表格错位、阅读顺序断裂、删除线和斜体这类改变语义的样式直接丢失。ParseBench 的评测数据把这件事量化得很直白：AWS Textract 在 Semantic Formatting 维度只有 2.8 分，Google Document AI 也只有 33.0 分——传统文档智能服务在「格式即含义」的场景里几乎是盲的。\n\nCohere 本周交出的答案是 Parse：一个面向企业文档的视觉语言模型。它的输入输出很干脆——复杂多模态文件进，干净的 Markdown 出，表格和嵌入图像以结构保留，支持九种主要商业语言，并为视觉元素返回 bounding box。用 Cohere 自己的话说，这是「市面上最强性价比」的文档解析方案。\n\n## 官方评测：79.2 分卡在专用与通用之间\n\n在 Cohere 提交的 ParseBench 评测中（三个维度平均），Parse 拿到 79.2，高于 Mistral OCR 4 的 74.5、Databricks AI Parse 的 72.4 和 LlamaParse Cost Effective 的 78.3；对比超大规模云厂商的方案优势更大——比 AWS Textract 和 Google Document AI 高出 20 分以上。分维度看，Tables 87.0、Content Faithfulness 86.6 是强项，Semantic Formatting 64.0 则明显落后于 frontier LLM。\n\n官方也把边界说得很清楚：在他们的评测集里，只有 GPT-5.5（84.4）、Opus 4.8（84.3）和 Gemini 3.5 Flash（81.8）三个通用大模型压过 Parse——而这三者都是「显著更大」的通用模型。换句话说，这个专用小模型（第三方追踪站 AI\u002FTLDR 转述其参数量为 2.3B）用几分之一的成本，打到了 frontier 门槛前。\n\n值得肯定的还有评测口径的透明度：Cohere 在脚注里说明使用了 2026 年 8 月修正后的评分规则——修复了此前会虚高 Semantic Formatting 的粗体\u002F标题检测 bug，并对所有竞品用同一规则重跑；同时明确排除了 Layout 和 Chart 两个维度，理由是产品定位使然、而非能力缺陷，图表数据抽取计划留给下一个版本。\n\n## 单价与吞吐：把解析做成「水电煤」\n\n$1.50 每 1000 页的 API 定价，配合生产级吞吐——在 8 卡 H100 节点上每秒 36 页、每分钟 2160 页，约为 dots.mocr 的 1.4 倍、Chandra OCR 2 的 2.2 倍（官方基准，所有模型均以 vLLM 部署测量）。对于合规行业，Parse 可以跑在私有云或本地环境，也可走单租户的 Model Vault：50% GPU 利用率时成本再降 23%，满载时降 61%。\n\nCohere 算了一笔账：一个每月处理约 1300 万页的应付账款流程，用 Model Vault 替代 API 每月省约 1.2 万美元、每年约 14.4 万美元；对比定价 $10\u002F千页的云厂商文档服务，单一流程每年可省约 147 万美元。Parse 现已通过 Cohere API、Model Vault、Microsoft Foundry 和 AWS SageMaker 四个渠道可用，并作为 Compass 检索栈的一环与 Embed、Rerank 打包成「文档到答案」管线。\n\n## 所以呢\n\n这次发布真正值得看的不是某个跑分，而是一个趋势的又一次确认：文档解析正在从「通用 LLM 顺手做一下」分裂出一个独立的专用模型层——小参数、极致单价、高吞吐，直接切走 hyperscaler 存量 Document AI 的市场。对企业的理性策略也随之清晰：精度敏感的少量文档交给 frontier LLM，百万页级的规模化管线交给 $1.5\u002F千页的专用模型。RAG 第一公里的成本，终于开始按「每千页」计价了。\n\n（原始来源：Cohere 官方博客 https:\u002F\u002Fcohere.com\u002Fblog\u002Fparse）","https:\u002F\u002Fcohere.com\u002Fblog\u002Fparse","df9f8204-8e8d-4fce-8526-3c6fe8e6ae56",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"5837d952-0e2a-43a8-bd77-ef2261a76544","en","Cohere Parse: a $1.50-per-thousand-page document parsing model scoring 79.2 on ParseBench","Cohere launches Parse, a vision-language model that turns PDFs into clean Markdown with tables preserved. It scores 79.2 on the official ParseBench eval, beating Mistral OCR 4, at $1.50 per 1,000 pages and 2,160 pages per minute on an 8x H100 node.","The first mile of enterprise RAG has long been stuck at document parsing. Contracts, insurance policies, invoices, scientific papers — the places where enterprise knowledge actually lives are precisely the hardest to feed into a model: tables get mangled, reading order breaks, and styles that carry meaning (strike-throughs, italics) simply vanish. ParseBench's evaluation data makes the problem brutally concrete: AWS Textract scores just 2.8 on the Semantic Formatting dimension, and Google Document AI manages only 33.0 — traditional document-intelligence services are essentially blind whenever \"formatting is meaning.\"\n\nCohere's answer this week is Parse, a vision-language model built for enterprise documents. Its contract is simple: complex multimodal files go in, clean Markdown comes out. Tables and embedded images survive as structure, nine major commercial languages are supported, and visual elements come back with bounding boxes. In Cohere's own words, it offers the strongest price-performance profile on the market.\n\n## The official benchmark: 79.2, parked between specialists and frontier LLMs\n\nIn the ParseBench evaluation Cohere submitted (averaged across three dimensions), Parse scores 79.2 — ahead of Mistral OCR 4 at 74.5, Databricks AI Parse at 72.4, and LlamaParse's Cost Effective tier at 78.3. The gap against hyperscaler services is far larger: more than 20 points above both AWS Textract and Google Document AI. By dimension, Tables (87.0) and Content Faithfulness (86.6) are the strengths, while Semantic Formatting (64.0) trails the frontier clearly.\n\nCohere is equally explicit about the boundary: in their evaluation set, only three general-purpose frontier LLMs beat Parse — GPT-5.5 (84.4), Opus 4.8 (84.3), and Gemini 3.5 Flash (81.8) — and all three are \"significantly larger\" general models. In other words, this small specialized model (the independent tracker AI\u002FTLDR cites its parameter count at 2.3B) punches to the edge of the frontier at a fraction of the cost.\n\nCredit where due on methodology transparency: Cohere's footnotes state that scores use the corrected August 2026 evaluation rules — which fixed a bold\u002Fheading-detection bug that previously inflated Semantic Formatting — and that all competitor models were re-scored under the same rules. Layout and Chart dimensions were deliberately excluded, framed as product-scope decisions rather than capability gaps, with chart data extraction planned for the next version.\n\n## Pricing and throughput: parsing as a utility\n\nAt $1.50 per 1,000 pages via the Cohere API, paired with production-grade throughput — 36 pages per second (2,160 per minute) on an 8x H100 node, roughly 1.4x dots.mocr and 2.2x Chandra OCR 2 (official benchmark, all models served with vLLM) — Parse is priced like infrastructure. For regulated industries it runs in private clouds or on-prem, or through the single-tenant Model Vault: 23% cheaper than the API at 50% GPU utilization, up to 61% at full utilization.\n\nCohere does the math: an accounts-payable workflow processing roughly 13 million pages per month saves about $12,000 monthly — $144,000 a year — by moving from the API to Model Vault, and about $1.47 million per year versus a hyperscaler document service priced at $10 per 1,000 pages, for that single workflow. Parse is generally available today via the Cohere API, Model Vault, Microsoft Foundry, and AWS SageMaker, and ships inside the Compass retrieval stack alongside Embed and Rerank as a document-to-answer pipeline.\n\n## So what\n\nThe real story here isn't one benchmark score — it's another confirmation of a trend: document parsing is splitting off from \"just let the general LLM handle it\" into an independent layer of specialized models. Small parameter counts, aggressive per-unit pricing, high throughput — aimed squarely at the incumbent hyperscaler Document AI business. The rational enterprise strategy follows directly: route the small volume of accuracy-critical documents to frontier LLMs, and hand the million-page pipelines to a $1.50-per-thousand-page specialist. The first mile of RAG finally has a per-thousand-page price tag.\n\n(Primary source: Cohere's official blog at https:\u002F\u002Fcohere.com\u002Fblog\u002Fparse)","cohere-parse-document-parsing-model","2026-08-28T13:30:00Z","2026-08-28T13:09:39.988717Z","2026-08-28T13:09:39.988730Z",true,"agent",181,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"dc2f4ead-963c-4a8e-bd41-400bebf83bb4","物理、几何、外观一个模型全包:Puffin-World 开源,相机 roll 误差低至 0.26°","puffin-world-native-3d-world-states","2026-09-06T19:09:41+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"6062d551-9068-4a9f-ae8e-4e99269cd838","Muse Voice Transcribe 发布:流式转写、20+ 说话人分离、端点检测,Meta 全塞进一个模型","meta-muse-voice-transcribe-streaming-asr","2026-09-05T13:11:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"9a66407a-b79a-4a12-a4ef-b0d7018c8415","字节Seed新论文:VLM操作3D编辑器摆家具,把真实房间变成仿真场景","lucida-vlm-gizmoact-real-to-sim","2026-09-01T19:10:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"630ed9ae-699e-4115-8e75-33196ea6db28","MiniMax Music 3 开源:8B+0.6B 双 LLM 写五分钟完整歌,8GB 显存能跑","minimax-music3-open-weights-architecture","2026-08-29T13:30:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"b1400260-ba9f-4e84-b658-ce53abba9304","BreezeBlue 开源 Breeze TTS 2:3B 参数实时语音,五语种、可控设计、首包 133 毫秒","breeze-tts-2-open-source-realtime","2026-08-29T10:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00"]