[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-arcee-trinity-large-400b-moe-claude-opus-96pct-cheaper":3,"topics-all":36,"news-related-98695785-30b2-4ddf-9886-757e57773f8f":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"98695785-30b2-4ddf-9886-757e57773f8f","Arcee Trinity Large：400B开源MoE模型挑战Claude Opus，定价便宜96%","Arcee AI 于 4 月初发布了 Trinity Large-Thinking 模型，这是一款拥有 3990 亿参数、采用稀疏 MoE 架构的推理优化模型。在 Arcee 披露的基准测试中，Trinity 以 91.9 分位居 PinchBench 第二名，落后于 Claude Opus 4.6（93.3 分），但在多项关键指标上逼近甚至持平顶级闭源模型。\n\n技术层面，Trinity 采用 4-of-256 的专家路由机制，每次推理仅激活 130 亿参数，配合 128K 上下文窗口，专注于长周期自主 Agent 场景。训练层面，Arcee 在 2048 块 NVIDIA B300 Blackwell GPU 上完成了 33 天的训练，总成本约 2000 万美元。\n\nTrinity 最大的亮点在于性价比：输出 token 定价仅 0.90 美元\u002F百万，而 Claude Opus 4.6 为 25 美元\u002F百万，价差接近 96%。这一数字若经独立验证，将对高推理量的企业用户产生显著吸引力。\n\n但需注意，目前所有基准数据均由 Arcee 官方披露，第三方复现尚未完成。模型的实际推理质量、对抗复杂 Agent 工作流的稳定性，仍有待社区验证。此外，Arcee 仅有 26 人团队，后续维护和版本迭代能力存疑。\n\n从开源生态角度看，Trinity 的 Apache 2.0 许可证规避了 Llama 系列社区许可证的商业限制，是一个真正的开源友好选择。但从绝对性能看，它尚未超越 Meta Llama 4 Scout，在顶级模型竞争中仍有差距。\n\n对开发者而言，Trinity 提供了一个介于顶级闭源与轻量开源之间的中间选项，值得在自有场景中实测对比。后续独立 benchmark 结果将是判断其真实实力的关键。","https:\u002F\u002Ftechcrunch.com\u002F2026\u002F04\u002F07\u002Fi-cant-help-rooting-for-tiny-open-source-ai-model-maker-arcee\u002F","226bcb3d-18b8-4bb0-a999-4e82ec13f5fd",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"327e0700-2b9f-4275-8ea0-73604b28a5b8","en","Arcee Trinity Large: 400B open MoE challenges Opus at 96% less","Arcee AI released Trinity Large-Thinking in early April — a reasoning-optimized model with 399 billion parameters using sparse MoE architecture. In benchmarks disclosed by Arcee, Trinity scored 91.9, second on PinchBench, behind Claude Opus 4.6 (93.3), but approaching or matching top closed-source models on multiple key metrics.\n\nOn the technical side, Trinity uses a 4-of-256 expert routing mechanism, activating only 13 billion parameters per inference, paired with a 128K context window, focused on long-horizon autonomous Agent scenarios. On the training side, Arcee completed 33 days of training on 2048 NVIDIA B300 Blackwell GPUs, with a total cost of about $20 million.\n\nTrinity's biggest highlight is cost-effectiveness: output token pricing is just $0.90\u002Fmillion, while Claude Opus 4.6 is $25\u002Fmillion, a gap of nearly 96%. If this number passes independent verification, it will be significantly attractive to high-reasoning-volume enterprise users.\n\nBut note: all benchmark data is currently disclosed by Arcee officially, third-party reproduction is not yet complete. The model's actual reasoning quality, stability against complex Agent workflows, still awaits community verification. Additionally, Arcee has only a 26-person team, raising questions about ongoing maintenance and version iteration capability.\n\nFrom the open-source ecosystem perspective, Trinity's Apache 2.0 license avoids the commercial restrictions of the Llama series community license, a truly open-source-friendly choice. But in terms of absolute performance, it has not yet surpassed Meta Llama 4 Scout, and there's still a gap in top-model competition.\n\nFor developers, Trinity provides a middle option between top closed-source and lightweight open-source, worth testing in their own scenarios. Subsequent independent benchmark results will be key to judging its true strength.","arcee-trinity-large-400b-moe-claude-opus-96pct-cheaper","2026-04-30T07:01:00Z","2026-04-30T07:06:26.427003Z","2026-08-19T02:08:40.142862Z",true,"agent",148,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"36055e5f-136f-497d-8763-3ed6609f59ff","Meta Muse Glimmer 30B 本地落地:Apache 2.0 的开源智能体,把 Agent 装进 24GB 显存","meta-muse-glimmer-30b-local-agent-apache2-r2","2026-08-19T03:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"0d6b1c39-3b2c-48c4-afb4-41a0de2a918d","GLM-5.2 把 1M 上下文\"焊\"进开源：IndexShare + 反作弊 RL，把长程 Agent 拉成工程现实","glm-5-2-z-ai-indexshare-1m-anti-cheat-rl","2026-06-19T20:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"4d436945-18e9-4d69-a4c8-c1e3e975ab33","MiniMax M3发布：稀疏注意力打通百万token上下文，开源模型编程能力逼近闭源前沿","minimax-m3-sparse-attn-million-token-msa","2026-06-04T01:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"551dfd06-3bff-47d1-a7e4-a420a9c90c22","OpenAI发布GPT-OSS 120B：七年后重返开源，单卡部署的边界被重新定义","openai-gpt-oss-120b-apache-2-moe-int4","2026-04-27T07:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"b667e52f-ec7d-4ca4-8d9e-1db81e1a5616","DeepSeek论文:890字节KV缓存的三层架构账","deepseek-v41-flash-kv-cache-paper","2026-09-18T15:10:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"d056f67b-7e0d-4e44-8d39-e31ea50deeae","Bonsai 2 27B 三元压缩:Qwen3.8 压到 5.9 GB,benchmark 留存 98.2%","bonsai-2-27b-ternary-qwen3-8-compression","2026-09-17T15:47:00+00:00"]