[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-openai-gpt-oss-120b-apache-2-moe-int4":3,"topics-all":36,"news-related-551dfd06-3bff-47d1-a7e4-a420a9c90c22":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"551dfd06-3bff-47d1-a7e4-a420a9c90c22","OpenAI发布GPT-OSS 120B：七年后重返开源，单卡部署的边界被重新定义","2019年，OpenAI发布GPT-2之后就彻底关上了开源的大门。七年后，这家公司做了一件几乎没人预料到的事——发布开源权重模型。\n\nGPT-OSS 120B于2026年4月正式亮相，包含gpt-oss-120b和gpt-oss-20b两个版本，全部采用Apache 2.0许可证，可自由下载、商业使用、修改和再分发。这不仅是OpenAI自GPT-2以来首次开源模型，更是其首次真正意义上进入开源权重模型的竞争格局。\n\n技术架构上，GPT-OSS 120B采用了改进的MoE（混合专家）架构。1170亿参数总量中，每次前向传播仅激活约390亿参数，这一设计使模型能在单张80GB显存的GPU上通过INT4量化运行。相比之下，GLM-5.1需要4张H200才能运行，DeepSeek V4则需要8张GPU。在硬件成本上，GPT-OSS 120B的部署门槛是三款顶级开源模型中最低的。\n\n另一个差异化亮点是链式思维（CoT）的完全可视化。GPT-OSS 120B提供了三个推理档位——快速、平衡、深度，用户可以实时观察模型的完整推理过程，并精细控制推理深度。这种透明度在开源模型中前所未有。\n\n基准测试方面，GPT-OSS 120B在编程和数学任务上与GPT-4o相当，并超越参数量为其三倍的Llama 3.1 405B。但受限于Apache 2.0许可证的月活7亿门槛条款，其商业应用存在一定限制。\n\nOpenAI的入局打破了原本Meta、阿里、DeepSeek三方竞争的开源格局。GPT-OSS 120B的真正意义不在于性能全面超越，而在于它改变了游戏规则：单卡部署顶级模型的能力，将AI推理的硬件门槛从集群级降低到了单卡级别。","https:\u002F\u002Fgithub.com\u002Fopenai\u002Fgpt-oss","998df6db-96e6-4b8e-8be1-cfa00a6cd177",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"ebf11195-f515-4f45-bc25-ae8266661c0f","en","GPT-OSS 120B: OpenAI returns to open source after seven years","In 2019, OpenAI released GPT-2 and then completely shut the open-source door. Seven years later, this company did something almost no one anticipated — releasing open-source-weight models.\n\nGPT-OSS 120B officially debuted in April 2026, including gpt-oss-120b and gpt-oss-20b versions, all under Apache 2.0 license, freely downloadable, commercially usable, modifiable, and redistributable. This isn't just OpenAI's first open-source model since GPT-2, but its first real entry into the open-source-weight-model competitive landscape.\n\nOn technical architecture, GPT-OSS 120B adopts an improved MoE (Mixture of Experts) architecture. Of the 117B total parameters, only about 39B are activated per forward pass — a design that lets the model run on a single 80GB VRAM GPU with INT4 quantization. In contrast, GLM-5.1 requires 4×H200 to run, DeepSeek V4 requires 8 GPUs. On hardware cost, GPT-OSS 120B has the lowest deployment threshold of the three top open-source models.\n\nAnother differentiation highlight is the complete visualization of chain-of-thought (CoT). GPT-OSS 120B offers three reasoning tiers — fast, balanced, deep — letting users observe the model's complete reasoning process in real time, with fine-grained control over reasoning depth. This transparency is unprecedented in open-source models.\n\nOn benchmarks, GPT-OSS 120B is on par with GPT-4o on coding and math tasks, and surpasses Llama 3.1 405B with three times its parameter count. But limited by the Apache 2.0 license's 700M MAU threshold clause, its commercial application has certain limits.\n\nOpenAI's entry breaks the open-source competitive landscape previously dominated by Meta, Alibaba, and DeepSeek. GPT-OSS 120B's real significance isn't that it fully surpasses in performance, but that it changes the rules: the ability to deploy top-tier models on a single card has lowered AI inference's hardware threshold from cluster-level to single-card level.","openai-gpt-oss-120b-apache-2-moe-int4","2026-04-27T07:10:00Z","2026-04-27T07:12:06.347495Z","2026-08-19T02:08:40.142862Z",true,"agent",141,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"36055e5f-136f-497d-8763-3ed6609f59ff","Meta Muse Glimmer 30B 本地落地:Apache 2.0 的开源智能体,把 Agent 装进 24GB 显存","meta-muse-glimmer-30b-local-agent-apache2-r2","2026-08-19T03:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"0d6b1c39-3b2c-48c4-afb4-41a0de2a918d","GLM-5.2 把 1M 上下文\"焊\"进开源：IndexShare + 反作弊 RL，把长程 Agent 拉成工程现实","glm-5-2-z-ai-indexshare-1m-anti-cheat-rl","2026-06-19T20:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"4d436945-18e9-4d69-a4c8-c1e3e975ab33","MiniMax M3发布：稀疏注意力打通百万token上下文，开源模型编程能力逼近闭源前沿","minimax-m3-sparse-attn-million-token-msa","2026-06-04T01:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"98695785-30b2-4ddf-9886-757e57773f8f","Arcee Trinity Large：400B开源MoE模型挑战Claude Opus，定价便宜96%","arcee-trinity-large-400b-moe-claude-opus-96pct-cheaper","2026-04-30T07:01:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"b667e52f-ec7d-4ca4-8d9e-1db81e1a5616","DeepSeek论文:890字节KV缓存的三层架构账","deepseek-v41-flash-kv-cache-paper","2026-09-18T15:10:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"d056f67b-7e0d-4e44-8d39-e31ea50deeae","Bonsai 2 27B 三元压缩:Qwen3.8 压到 5.9 GB,benchmark 留存 98.2%","bonsai-2-27b-ternary-qwen3-8-compression","2026-09-17T15:47:00+00:00"]