[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-openai-jalapeno-hot-chips-broadcom-blackwell":3,"topics-all":39,"news-related-97eb895f-8294-47b9-a9e0-5f165a75cb5b":58},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":37,"view_count":38},"97eb895f-8294-47b9-a9e0-5f165a75cb5b","OpenAI Jalapeño 实测:自研推理芯片在 Hot Chips 上跑赢 Blackwell","OpenAI 公布 Jalapeño 首批 InferenceX 实测:峰值每瓦 1.5–1.9× 于对比系统、延迟低 1.7–3.6×;AI 编程内核部分模块比人工实现快 1.5–1.8×。2026 年底小规模上线,Gen 2\u002F3 已启动。","OpenAI 在 Hot Chips 2026 公布了旗下首款自研推理芯片 Jalapeño 的首批实测数据。这是 OpenAI 与 Broadcom 联合研发的通用推理芯片,瞄准的是 AI 算力真正吃紧的「用户响应」一环。\n\n## Hot Chips 首测:1.5–1.9× 每瓦,延迟压到 1 秒\n\n基准是 SemiAnalysis 的 InferenceX。Jalapeño 在三个公开模型上跑出稳定领先:在 GPT‑OSS 120B(8k\u002F1k)上,峰值每瓦对比 GB200 高约 1.9×,端到端延迟从 1.80s 降到 1.03s;在 DeepSeek R1 670B(MXFP4)上,峰值每瓦 1.7×,延迟从 5.99s 降到 1.65s;在 Kimi K2.5 1T 上,峰值每瓦 1.5×,延迟从 5.31s 降到 1.56s。从超低延迟对话到高吞吐批处理,Jalapeño 处于 Pareto 前沿。\n\n数字之外,架构选择更值得关注。OpenAI 把芯片定位为「为 LLM 推理而生」,围绕 prefill、decode、通信三段瓶颈做了不同资源配比:prefill 计算重,decode 受限于显存带宽,跨芯片通信又会拖累算力闲置。Jalapeño 用一张统一架构同时吃下三段:把 KV cache 显式放本地,用大域网络把整个请求留在一个连通系统里,降低数据搬移和同步开销。\n\n## AI 写内核,九个月从设计到流片\n\n另一个细节,是 AI 在芯片研发自身的角色。OpenAI 用自家模型加速了设计、验证与迭代,从立项到 tapeout 只用了九个月;同时把芯片做得「人和 AI 都能用」——用 Codex + GPT‑Astra 做内核生成,部分 GPT‑OSS attention 与 MoE 模块,AI 生成的内核跑出比人类专家手写版本快 1.5–1.8× 的速度。数字虽只针对「选定模块」而非全模型,但方向意味很强:AI 不只是被推理芯片服务,也在反向成为芯片开发的「编译器」。\n\n## 时间表与对英伟达的真正含义\n\n硬件负责人 Richard Ho 给的节奏是:Jalapeño 2026 年底「很小规模」上线,真正大规模部署在 2027;Gen 2 深入开发,Gen 3 已经成形。这与英伟达在 Hot Chips 上谈「电力供应瓶颈」的发言在同一周,两家都把瓦数变成头号约束变量。OpenAI 也重申「不会脱离英伟达」:将继续广泛部署英伟达与其他合作伙伴的加速器。Jalapeño 不是替代,而是补位——补的恰恰是英伟达体系内瓦数成本最低的那一段,也即模型上线后被反复调用的那一段。\n\n## 所以呢:硅片卷完之后,真正卷的是「每瓦」\n\n这件事给出清晰信号——AI 算力的下一阶段不在「谁的 GPU 更大」,而在「谁能在每一瓦电里挤出更多 token」。Agent、长上下文、实时音视频这类对延迟和持续功耗敏感的场景,会最先吃到红利;反过来,继续靠「堆卡换智能」的玩家,边际收益正在肉眼可见地收窄。\n\n参考:[Jalapeño first results — OpenAI](https:\u002F\u002Fopenai.com\u002Findex\u002Fjalapeno-first-results\u002F)、[OpenAI's Jalapeño chip — TechCrunch](https:\u002F\u002Ftechcrunch.com\u002F2026\u002F08\u002F25\u002Fopenais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show\u002F)","https:\u002F\u002Fopenai.com\u002Findex\u002Fopenai-broadcom-jalapeno-inference-chip\u002F","15975962-b5fe-49e5-ae68-687ba6cb7015",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"15f54ac7-cf59-4867-8474-ef78970b9658","en","OpenAI Jalapeño benchmarked: in-house inference chip beats Blackwell on per-watt at Hot Chips","OpenAI released Jalapeño's first InferenceX benchmarks at Hot Chips 2026: peak performance per watt 1.5–1.9× the comparison system and end-to-end latency 1.7–3.6× lower across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T; AI-generated kernels for selected blocks ran 1.5–1.8× faster than human-expert versions. Small-volume deployment is slated for end of 2026, with Gen 2 and Gen 3 already in development.","At Hot Chips 2026, OpenAI released the first batch of measured data for its in-house inference chip, Jalapeño. The chip is a general-purpose inference accelerator co-developed with Broadcom, aimed squarely at the part of AI compute that is truly tight today: the user-facing response loop.\n\n## Hot Chips first results: 1.5–1.9× per watt, latency compressed to ~1 second\n\nThe benchmark is SemiAnalysis's InferenceX. Jalapeño turned in a stable lead across three public models: on GPT-OSS 120B (nominal 8k\u002F1k), peak mixed TPS per kilowatt ran roughly 1.9× the GB200 comparison, with end-to-end latency falling from 1.80s to 1.03s; on DeepSeek R1 670B (MXFP4), peak per-watt was 1.7× and latency dropped from 5.99s to 1.65s; on Kimi K2.5 1T, peak per-watt was 1.5× and latency dropped from 5.31s to 1.56s. Across the full operating range, from ultra-low-latency conversation to high-throughput batch, Jalapeño sits on the Pareto frontier.\n\nBeyond the numbers, the architectural choices matter more. OpenAI positioned the chip as \"built for LLM inference from day one,\" with different resource balances for the three bottleneck phases of inference: prefill is compute-heavy, decode is constrained by memory bandwidth, and inter-chip communication can leave compute units idle while waiting on data. Jalapeño tackles all three with a single architecture: KV cache is placed explicitly and kept local, and a large-domain network keeps the entire request inside one connected system, reducing data movement and synchronization overhead.\n\n## AI-written kernels, nine months from design to tapeout\n\nAnother detail worth flagging is AI's place in the chip development loop itself. OpenAI used its own models to compress the design, verification, and iteration cycle — from project kickoff to tapeout took only nine months. The team also made the hardware a programming target that \"both humans and AI can use\": with Codex plus GPT-Astra for kernel generation, AI-produced kernels for selected GPT-OSS attention and MoE blocks ran 1.5–1.8× faster than human-expert-written versions. The figures apply to selected blocks, not the full model, but the direction is strong: AI is not only being served by inference chips, it is becoming the compiler of chip development in return.\n\n## Timeline and what it actually means for Nvidia\n\nHardware lead Richard Ho set the cadence: Jalapeño deploys at the end of 2026 in \"very small volumes,\" with meaningful scale coming in 2027; Gen 2 is deep in development and Gen 3 is already taking shape. That timing lines up with Nvidia's Hot Chips talk on the same week framing power supply as the binding constraint — both companies have made watts the headline knob. Notably, OpenAI also reiterated it \"is not decoupling from Nvidia\": it will keep deploying Nvidia and other partner accelerators broadly for both training and inference. Jalapeño is not a replacement, it is a complement — filling exactly the lowest-cost-per-watt slot in Nvidia's stack, the slot that gets hammered once a model is in production and being called over and over.\n\n## So what: once silicon is a solved game, the real race is \"per watt\"\n\nThe signal for practitioners is clear: the next stage of AI compute is not about \"who has the bigger GPU,\" it is about \"who can squeeze more tokens out of every watt.\" Agentic workloads, long context, and real-time audio\u002Fvideo — the latency- and sustained-power-sensitive regimes — will be the first to feel the benefit. Conversely, players leaning on \"more cards for more intelligence\" are seeing their marginal returns visibly compress.\n\nReferences: [Jalapeño first results — OpenAI](https:\u002F\u002Fopenai.com\u002Findex\u002Fjalapeno-first-results\u002F), [OpenAI's Jalapeño chip — TechCrunch](https:\u002F\u002Ftechcrunch.com\u002F2026\u002F08\u002F25\u002Fopenais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show\u002F)","openai-jalapeno-hot-chips-broadcom-blackwell","2026-08-26T08:00:00Z","2026-08-27T09:06:17.944179Z","2026-08-27T09:06:17.944194Z",true,"agent","https:\u002F\u002Fimages.ctfassets.net\u002Fkftzwdyauwt9\u002F5OsuPfmq6Km8mvw4podaWa\u002Fbe9c8ec1bc37f6a3bd177da03a7fe03f\u002FSEO_Card__2_.png?w=1600&h=900&fit=fill",300,[40,49],{"slug":41,"tag_slug":41,"title_zh":42,"title_en":43,"intro_zh":44,"intro_en":45,"id":46,"is_active":35,"created_at":47,"modified_at":48},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":50,"tag_slug":50,"title_zh":51,"title_en":52,"intro_zh":53,"intro_en":54,"id":55,"is_active":35,"created_at":56,"modified_at":57},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":59},[60,65,70,75,80,85],{"id":61,"title":62,"news_slug":63,"published_at":64},"4aa9534a-778e-4cd7-8194-fdf3097249b8","OpenAI Jalapeño Hot Chips 实测:峰值每瓦 1.9×,延迟压到 1 秒","openai-jalapeno-hot-chips-benchmark-2026","2026-08-26T02:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"54d25a48-2524-41d5-8043-da7d3d88c7cd","OpenAI 联合 Broadcom 推出 Jalapeño：专为 LLM 推理自研，9 个月从设计到流片","openai-broadcom-jalapeno-llm-inference-chip","2026-06-24T14:00:00+00:00",{"id":71,"title":72,"news_slug":73,"published_at":74},"5d452086-ecb7-494d-82b9-d962664aa243","IBM 把 Arm 核塞进 Z 大型机:Hot Chips 2026 公布业界首款双指令集处理器","ibm-z-arm-dual-isa-hot-chips-aug-2026","2026-08-29T06:00:00+00:00",{"id":76,"title":77,"news_slug":78,"published_at":79},"7bb4f5ec-14e0-43b6-9913-07cad82a520b","微软内部 AI 账单失控:单员工月烧 2.8 万美元,倒逼默认模型换人","microsoft-internal-ai-bill-explode-default-model-swap","2026-08-28T04:00:00+00:00",{"id":81,"title":82,"news_slug":83,"published_at":84},"bce0fe8f-14de-4ffc-8c22-2d798e711e73","Kimi K3 上线 48 小时打满集群:开源旗舰正在把推理算力拖进新一轮\"卖方周期\"","kimi-k3-48h-saturate-chinese-compute-supernode","2026-08-02T06:04:11+00:00",{"id":86,"title":87,"news_slug":88,"published_at":89},"f4aad332-1f97-4cb2-96a4-37d8d2980728","英伟达BW首秀RTX Spark：笔记本本地跑120B大模型","nvidia-rtx-spark-120b-laptop","2026-07-12T12:00:00+00:00"]