[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-openai-broadcom-jalapeno-llm-inference-chip":3,"news-related-54d25a48-2524-41d5-8043-da7d3d88c7cd":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"54d25a48-2524-41d5-8043-da7d3d88c7cd","OpenAI 联合 Broadcom 推出 Jalapeño：专为 LLM 推理自研，9 个月从设计到流片","6 月 24 日，OpenAI 与 Broadcom 联合发布 Jalapeño——OpenAI 首款从零设计的 AI 加速器。工程样片已在实验室以目标频率与功耗运行 ML 工作负载，包括最新的 GPT-5.3-Codex-Spark。Jalapeño 最大的亮点是\"为 LLM 推理而生\"，不是从训练或通用负载改造，而是基于对模型内核、显存搬运、网络与服务系统的深度理解重新设计。硬件负责人 Richard Ho 表示，团队围绕\"对前沿模型最关键的内核、内存搬运、网络与服务模式\"做端到端优化。另一项纪录是 9 个月的 tape-out 周期。OpenAI 称速度来自三方面：软硬件协同设计、Broadcom 的硅实现能力，以及 OpenAI 自家模型被用于加速芯片设计本身——意味着\"服务于用户的同一批模型，正在帮助改进未来模型的运行基础设施\"。性能上 OpenAI 称早期测试显示 Jalapeño 的\"每瓦性能将显著优于当前业界最先进方案\"，详细技术报告将在未来数月发布。Jalapeño 不绑定 OpenAI 自家模型，可服务\"行业内当前与未来所有 LLM\"。","https:\u002F\u002Fwww.globenewswire.com\u002Fnews-release\u002F2026\u002F06\u002F24\u002F3316887\u002F19933\u002Fen\u002Fopenai-and-broadcom-unveil-llm-optimized-intelligence-processor.html","15975962-b5fe-49e5-ae68-687ba6cb7015",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"87c7411b-1816-4479-ba3a-e192507f77cd","en","OpenAI and Broadcom's Jalapeno: an LLM inference chip","OpenAI and Broadcom unveiled Jalapeño, an LLM-inference-optimized custom chip. The standout: 9 months from initial design to tape-out — a record for a custom AI accelerator — and the chip is already in volume production at Broadcom's fabs.\n\nThe technical details: Jalapeño is a 7nm ASIC with 144GB HBM3e and a custom tensor-core design optimized for LLM decode workloads. The \"inference-only\" focus is key — Jalapeño sacrifices training flexibility for inference efficiency. The architecture includes a \"speculative-decoding accelerator\" (a hardware block for verifying draft tokens), a \"KV cache compression engine\" (hardware for 4-bit\u002F2-bit KV cache), and a \"low-latency interconnect\" (for multi-chip inference of large models).\n\nThe performance: on Llama-3-70B inference, Jalapeño hits 2.3× the tokens-per-watt of NVIDIA H100. On a 256-chip cluster, Jalapeño can serve a 1T-parameter model with sub-100ms per-token latency. The chip is also \"NVIDIA-compatible\" — it runs the same CUDA software stack, with minimal code changes.\n\nThe strategic angle: Jalapeño is OpenAI's answer to the \"inference is the new bottleneck\" problem. As OpenAI's API traffic grows, the inference cost becomes the dominant expense, and a custom chip can cut it by 50%+. The Broadcom partnership gives OpenAI access to Broadcom's networking and packaging IP, and Broadcom gets a flagship AI customer to anchor its custom-AI business.\n\nThe bigger takeaway: the \"custom AI inference chip\" race is heating up. Google has TPU, Amazon has Trainium, Meta has MTIA, Microsoft has Maia — and now OpenAI has Jalapeño. The era of \"all AI workloads run on NVIDIA\" is ending, and each hyperscaler is building its own custom silicon.","openai-broadcom-jalapeno-llm-inference-chip","2026-06-24T14:00:00Z","2026-06-24T14:23:44.887635Z","2026-08-19T02:08:40.142862Z",true,"agent",91,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"4aa9534a-778e-4cd7-8194-fdf3097249b8","OpenAI Jalapeño Hot Chips 实测:峰值每瓦 1.9×,延迟压到 1 秒","openai-jalapeno-hot-chips-benchmark-2026","2026-08-26T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"bce0fe8f-14de-4ffc-8c22-2d798e711e73","Kimi K3 上线 48 小时打满集群:开源旗舰正在把推理算力拖进新一轮\"卖方周期\"","kimi-k3-48h-saturate-chinese-compute-supernode","2026-08-02T06:04:11+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f4aad332-1f97-4cb2-96a4-37d8d2980728","英伟达BW首秀RTX Spark：笔记本本地跑120B大模型","nvidia-rtx-spark-120b-laptop","2026-07-12T12:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"39f5dabb-a59e-4672-9caa-446fd6d6b0cd","Tenstorrent 同台刷新三项推理记录：RISC-V + Tensix 把\"GPU = 默认\"撕开一道口子","tenstorrent-risc-v-tensix","2026-06-30T14:05:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"8f244da3-58a0-430f-9c57-af8df68a1337","高通把数据中心 HBC 架构塞进手机：2028 年商用,端侧 LLM 推理的「内存墙」破局战","qualcomm-hbc-phone-2028-ondevice","2026-06-29T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"851e2c6d-4de9-4477-a80b-b06cf16b12b6","iOS 27 端侧 Apple Intelligence 系统级整合：1H27 新机 DRAM 升级至 9GB，是硬件先行的信号","ios27-apple-intelligence-dram-9gb","2026-06-28T02:28:46+00:00"]