[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-anthropic-clive-chan-perplexity-picojoule":3,"topics-all":36,"news-related-02f573b0-2cae-45dc-8bf5-b226143f79ec":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"02f573b0-2cae-45dc-8bf5-b226143f79ec","Anthropic 招揽 OpenAI 芯片元老：「Perplexity per Picojoule」开启大模型能效新范式","Anthropic 近日从 OpenAI 挖走自研芯片项目\"002 号员工\" Clive Chan，他在 LinkedIn 上的职位描述只有一句：\"perplexity per picojoule\"——把模型预测能力与单位能耗放进同一个优化目标。\n\n这句话折射出大模型评估范式的迁移。传统指标 FLOPS、tokens\u002Fsec、MMLU 关注\"算得快\"，而 perplexity per picojoule 把能耗摆到一等公民位置。背后有三股力：规模撞上电力墙——GPT-5.4、Claude Mythos 在 256K 上下文下推理能耗已逼近数据中心承载上限；硬件-软件协同设计回归——Anthropic 评估自研 ASIC 加上 Chan 熟稔 OpenAI-Broadcom 自研芯片项目；端侧 AI 倒逼能效优先——Anemll、Ollama MLX、WWDC 押注的端侧模型让\"每焦耳 token 数\"成为产品级指标。\n\nperplexity per joule 类指标 2025 年已在 arXiv 出现（d-Matrix 的 roofline 建模与硬件协同设计论文），并非 Anthropic 首创。但当顶级实验室把它写进招聘 JD 并组建专门团队，意味着它已从学术讨论进入工业级落地。\n\n未来模型选择标准可能从\"MMLU 多少分\"或\"每千 token 成本\"转向\"固定功耗预算下能跑多准\"，反向推动稀疏 MoE、低秩近似、4\u002F2\u002F1.58-bit 量化与投机解码的协同进化。可以预期，2026 下半年起，\"perplexity per joule\" 会像当年的 cost-per-token 一样，成为云厂商比较 LLM 推理性价比的新基准。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3842586501466625","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"47a21a92-9477-4b7b-acb5-18b6d6991ed7","en","Anthropic hires OpenAI chip veteran: perplexity per picojoule","Anthropic recently poached \"employee #2\" of OpenAI's in-house chip project, Clive Chan, from OpenAI. His LinkedIn position description is just one sentence: \"perplexity per picojoule\" — putting the model's predictive capability and the unit of energy consumption into the same optimization target.\n\nThis sentence reflects the migration of the large-model evaluation paradigm. Traditional metrics FLOPS, tokens\u002Fsec, and MMLU focus on \"compute fast,\" while perplexity per picojoule puts energy consumption front and center. Three forces are behind it: scale is hitting the power wall — GPT-5.4 and Claude Mythos at 256K context have inference energy consumption approaching the data-center carrying capacity; hardware-software co-design is back — Anthropic is evaluating in-house ASICs plus Chan's familiarity with the OpenAI-Broadcom in-house chip project; and on-device AI forces energy efficiency first — Anemll, Ollama MLX, and the on-device models WWDC is betting on make \"tokens per joule\" a product-level metric.\n\nperplexity per joule-class metrics already appeared on arXiv in 2025 (d-Matrix's roofline modeling and hardware co-design papers), not first coined by Anthropic. But when a top lab writes it into a recruiting JD and assembles a dedicated team, it means it has moved from academic discussion to industrial-scale deployment.\n\nThe future model selection criterion may shift from \"how many MMLU points\" or \"cost per thousand tokens\" to \"how accurate can you run under a fixed power budget,\" in turn driving the co-evolution of sparse MoE, low-rank approximation, 4\u002F2\u002F1.58-bit quantization, and speculative decoding. It can be expected that from the second half of 2026, \"perplexity per joule\" will, like the cost-per-token of yesteryear, become the new benchmark for cloud-vendor comparison of LLM inference cost-performance.","anthropic-clive-chan-perplexity-picojoule","2026-06-08T00:30:00Z","2026-06-08T00:28:15.552665Z","2026-08-19T02:08:40.142862Z",true,"agent",169,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"511ee44b-41e1-4de8-a437-9cbb2ff22cf2","Google 收紧 Android 内存红线:AI 数据中心抢走 DRAM,手机 App 也要瘦身","google-android-memory-limit-ai-dram-crunch","2026-08-29T08:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"5d452086-ecb7-494d-82b9-d962664aa243","IBM 把 Arm 核塞进 Z 大型机:Hot Chips 2026 公布业界首款双指令集处理器","ibm-z-arm-dual-isa-hot-chips-aug-2026","2026-08-29T06:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"5da227db-f53d-4b07-a0c6-4ea16e04cd4d","CURE 用不确定性焦点做「block-parallel 投机解码」：端到端 2.66–3.49×、接受长度涨 4.2–7.5%","cure-block-parallel-speculative-decoding","2026-08-08T02:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"183fb3be-e062-47e7-9591-7c2372e116c1","LLM 蒸馏的显存瓶颈不只在教师模型：离线 Top-K 与分块 KL 把长上下文训练装回单卡","llm-distillation-offline-top-k-chunked-kl","2026-08-05T20:08:13+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"e3f049f5-2f0d-48d2-8e88-246ef006fa16","LoopMTP 给循环 Transformer 装上前瞻路标：固定参数下让每一轮都做不同的事","loopmtp-latent-multi-token-loop-guidance","2026-08-04T13:13:09+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"cdc8e3ce-b1aa-4348-9436-04763179af9c","AMD MI455X：Transformers 99.5% 通过率，432GB HBM4","amd-mi455x-huggingface-99-5","2026-07-27T10:30:00+00:00"]