[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-hf-kernels-hub-first-class":3,"topics-all":36,"news-related-0049e060-f142-4b32-b655-5734b67b346f":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"0049e060-f142-4b32-b655-5734b67b346f","Kernels 大重构:把 GPU kernel 升级为 Hub 一等公民,LLM 基础设施开始标准化","Hugging Face 重构 🤗 Kernels 项目,把 GPU kernel 从零散脚本升级为 Hub 一等公民:新增 kernel 仓库类型,所有 kernel 挂在 huggingface.co\u002Fkernels 下,可直接看到加速器(CUDA\u002FROCm\u002FXPU)、OS、后端版本支持矩阵,源码 Git SHA1 内嵌编译产物,搭配 Nix hermetic build,任何发布版本可被独立复现,告别\"我编出来跟你不一样\"的玄学问题。安全升到协议层:默认只允许 trusted publishers 发布,同时用 Sigstore cosign 内核签名 + GitHub Actions workflow 双重验证,即便账号泄漏攻击者也无法签出恶意版本——把 OCI 供应链安全经验搬到了 ML kernel 生态。为 Agent 工程化铺路:`kernels` 与 `kernel-builder` 两个 CLI 拆分干净,后者刻意做成\"agent-optimized\"(非交互、输出可解析),配合后端 skills 文件,Agent 可端到端脚手架、构建、benchmark 并迭代优化一个 flash-attention 类内核;新增 Torch Stable ABI 向下兼容约 2 年,Apache TVM FFI 让同一份 kernel 跨 PyTorch \u002F JAX \u002F CuPy 运行。flash-attention、Fused MoE、MLA decoding 等高频内核从此 install 即用、可信可复现,不必再\"上传 PyPI + 配 CMake + 写 gotcha 文档\"。Agent 自动生成并验证 kernel 的闭环有了真实落地底座,2026 年下半年 LLM 训练推理栈的\"基础设施标准化\"才刚开跑。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Frevamped-kernels","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"b26792c6-7c8d-403e-88d1-97d54726315e","en","Kernels rebuild: GPU kernels become Hub first-class citizens","Hugging Face refactors the 🤗 Kernels project, upgrading GPU kernels from scattered scripts to first-class Hub citizens: a new kernel repository type, all kernels listed under huggingface.co\u002Fkernels, with a directly visible accelerator (CUDA\u002FROCm\u002FXPU), OS, and backend-version support matrix, source Git SHA1 embedded in compiled artifacts, paired with Nix hermetic build, any release version can be independently reproduced, bidding farewell to the \"I built it different from you\" mystic problem. Security is raised to the protocol layer: only trusted publishers are allowed to release by default, with Sigstore cosign kernel signing + GitHub Actions workflow dual verification — even if an account is compromised, the attacker cannot sign out a malicious version. This ports OCI supply-chain security experience to the ML kernel ecosystem. Paving the way for Agent engineering: the `kernels` and `kernel-builder` CLIs are cleanly split, the latter intentionally \"agent-optimized\" (non-interactive, parseable output), paired with backend skills files, Agents can end-to-end scaffold, build, benchmark, and iteratively optimize a flash-attention-class kernel; a new Torch Stable ABI gives about 2 years of backward compatibility, Apache TVM FFI lets the same kernel run across PyTorch \u002F JAX \u002F CuPy. flash-attention, Fused MoE, MLA decoding and other high-frequency kernels are now install-and-use, trustworthy, and reproducible — no more \"upload to PyPI + write CMake + write gotcha docs\". The closed loop of Agent auto-generating and verifying kernels has a real landing foundation; the \"infrastructure standardization\" of the LLM training-inference stack in the second half of 2026 has just begun running.","hf-kernels-hub-first-class","2026-07-10T04:01:00Z","2026-07-10T04:09:45.499895Z","2026-08-19T02:08:40.142862Z",true,"agent",158,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"d3db2e5d-c2b0-457c-a874-47a8d48ce42f","vLLM Semantic Router v0.3 \"Themis\" 发布：把 LLM 推理路由从「能跑」推进到「可治理」","vllm-semantic-router-themis-0-3-governable","2026-06-05T22:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"12c67d52-17a2-4df5-8386-35d18ffd221a","JEPA-Anything:一套预测框架打通七个领域,湿实验也给了背书","jepa-anything-orthogonal-predictive-factorization","2026-09-19T23:10:37+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"28c41f06-d20f-481c-b133-cd109af3aed1","答对之后停不下来:微软团队揪出在线蒸馏的 EOS 错配元凶","eos-mismatch-opd-length-inflation","2026-09-18T21:09:06+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"cb1e799d-d6d7-4ab9-9eaf-bea0aa432b06","Mistral 模型进驻 Firefox:119B 开放权重模型驱动浏览器 AI 助手","mistral-small-4-firefox-smart-window","2026-09-16T17:07:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"5535b4e4-21de-4ed6-a9bb-2b4d824e6568","F-Droid 一次更新的 102 款应用,72.5% 主要是 AI 写的","f-droid-72-percent-ai-written-audit","2026-09-15T15:12:16+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"dcd8b3e1-a3c7-4614-aba4-9002219ea5f6","LibreDB Studio 0.15 发布:本地 LLM 接管数据库交互","libredb-studio-local-llm-agent","2026-09-15T00:00:00+00:00"]