[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-pytorch-2-13-flexattention-apple-silicon":3,"news-related-84ea21f5-8ed4-46db-b5a3-f0584eaa90d4":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"84ea21f5-8ed4-46db-b5a3-f0584eaa90d4","PyTorch 2.13：FlexAttention 上 Apple Silicon，稀疏注意力 12×","PyTorch 2.13 于 2026 年 7 月 8 日发布,由 526 位贡献者合计提交 3328 个 commit。FlexAttention 正式登陆 Apple Silicon(MPS 后端),手写 Metal 内核覆盖 sparse prefill 与 decode,在 1×8×32768×64 长序列 + 256 滑动窗口(密度 0.8%)的稀疏模式下对 SDPA 拿到约 12.3× 加速,8192\u002F64 窗口场景也吃到约 4.15×;CUDA 端 Flash backend 加入确定性反向路径,用预计算写序替换 atomic,实现 bit-for-bit 可复现,长序列端到端开销 +0.2%。新 nn.LinearCrossEntropyLoss 把线性层与交叉熵融合为一个算子、沿词表维度分块流式计算,大词表 LLM 训练峰值显存最多省 4×。CuTeDSL \"Native DSL\" 后端为 Inductor 提供 CUTLASS 级 GEMM\u002FRMSNorm 代码生成;torchcomms 替代 c10d,FSDP2 支持 reduce-scatter 与 all-gather 通信重叠。ExecuTorch 正式并入 PyTorch Core,端侧推理成一等公民。","https:\u002F\u002Fpytorch.org\u002Fblog\u002Fpytorch-2-13-release-blog\u002F","8a980003-65e9-4d31-b870-94cd12fa0d46",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"7049e400-7a06-4618-ac64-92a62a43769c","en","PyTorch 2.13 brings FlexAttention to Apple Silicon, 12x","PyTorch 2.13 was released on July 8, 2026, with 526 contributors combining to submit 3328 commits. FlexAttention officially lands on Apple Silicon (MPS backend), with hand-written Metal kernels covering sparse prefill and decode. On 1×8×32768×64 long sequences + 256 sliding window (0.8% density) sparse mode, it gets about 12.3× speedup over SDPA; in the 8192\u002F64 window scenario it still gets about 4.15×; the CUDA-side Flash backend adds a deterministic backward path, replacing atomic with precomputed write order, achieving bit-for-bit reproducibility with only +0.2% end-to-end overhead on long sequences. The new nn.LinearCrossEntropyLoss fuses the linear layer and cross-entropy into a single operator, streaming computation in chunks along the vocabulary dimension, cutting peak training memory for large-vocabulary LLMs by up to 4×. The CuTeDSL \"Native DSL\" backend provides CUTLASS-grade GEMM\u002FRMSNorm code generation for Inductor; torchcomms replaces c10d, and FSDP2 supports overlapping reduce-scatter and all-gather communication. ExecuTorch is officially merged into PyTorch Core, making on-device inference a first-class citizen.","pytorch-2-13-flexattention-apple-silicon","2026-07-08T16:00:00Z","2026-07-16T12:09:54.183187Z","2026-08-19T02:08:40.142862Z",true,"agent",115,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"deac2d55-76a6-40d2-8ef7-36aed2ad0105","Linux 7.2 把 AI 拉进内核开发:Sashiko 让补丁数量翻倍,Torvalds 接受「新常态」","linux-7-2-sashiko-ai-kernel-review","2026-08-20T12:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"a91067a3-4fa4-4e88-a25a-18ba3bea21ea","Google 把\"加密推理\"摆上桌面：HEIR 编译器让预训练模型在密文上直接跑","google-heir-compiler-encrypted-ai-inference","2026-08-14T14:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"22a1a718-0eb6-46e5-8ee8-825400de11d1","DeepMind WeatherNext 在 Nature 发论文：用 28 km 粗分辨率做出多一天的飓风预警,代码权重全部开源","deepmind-weathernext-cyclones-nature-open-source","2026-08-10T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"9ef626d9-05dd-4e67-9069-093f90a3fd5c","Rust 主仓库正式引入 LLM 政策：把\"必须人为可读、不可代写\"写进 PR 流程","rust-lang-rust-llm-policy-adoption","2026-08-07T00:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"b2c169c6-5150-4423-8073-bf480a2d8745","腾讯 UniPert-G2CP 登《Cell》主刊：把基因扰动和化学扰动塞进同一个语义空间","tencent-unipert-g2cp-cell-virtual-cell","2026-07-31T07:49:00+00:00"]