[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-hf-skypilot-hf-storage":3,"topics-all":36,"news-related-1edd86ea-3eb0-4424-9f57-add15c08d891":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"1edd86ea-3eb0-4424-9f57-add15c08d891","HF Hub 接入 SkyPilot：Xet 去重让 20+ 云共享模型数据","Hugging Face 与 SkyPilot 联合把 Hub 升级为 SkyPilot 的 `store: hf` 一等存储后端,通过 `hf:\u002F\u002F` URL 把任意 model、dataset、Space repo 直接挂载到 AWS、GCP、Lambda、CoreWeave 等 20+ 家云以及 Kubernetes\u002FSlurm 集群的 GPU 任务里,核心由两端实现:(1) 新建 `hf-mount` FUSE 后端 — 每次 `read()` 仅拉需要的那几个字节,避免完整复制,GPU 几乎\"立马上工\",并保留本地 on-disk 缓存让 repeat read 命中本地;(2) 基于 Xet 的内容定义分块 (CDC) — 把模型权重切成约 64 KB 的 chunk,只对真正修改部分重传,统一以 `store: hf` 暴露在 SkyPilot 的 file_mounts 里。基准测试用同一个 `qwen-sft.yaml` 在三家云切换 `--infra` 跑 Qwen3.5-4B SFT,模型首读约 30 秒(峰值 ~500 MB\u002Fs),8.43 GB checkpoint 写入 bucket 在 AWS L40S 上跑到 ~168 MB\u002Fs、GCP L4 约 123 MB\u002Fs、Lambda H100 约 112 MB\u002Fs。Xet 把 dedup 推到 Parquet 行级 append — 内部测试追加 10K 行到 100K 行表只上传约 10 MB,二次上传一个 已存的 8.43 GB blob 只需约 8 秒。这一改动是首次把\"数据稳坐 Hub、算力随便跑\"做成产品级契约,打破了长期以来\"对象存储在哪个云,GPU 调度就被钉死在哪个云\"的部署约束,跨云训练与推理的数据搬运成本被压到接近零。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fskypilot-hf-storage","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"ef347839-b339-4d26-9eb9-bc974fa467bc","en","HF Hub joins SkyPilot: Xet dedup shares models across clouds","Hugging Face and SkyPilot jointly upgrade Hub to SkyPilot's `store: hf` first-class storage backend, mounting any model, dataset, or Space repo directly to GPU tasks on 20+ clouds (AWS, GCP, Lambda, CoreWeave, etc.) and Kubernetes\u002FSlurm clusters via `hf:\u002F\u002F` URLs, with the core implemented at two ends: (1) the new `hf-mount` FUSE backend — each `read()` only pulls the few bytes needed, avoiding full replication, so GPUs are \"almost immediately working\", with local on-disk cache keeping repeat reads hitting local; (2) Xet-based content-defined chunking (CDC) — splitting model weights into ~64 KB chunks, re-uploading only the truly modified parts, uniformly exposed as `store: hf` in SkyPilot's file_mounts. Benchmarks use the same `qwen-sft.yaml` switching `--infra` across three clouds to run Qwen3.5-4B SFT, model first read about 30 seconds (peak ~500 MB\u002Fs), 8.43 GB checkpoint written to bucket hits ~168 MB\u002Fs on AWS L40S, ~123 MB\u002Fs on GCP L4, ~112 MB\u002Fs on Lambda H100. Xet pushes dedup down to Parquet row-level append — internal tests appending 10K rows to a 100K-row table upload only ~10 MB, re-uploading an already-stored 8.43 GB blob only takes about 8 seconds. This change is the first time \"data lives on Hub, compute runs anywhere\" has been made a product-grade contract, breaking the long-standing deployment constraint that \"wherever the object storage is, GPU scheduling is pinned there\" — the data-movement cost of cross-cloud training and inference is compressed to near zero.","hf-skypilot-hf-storage","2026-07-12T06:30:00Z","2026-07-11T22:14:09.053006Z","2026-08-19T02:08:40.142862Z",true,"agent",157,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"f76e6f4b-1ae5-442a-abfa-823e7226ae81","Hugging Face 把 transformers 跑出 vLLM 原生速度：单 flag 让 235B MoE 直接吃到 EP 红利","hf-vllm-transformers-backend","2026-07-09T04:03:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"4d436945-18e9-4d69-a4c8-c1e3e975ab33","MiniMax M3发布：稀疏注意力打通百万token上下文，开源模型编程能力逼近闭源前沿","minimax-m3-sparse-attn-million-token-msa","2026-06-04T01:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"86c3465b-27f0-4f8c-83d1-9a8e1fa171c8","Compute Aligned Training：让模型学会协同推理的新训练范式","compute-aligned-training-pass-n-vote-michigan","2026-05-22T13:10:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"068bcc7a-d901-4d3d-924c-eefa7ced4467","TokenSpeed 开源推理引擎发布：剑指 Agentic Workloads 的高效推理","tokenspeed-agentic-inference-engine","2026-05-20T19:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"0565190a-0bcd-492f-934f-0ad2ab32f485","70万参数2.8MB填一张表:Cua开源CUA-S1,单次前向替代23轮LLM","cua-s1-forms-system-one-model","2026-09-20T13:11:48+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"b571067a-9fa8-42bf-9431-98f26ac78e03","伯克利把LLM推理搬进SSD:KV缓存压缩15倍","llm-inference-in-flash-cim-ssd","2026-09-19T21:10:00+00:00"]