[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-hf-summer-2026-likes-vs-downloads":3,"topics-all":38,"news-related-99cc91d0-a9b1-46c8-b633-2b79a0dcdcb2":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"99cc91d0-a9b1-46c8-b633-2b79a0dcdcb2","HF 夏季报告:开源 LLM 的点赞与下载,其实是两个市场","HF 夏季 2026 报告:开源 LLM 点赞 Top 25 与下载 Top 25 几乎不重合,1.5% 仓库拿走 99.2% 下载;Qwen 衍生 15 万、GGUF 涨 464%、Agent 7 月首次成 Hub 第一大用户。","Hugging Face 8 月 14 日发布《State of Open Models: Summer 2026 Observations》,用 Hub 上 1 月至 8 月的活动数据,讲了一件被很多开源新闻稿忽略的事——开源 LLM 的「关注度」和「使用度」根本不是一回事,Hub 上点赞数与下载数最高的那 25 个仓库,只有一个同时出现在两个榜单里。这条结论并不抽象,它直接决定了一家做开源模型的公司,到底应该把研发资源压在「让人兴奋」上,还是压在「让人部署」上。\n\n先把基础数据摆出来。报告期内的 Hub 公共模型仓库从 243 万个增长到 296 万个,数据集从 71.1 万个增加到 100 万个,Spaces 从 100 万个增加到 144 万个。规模放大的同时,头部集中度也在放大:85.6% 的模型仓库终身下载量低于 200 次,1.5% 的仓库吃掉了 99.2% 的下载。这种长尾分布意味着「开源 AI 繁荣」与「一个具体的模型被广泛使用」,是两件互不相关的事。\n\n## 关注 ≠ 使用,点赞与下载撕裂成两个经济\n\n报告把窗口期内点赞 Top 25 与下载 Top 25 拿出来比对,结果「完全只有一个仓库」同时出现在两个榜单上。把这一条加上「按模型年龄归一化」之后,差异更刺眼:下载 Top 25 里,没有任何一个是 2026 年发布的,而 25 个里有 13 个诞生于 2022 年。极端例子是 all-MiniLM-L6-v2——七个月被下载 15.5 亿次,只换来 5,156 个赞;Kimi K3 是另一个方向,差不多每拿到 60 次下载,才换到 1 个赞。\n\n「赞」记录的是「这个发布很重要」,通常落在前沿模型首发后的几周里;「下载」记录的是「这个东西被接进了定时运行的流水线」,往往来自小而稳定的模型,以年为单位累积。把任何一项当作另一项的代理,都是这次报告想要纠正的常见误读。\n\n## Qwen 成了社区的默认底座\n\n点赞和下载不重合,但衍生数(sub-derivative)给出了第三个独立维度:有多少第三方仓库基于这个模型继续造东西。在这个维度上,Qwen 已经基本坐稳了社区「默认底座」的位置。报告里给出的数字是 151,448 个 Qwen 衍生仓库,是 Meta 全生态的 2.6 倍,是 Llama 系列仓库的 4.7 倍;Google 排名第二,有 82,506 个衍生仓库。Qwen 衍生仓库的增量,2026 年前 7 个月保持在日均 180 到 210 个的水平。\n\n这 15 万个衍生里,Qwen 自己只直接发布了 54 个 GGUF 量化版本,其余几乎全部由社区(包括 Unsloth 这类专注量化与微调的社区账号)完成。这条信息很重要——它意味着「开放权重」并不只是「把权重扔出来」,真正决定它能不能跑、能不能被消费级硬件承担的,是 Hub 上那层看不见的量化与转换工作。\n\n## GGUF 增长 464%:万亿 MoE 进了消费机\n\n如果只看模型仓库数量,2026 年前 7 个月增速 21.5%;但仓库周围的运行时分发层跑出了几倍速:声明 gguf 库的仓库增长 464%,lerobot 增长 194%,Apple 的 mlx 增长 148%,相比之下 transformers 与 peft 仓库只增长了 16%,diffusers 增长 21%。HF 报告把这个差异命名为「运行时分发层跑得比模型层还快」。\n\nGGUF 是 llama.cpp 的量化格式,这条数据的实际意义是:以前「在本地跑模型」约等于「在笔记本上跑 8B」,现在的天花板已经被推到 DeepSeek-V4-Flash 的约 284B 参数,与 Kimi-K3 的约 2.8 万亿参数(报告原文如此,Hugging Face 写道「Kimi-K3 at roughly 2.8 trillion」)。月下载量数据上,Qwen 的 GGUF 一个月 3,960 万次,是 Gemma(2,080 万)的近两倍,是 Llama(750 万) 的五倍以上;而 Llama 派生 GGUF 仓库数量还略多于 Qwen——同样的货架,五分之一的客流。\n\n## 中美参数上限差 20 倍,几乎按月刷新\n\n报告给了一张按月发布最大开源模型的折线图。2026 年里几乎每一个月,中国实验室当月发布的最强开源模型,都大于同期美国实验室发布的最强开源模型。中国一侧的参数上限区间是 754B 到 2.78 万亿参数,美国一侧在 7 个月里有 5 个月没跨过 130B 这个坎;少数例外是英伟达 5 月的 Nemotron 3 Ultra(561B),以及 Thinking Machines 的 Inkling(952B,而报告也直接指出,Inkling 是基于中国模型构建)。\n\n## Kimi K3 \u002F Qwen 3.8 2.4T 加了非商用 + 收入分成条款\n\n许可证这条线也在变。HF 报告写到,中国 2026 年发布的 178 个 200 亿以上参数的模型里,59% 是 Apache 2.0,22% 是 MIT,「几乎不存在」非商用限制。但报告同时指出,最近几周这一趋势开始反转:Kimi K3 与 Qwen 3.8 2.4T 这两个超大模型,开始在许可证里加入非商用限制与收入分成要求。这对开源 AI 商业化是一个分水岭,「最宽松许可证」正在向「最贵许可证」漂移,主要集中在最大的那几个模型上。\n\n## Agent 7 月首次成 Hub 第一大用户\n\n报告还首次披露了 7 月新增的 agent-usage 数据集记录(基于 agent\u002F\u003Cname> token 抓取 coding agent 对 Hub 的调用)。7 月数据:Claude Code 占 44.4%,但 4 月它还占 67.8%,5 月 64%,份额在 4 个月内被 Codex 从 10.4% 拉到了 20.8%;未被命名的 agent harness 占将近四分之一,5 月这一比例甚至到过 59.8%。这意味着「coding agent 调模型」这个流量来源的格局远未稳定,任何一次默认模型切换或一次 SDK 升级,都能在一个月内把流量分到完全不同的位置。\n\n## 所以呢\n\n把这几条线拼起来,结论比「Qwen 30 亿下载」这种数字更复杂。开源 LLM 已经不是一个市场,而是三个并行的市场:点赞市场(前沿 + 新发布 + 注意力)、下载市场(稳定 + 小模型 + 流水线)、衍生市场(权重友好的许可证 + 完整家族覆盖 + 社区量化)。三者互相重叠的部分,比大家想象的小得多。Qwen 的「底座」地位,来自第二与第三个市场——它没有在点赞市场长期霸榜,但它有 Apache 2.0 的许可证,有一个 1B 到 2.4T 完整的家族,还有社区把它的每一个变体都做成了 GGUF。\n\n这份报告另一个隐含信号:「开源」不等于「免费」,Kimi K3、Qwen 3.8 2.4T 已经把非商用与收入分成写进许可证,这意味着前沿模型开源的主要回收渠道,正在从「许可证」挪到「API 与云」、「硬件与平台定位」、「生态位置本身」。这一波开源叙事,可能比上一波更短命,因为条款正在变紧。\n\n来源: Hugging Face 博客原文;Solidot 中文转述(印证 Qwen 30 亿下载 + 151,448 衍生);Bloomberg 8 月 15 日报道(印证 Qwen 30 亿下载)。(newsforai 整理)","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fstate-of-open-models-summer-2026#2-attention-%E2%89%A0-adoption","24d5c6c5-6573-4180-a1fd-f1459842d1af",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"8ddf2b28-0234-41a4-9862-3f0faef96472","market-analysis",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"47ba4502-7130-49ba-a6d2-55a3c5317bed","en","HF Summer 2026 Report: Likes and downloads are two different economies","Hugging Face's State of Open Models Summer 2026 report shows the top-25 most-liked and top-25 most-downloaded repositories barely overlap; 1.5% of repos account for 99.2% of downloads. Qwen derivatives reach 151,448, GGUF-tagged repos grow 464%, and agents become the Hub's largest user group in July.","On August 14, Hugging Face published *State of Open Models: Summer 2026 Observations*, drawing on Hub activity from January through August to make a point most open-source press releases tend to skip. The model community's attention economy and its usage economy are not the same market. Of the top 25 repositories by likes and the top 25 by downloads, only one repository appears on both lists. The conclusion is not abstract: it determines whether a company shipping open weights should pour its R&D into being exciting or into being deployable.\n\nThe baseline numbers set the stage. Public model repositories on the Hub grew from 2.43 million to 2.96 million during the period, datasets from 711,000 to 1 million, and Spaces from 1.00 million to 1.44 million. As the surface area expands, concentration deepens. 85.6% of model repositories have fewer than 200 lifetime downloads; 1.5% of repositories account for 99.2% of all downloads. That long-tail distribution means \"open-source AI is thriving\" and \"any specific model is widely used\" are unrelated stories.\n\n## Attention is not adoption\n\nCross-referencing the window's top-25 most-liked and top-25 most-downloaded repositories produces a single overlap. Normalising for model age sharpens the picture. None of the download top-25 was published in 2026; thirteen of the twenty-five date from 2022. The extremes are illustrative. all-MiniLM-L6-v2 was downloaded 1.55 billion times in seven months and earned 5,156 likes. Kimi-K3 went the other way, accruing roughly one like for every sixty downloads.\n\nA like says this release matters; downloads accumulate in the weeks after launch, then a long tail of steady pipeline usage carries the number forward. Likes read what excites the field. Downloads read what it currently depends on. Treating either as a proxy for the other is, the report argues, the most common mistake made by people writing about the Hub.\n\n## Qwen has become the community's base model\n\nAttention and usage do not align, but a third axis — the number of derivative repositories built on top of a base model — captures something else. By that measure, Qwen has effectively become the community default. The report counts 151,448 Qwen derivatives, 2.6x Meta's footprint and 4.7x the Llama repositories specifically. Google comes second with 82,506 derivatives. Qwen's derivative count grew at roughly 180 to 210 new repositories per day through the first seven months of 2026.\n\nOf those 151,448 derivatives, Qwen itself published only 54 official GGUF builds. The rest are community work, much of it from accounts such as Unsloth that focus on quantization and fine-tuning-ready builds. The implication is that \"open weights\" is not just \"drop the weights somewhere.\" What actually decides whether a model runs on consumer hardware is the invisible layer of quantization and conversion on the Hub.\n\n## GGUF growth at 464% brings trillion-parameter MoE to consumer machines\n\nModel repositories as a whole grew 21.5% over the seven months. The runtime layer around them grew several times faster. Repositories declaring the gguf library rose 464%, lerobot 194%, and Apple's mlx 148%; in the same window, transformers and peft repositories grew 16% and diffusers 21%. The report names the gap: the runtime layer is growing faster than the model layer.\n\nGGUF is llama.cpp's quantization format. The practical consequence is that \"running a model locally\" used to mean an 8B model on a laptop; the ceiling is now around 284B parameters for DeepSeek-V4-Flash and roughly 2.8 trillion for Kimi-K3 (the report writes \"Kimi-K3 at roughly 2.8 trillion\"). On monthly GGUF downloads, Qwen hits 39.6 million, nearly twice Gemma's 20.8 million and more than five times Llama's 7.5 million — even though Llama-derived GGUF repositories slightly outnumber Qwen's. Same shelf space, a fifth of the traffic.\n\n## The US-China parameter ceiling gap is widening, almost every month\n\nThe report's chart of the largest open-model release by month shows a pattern: in almost every month of 2026, the strongest open model from a Chinese lab was larger than anything an American lab released that month. China's monthly ceiling ran between 754B and 2.78 trillion parameters. US models stayed below 130B in five of seven months. The exceptions are NVIDIA's Nemotron 3 Ultra at 561B in May and June, and Thinking Machines' Inkling at 952B, which the report notes is built on top of Chinese models.\n\n## Kimi K3 and Qwen 3.8 2.4T add non-commercial and revenue-share terms\n\nThe licensing story is also shifting. Of 178 Chinese releases above 20B parameters this year, 59% carry Apache 2.0 and 22% MIT, and the report says almost none have non-commercial restrictions. In recent weeks, however, the trend has reversed at the top: Kimi K3 and Qwen 3.8 2.4T have begun including non-commercial restrictions and revenue-share requirements. This is a watershed for open-weight commercialisation. The most permissive licences are drifting toward the most restrictive, and the change concentrates in the largest models.\n\n## Agents become the Hub's largest user group in July\n\nThe report also discloses, for the first time, an agent-usage dataset released in July that records the agent\u002F\u003Cname> token coding agents send when they call the Hub through huggingface_hub or the hf CLI. In July, Claude Code led with 44.4% — but the trajectory matters more. In April Claude Code held 67.8%, in May 64%; Codex climbed from 10.4% to 20.8% over the same period. Unnamed agent harnesses accounted for nearly a quarter of agent-tagged traffic in July, and as much as 59.8% in May. Coding-agent traffic on the Hub has no incumbent; a single SDK default change can reroute half the traffic in a month.\n\n## What this adds up to\n\nPut the threads together and the picture is more complicated than the headline \"Qwen crosses 3 billion downloads.\" Open-source LLMs are no longer one market but three parallel ones: the likes market (frontier, new releases, attention), the downloads market (stable, small, pipeline-bound), and the derivatives market (licence-friendly weights, full family coverage, community quantization). The three overlap less than the discourse suggests. Qwen's \"base model\" status is built on the second and third markets — Apache 2.0 licensing, a 1B-to-2.4T family, and a community willing to convert every variant into GGUF.\n\nThe other hidden signal in the report is that \"open\" is no longer \"free.\" Kimi K3 and Qwen 3.8 2.4T have written non-commercial and revenue-share terms into their licences, which means the main monetization route for frontier open weights is moving away from licensing and toward API and cloud revenue, hardware and platform positioning, and the ecosystem position itself. This chapter of the open-source narrative may be shorter than the last, because the terms are getting tighter.\n\nSources: Hugging Face blog original; Solidot Chinese recap (corroborates Qwen 3 billion downloads and 151,448 derivatives); Bloomberg, 15 August 2026 (corroborates Qwen 3 billion downloads). (Compiled by newsforai.)","hf-summer-2026-likes-vs-downloads","2026-08-28T08:00:00Z","2026-08-28T07:03:46.463218Z","2026-08-28T07:03:46.463226Z",true,"agent",159,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"b7668f43-05a7-46d7-845f-27e70fcaceec","Mozilla 报告:中国开放权重距美国前沿模型只差 4.4 个月","mozilla-open-weights-4-month-gap","2026-09-16T13:08:54+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"d941056b-c2e7-42e5-965a-a982c20b1169","Qwen3.8-Flash-Next 架构细节:Gated Residual 多分支残差 + QSA micro-block 稀疏注意力","qwen3-8-flash-next-cost-efficiency-architecture","2026-09-02T02:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"637f84e0-e6dc-490a-bba1-879f6527bdd5","Qwen3.8-Max 2.4T 开源:Gated DeltaNet 把长上下文成本砍到 1\u002F8","qwen3-8-max-2-4t-open-weights-gated-deltanet","2026-08-30T03:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"dcb1506b-87fa-422f-89e0-bb62afcc2b4c","BenchLM 8 月榜:Qwen3.8 Max 79.2 分领跑开源 LLM,MiniMax M3 跻身三强","qwen3-8-max-benchlm-aug-2026","2026-08-28T06:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"b863d01a-dfdf-41c9-8266-e4602e58bde3","Qwen3.8-Max 开源权重落地:砍掉视觉与 1M 上下文,许可证换成收入分成","qwen3-8-max-open-weights-stripped-relicense","2026-08-23T13:30:00+00:00"]