[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-lora-over-gguf-low-vram-training":3,"topics-all":38,"news-related-4ccc491b-dbee-4c84-beb6-7cf519f76320":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"4ccc491b-dbee-4c84-beb6-7cf519f76320","LoRA 基座换 GGUF:40G 显存训 125B","bnb 不支持 MoE,近几代开源模型的本地训练停滞。woct0rdho 开源训练配方,把低显存 LoRA 基座换成 GGUF:Strix Halo 上 40 GiB 显存训 Qwen3.8-Flash-Next 125B,200 token\u002Fs;90 GiB 可训 284B DeepSeek-V4-Flash。","bitsandbytes 的 4-bit 量化曾是本地 LoRA 微调的默认底座,但它至今不支持 MoE——而近几代新开源模型大量转向了 MoE。10 月 10 日,开发者 woct0rdho 在 Hacker News 放出 transformers5-qwen3.5-recipe,解法干脆:基座不用 bnb,直接用 GGUF。\n\n## 卡点:bnb 的 MoE 空窗\n\n过去几年本地训练的通行方案是 QLoRA:bnb 4-bit 基座加 LoRA 适配器,HuggingFace Transformers 做底座,Unsloth、Axolotl 在其上封装。bnb 迟迟没有 MoE 支持,新模型的本地训练因此停滞。另一边,GGUF 是 llama.cpp 生态的容器格式,4 bpw 以下量化质量依然可用,还装得下 MoE 与稀疏注意力,只是从来没人把它当训练基座。Transformers 5.18 开始原生支持 GGUF,作者实测在 Mac 上已快过 llama.cpp——这扇门才算打开。\n\n## GGUF 上位:一整套 kernel 栈\n\n这个 repo 给的是一整条训练链路:GGUF 反量化函数过 torch.compile 省显存;仿 llama.cpp 的 MMQ 与分组 MMQ;线性层和 MoE 层的快 LoRA 反传公式;Liger 的 RMSNorm 与 chunked cross entropy;8-bit AdamW。模型侧适配覆盖 DeepSeek 的滑动注意力、CSA 与 HCA,以及 GatedDeltaNet 的 Triton 前反向 kernel。llama.cpp 目前没有快 LoRA kernel,作者顺手写了一个提交上去。\n\n## 显存账本:16 GiB 到 192 GiB\n\n数字谱系(全部无 CPU offload):\n\n- Qwen3.6-35B-A3B:16 GiB(APEX-I-Mini 量化占 13.3 GiB),推得 122B-A10B 约 64 GiB、397B-A17B 约 192 GiB\n- Qwen3.8-Flash-Next(125B-A6B 加 51B engram):40 GiB,GSQ-RCO Q2_0 量化 35 GiB\n- DeepSeek-V4-Flash(284B-A13B):90 GiB,IQ2_XXS 占 81 GiB\n\n在 Strix Halo 统一内存上,Qwen3.8-Flash-Next 训练 200 token\u002Fs,prompt 处理超 1600 token\u002Fs;DeepSeek-V4-Flash 也能 100 token\u002Fs 训练,但作者认为本地场景不如前者实用。\n\n## 冷水与边界\n\n目前配方里全部 kernel 参数都为 Strix Halo 调优,其他 GPU 还需适配;GGUF 支持仍在向 transformers 主线合并的 issue 里,合不进去就以 monkey patch 留在 repo。「GGUF 将取代 bnb 成为低显存 LoRA 基座」是方向性预测,不是既成事实。repo 以 Apache-2.0 开源,22 star、7 fork,很早期。\n\n对开源权重社区,意义比数字大:open weights 的完整闭环不只是「能跑」,还包括「能改」。bnb 的 MoE 空窗说明把训练基座押注单一库有风险,GGUF 路径则证明推理生态的量化工艺可以反向输入训练侧。下次看到 125B 开放权重,先看显存再看教程——自己微调,可能比想象中近。\n\n参考:[transformers5-qwen3.5-recipe](https:\u002F\u002Fgithub.com\u002Fwoct0rdho\u002Ftransformers5-qwen3.5-recipe)","https:\u002F\u002Fgithub.com\u002Fwoct0rdho\u002Ftransformers5-qwen3.5-recipe","998df6db-96e6-4b8e-8be1-cfa00a6cd177",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"a5c2f096-5e05-4633-88f3-b4c598aea9ea","en","LoRA over GGUF: Train a 125B MoE in 40 GiB VRAM","An open-source recipe runs LoRA over GGUF bases, bypassing bitsandbytes' MoE gap: Qwen3.8-Flash-Next 125B trains in 40 GiB VRAM at 200 tok\u002Fs locally.","bitsandbytes' 4-bit quantization used to be the default base for local LoRA fine-tuning, but it still has no MoE support — and recent open-weight models have largely moved to MoE. That mismatch froze \"fine-tune a large model on your own GPU\" for a while. On October 10, developer woct0rdho posted transformers5-qwen3.5-recipe on Hacker News with a blunt fix: drop bnb as the base and use GGUF directly.\n\n## The bottleneck: bnb's MoE gap\n\nFor years the standard local setup was QLoRA — a bitsandbytes 4-bit base with LoRA adapters, built on HuggingFace Transformers and wrapped by frameworks like Unsloth and Axolotl. But bnb never shipped MoE support, so local training of recent MoE models stalled. GGUF, meanwhile, is the llama.cpp ecosystem's container format: it stays usable below 4 bpw with surprisingly good quantization quality, holds MoE and sparse-attention models, and — since Transformers 5.18 — loads natively, which the author reports already runs faster than llama.cpp on Mac. Nobody had treated it as a training base until now.\n\n## GGUF steps up: a full kernel stack\n\nThe repo is not a config file but an entire training path: the GGUF dequantization function passed through torch.compile to save VRAM; llama.cpp-style MMQ and grouped MMQ; a fast LoRA backward formula for linear and MoE layers; Liger's RMSNorm and chunked cross-entropy; an 8-bit AdamW optimizer; non-reentrant gradient checkpointing. Model-specific adaptations include DeepSeek's sliding attention and CSA\u002FHCA, plus GatedDeltaNet Triton kernels with backward passes. Because llama.cpp has no fast LoRA kernel today, the author wrote one and submitted it upstream.\n\n## The VRAM ledger: 16 GiB to 192 GiB\n\nThe numbers, all without CPU offload:\n\n- Qwen3.6-35B-A3B: 16 GiB, of which the APEX-I-Mini quantization takes just 13.3 GiB — implying roughly 64 GiB for 122B-A10B and 192 GiB for 397B-A17B\n- Qwen3.8-Flash-Next (125B-A6B plus a 51B engram): 40 GiB, with GSQ-RCO Q2_0 quantization at 35 GiB and 27 GiB of engram\n- DeepSeek-V4-Flash (284B-A13B): 90 GiB, of which the IQ2_XXS quantization takes 81 GiB\n\nOn Strix Halo's unified memory, Qwen3.8-Flash-Next trains at 200 token\u002Fs against prompt processing above 1600 token\u002Fs; DeepSeek-V4-Flash also trains at 100 token\u002Fs, though the author considers it less practical than the Qwen model for local use.\n\n## Caveats and boundaries\n\nEvery kernel parameter in the recipe is currently tuned for Strix Halo; other GPUs still need adaptation work. The GGUF support remains tracked in an open issue toward the transformers mainline — the author says that if it cannot merge, it will live on as monkey patches in the repo. And his framing, \"GGUF is going to replace bitsandbytes as the base model format for low-VRAM LoRA training\", is a directional prediction, not an accomplished fact. The repo is Apache-2.0, with 22 stars and 7 forks: early days.\n\nFor the open-weights community the significance goes beyond the numbers. The full loop of open weights is not just \"can run\" but \"can modify\". bnb's MoE gap shows the risk of betting the training base on a single library, while this GGUF path proves that quantization craft from the inference ecosystem can flow back into training. Next time a 125B open-weight model drops, check your VRAM before the tutorials — fine-tuning it yourself may be closer than you think.\n\nReference: [transformers5-qwen3.5-recipe](https:\u002F\u002Fgithub.com\u002Fwoct0rdho\u002Ftransformers5-qwen3.5-recipe)","lora-over-gguf-low-vram-training","2026-10-10T21:08:22Z","2026-10-10T21:09:03.049730Z","2026-10-10T21:09:03.049756Z",true,"agent",50,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"031e715e-c9d6-4855-83da-0515f33f0e3c","POCKET：35B MoE 1-bit 跑进 iPhone，27 tok\u002Fs","pocket-35b-moe-iphone-edge","2026-07-28T04:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"530e5aa5-2026-4c30-b4e6-421caca907b2","Transformer 提前罢工:13 个基座模型跟不住引用链,一个 rank-8 LoRA 修好","tiny-lora-frozen-transformer-chain-relay","2026-10-03T21:05:16+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"474e602e-a509-4973-8d44-eafc532138a4","GLM-5.3-Flash 一个月:Wagtail 半程失守","wagtail-glm-5-3-flash-month","2026-10-03T17:08:59+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"32aa44d5-81f1-4717-8d17-ff713101b725","Phonon-2:2.1比特量化ASR,164MB逼近全精度","phonon-2-2-bit-parakeet-asr","2026-10-03T13:10:31+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"9a77e0d7-8164-494e-984d-54cb57a7a0dd","Magnitude 开源:本地 Agent 专用推理引擎","magnitude-self-tuning-local-agent-inference-engine","2026-10-01T17:08:35+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"c23c2fca-29cc-40e8-9ff2-aebb32aba767","3.21 比特量化:27B 模型从 54GB 压到 12.3GB","orcasaq2-3-bit-qwen-quantization","2026-09-30T21:20:00+00:00"]