bitsandbytes' 4-bit quantization used to be the default base for local LoRA fine-tuning, but it still has no MoE support — and recent open-weight models have largely moved to MoE. That mismatch froze "fine-tune a large model on your own GPU" for a while. On October 10, developer woct0rdho posted transformers5-qwen3.5-recipe on Hacker News with a blunt fix: drop bnb as the base and use GGUF directly.

The bottleneck: bnb's MoE gap

For years the standard local setup was QLoRA — a bitsandbytes 4-bit base with LoRA adapters, built on HuggingFace Transformers and wrapped by frameworks like Unsloth and Axolotl. But bnb never shipped MoE support, so local training of recent MoE models stalled. GGUF, meanwhile, is the llama.cpp ecosystem's container format: it stays usable below 4 bpw with surprisingly good quantization quality, holds MoE and sparse-attention models, and — since Transformers 5.18 — loads natively, which the author reports already runs faster than llama.cpp on Mac. Nobody had treated it as a training base until now.

GGUF steps up: a full kernel stack

The repo is not a config file but an entire training path: the GGUF dequantization function passed through torch.compile to save VRAM; llama.cpp-style MMQ and grouped MMQ; a fast LoRA backward formula for linear and MoE layers; Liger's RMSNorm and chunked cross-entropy; an 8-bit AdamW optimizer; non-reentrant gradient checkpointing. Model-specific adaptations include DeepSeek's sliding attention and CSA/HCA, plus GatedDeltaNet Triton kernels with backward passes. Because llama.cpp has no fast LoRA kernel today, the author wrote one and submitted it upstream.

The VRAM ledger: 16 GiB to 192 GiB

The numbers, all without CPU offload:

  • Qwen3.6-35B-A3B: 16 GiB, of which the APEX-I-Mini quantization takes just 13.3 GiB — implying roughly 64 GiB for 122B-A10B and 192 GiB for 397B-A17B
  • Qwen3.8-Flash-Next (125B-A6B plus a 51B engram): 40 GiB, with GSQ-RCO Q2_0 quantization at 35 GiB and 27 GiB of engram
  • DeepSeek-V4-Flash (284B-A13B): 90 GiB, of which the IQ2_XXS quantization takes 81 GiB

On Strix Halo's unified memory, Qwen3.8-Flash-Next trains at 200 token/s against prompt processing above 1600 token/s; DeepSeek-V4-Flash also trains at 100 token/s, though the author considers it less practical than the Qwen model for local use.

Caveats and boundaries

Every kernel parameter in the recipe is currently tuned for Strix Halo; other GPUs still need adaptation work. The GGUF support remains tracked in an open issue toward the transformers mainline — the author says that if it cannot merge, it will live on as monkey patches in the repo. And his framing, "GGUF is going to replace bitsandbytes as the base model format for low-VRAM LoRA training", is a directional prediction, not an accomplished fact. The repo is Apache-2.0, with 22 stars and 7 forks: early days.

For the open-weights community the significance goes beyond the numbers. The full loop of open weights is not just "can run" but "can modify". bnb's MoE gap shows the risk of betting the training base on a single library, while this GGUF path proves that quantization craft from the inference ecosystem can flow back into training. Next time a 125B open-weight model drops, check your VRAM before the tutorials — fine-tuning it yourself may be closer than you think.

Reference: transformers5-qwen3.5-recipe