Vibeleaderboard
← All Intel
Intel / repo

LoRA over GGUF – Train Qwen3.8-Flash-Next in 40G VRAM

Source
github.com
Author
woctordho
Date
Why it matters

Shows local fine-tuning of large MoE models is feasible on a single high-memory machine using GGUF , with measured memory and throughput figures.

Key takeaways · AI-distilled
  • The usual local QLoRA stack (Transformers with a bitsandbytes 4-bit base) stalled on models because bitsandbytes does not yet support MoE. The author says Transformers 5.18 added GGUF loading that covers recent MoE and sparse- models.
  • The author reports training Qwen3.8-Flash-Next (125B-A6B plus a 51B engram) in 40 GiB of VRAM without CPU offload, at 200 tokens/s on AMD Strix Halo, and DeepSeek-V4-Flash (284B-A13B) in 90 GiB at 100 tokens/s.
  • There is headroom: the same setup reaches over 1,600 tokens/s on prompt processing, and training with gradient checkpointing usually costs 4-5x prompt-processing work, so 200 tokens/s is below the expected ceiling. CPU/disk offload and multi-GPU are still unoptimized.
Terms in this piece · Glossary
  • mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
  • quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
  • attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
  • LoRA — A cheap way to fine-tune a model by training a small add-on layer instead of changing all of the model's weights.
Recommended reads
Comments

Checking sign-in…

Loading comments…