← All IntelIntel / repo
LoRA over GGUF – Train Qwen3.8-Flash-Next in 40G VRAM
- Source
- github.com
- Author
- woctordho
- Date
Why it matters
Shows local fine-tuning of large MoE models is feasible on a single high-memory machine using GGUF , with measured memory and throughput figures.
Key takeaways · AI-distilled
- The usual local QLoRA stack (Transformers with a bitsandbytes 4-bit base) stalled on models because bitsandbytes does not yet support MoE. The author says Transformers 5.18 added GGUF loading that covers recent MoE and sparse- models.
- The author reports training Qwen3.8-Flash-Next (125B-A6B plus a 51B engram) in 40 GiB of VRAM without CPU offload, at 200 tokens/s on AMD Strix Halo, and DeepSeek-V4-Flash (284B-A13B) in 90 GiB at 100 tokens/s.
- There is headroom: the same setup reaches over 1,600 tokens/s on prompt processing, and training with gradient checkpointing usually costs 4-5x prompt-processing work, so 200 tokens/s is below the expected ceiling. CPU/disk offload and multi-GPU are still unoptimized.
Terms in this piece · Glossary
- mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
- quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
- attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
- LoRA — A cheap way to fine-tune a model by training a small add-on layer instead of changing all of the model's weights.
Read the source github.com
Recommended reads
Comments
Checking sign-in…
Loading comments…

