← All IntelClip / OtherQLoRA: Fine-tuning on a 'Toaster' (Free-tier T4 GPU)
From Compression at the Edge — Chris Alexiuk, NVIDIA · ≈8:12
“we could actually fine-tune stuff on a toaster essentially”
“I could run open claw and hermes agent on like very long horizon tasks with Q1 3.6 quants which wouldn't be able to which wouldn't be possible before”
“it can do a lot of coding it can fix its own harness and stuff”
What’s in it
- Traces QLoRA's breakthrough: fine-tuning LLMs on a free Colab T4 GPU
- Details the llama.cpp acquisition and pivot to quantized local agents
- Shows Q1 3.5/3.6 quantized models handling long coding agent tasks
Clip transcript
that was a slowly growth for me. >> Should I? >> Yes please. So, um I think for me the biggest wow moment of quantization was back in the day there was something called Q Laura >> and and the fact that we could actually fine-tune stuff on a toaster essentially that is the collab tier T4 for me that was like a big wow although it's super slow it's it's okay um but on top of it um for instance like we have libraries called TRL and bits and bites and have that enables all of this and like you can even do this with the very advanced techniques like GRPO for instance actually you can train stuff on the um I mean you need to collocate on stuff but like with very small amounts of VRAM and then recently we uh in case you missed it we have acquired Lama CPPA kind of AUI hired llama CPP and then I kind of pivoted to the llama CPP world and I was like wow Because like the second wow moment for me was the fact that I could run open claw and hermes agent on like very long horizon tasks with Q1 3.6 quants which wouldn't be able to which wouldn't be possible before. Um I just know uh I tried many models actually with that and then Q1 3.5 was like wow it can do a lot of coding it can fix its own harness and stuff. So yeah, that was the two big moments for me actually.
Comments
Sign in to comment.
Loading comments…