← All IntelClip / OtherNVIDIA's Flex Run: distillation and flexible weight activation
From The State of Model Routing — NVIDIA, Cognition, OpenRouter · ≈20:12
“So we have a technology called Flex Run.”
“Most In most cases, you can essentially understand the novelty of a question to a model if you have access to the recipe with which it was trained.”
What’s in it
- Explains Nvidia's Flex Run tech for dynamic model-size switching
- Shows how distillation creates smaller models for different tasks
- Introduces the 'distillation gap' concept across training domains
Clip transcript
training at at Nvidia for this kind of purposes? >> Yeah, so we have a technology called >> The mic's go? >> Right. >> Hello. >> Testing. Oh, this one works. >> [laughter] >> Okay. Uh so we have a technology called Flex Run. So you have uh we we have a setup where uh there's a the uh there's a main model, then we distill it into smaller uh footprints. And then based on the based on the task at hand, you can switch which model does the decoding. Right? So there's a lot of fancy stuff you can do uh within a model artifact, too. Uh to essentially only activate a class of model or a section of weights, depending on the task at hand or the complexity at hand. Most In most cases, you can essentially understand the novelty of a question to a model if you have access to the recipe with which it was trained. So, this works very well for open models, right? Like or any model you have access to its data for, right? Because you can literally decide if it's in like see if it's in distribution or not. Uh again, if you have studies from when it was trained, you can also see how much essentially how much was your distillation gap across teachers and the artifact that you trained, right? Because sure, you have domain data from all all the different domains you're tuning, but it's not guaranteed that it uh absorbed all the that data evenly across the models, right? So, um it becomes it becomes very interesting uh to start thinking about these flexible weights and flexible model sizes
Comments
Checking sign-in…
Loading comments…