Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer
Source
Tanya Lenz
Author
Tanya Lenz
Date
Terms in this piece · Glossary
quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
distillation — Training a small, cheap model to imitate a big one's outputs, keeping much of the capability at a fraction of the cost.
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Recovering agentic and coding benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → accuracy at four-bit weights changes what a team can realistically self-host on one node instead of renting.