
Shows NVFP4 combined with can speed on-device reasoning-model by up to 6.28x on Jetson, with the optimal decoding method differing per model, concrete guidance for deploying agentic models at the edge.
articleDeveloping Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model OptimizerTanya Lenz
postNVIDIA has just released the first Nemotron 3.5 model: Nemotron 3.5 Lightning,…Artificial Analysis
blogNVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AIKari BriskiChecking sign-in…
Loading comments…