
Post-training is becoming something you delegate to an rather than babysit.
“However, since SFT updates all model weights, it requires extensive computing resources and may show regression on general knowledge.”
Tanya Lenz
“For the example detailed in this post, LoRA post-training required ~7x fewer GPU hours compared to full-parameter SFT for Cosmos 3 Nano, making a one-day post-training turnaround a reality for engineering teams.”
Tanya Lenz
“In just one run, and around 30 minutes of training time on eight NVIDIA A100 Tensor Core GPUs, the model jumps to 87.14% accuracy, a massive improvement of 32 percentage points, achieved completely hands-free.”
Tanya Lenz
“In this experiment environment, it took 19.5 hours to process the 43 parallel trials, running simultaneously across five parallel nodes (allocating one 8x A100 GPU node per search strategy).”
Tanya Lenz
“While a Llama NIM can serve a base model alongside standalone LoRA adapter directories, the Cosmos 3 Reasoner NIM expects a merged checkpoint (base weights fused with the LoRA adapter) rather than a bare adapter.”
Tanya Lenz
articleHow to Use AI Agents to Prepare 3D Scenes for Simulation
articleTranslating CUDA Tile Operations from Python to Rust Using Agentic AI
articleHow NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin
articleAccelerating Dropless MoE Training in JAX with NVIDIA Transformer EngineChecking sign-in…
Loading comments…