
🤔Can a lightweight model not only run locally, but also train locally? With AReno @ARenoTeam, we post-trained @AntLingAGI Ling-3.0-tiny on DGX Spark using an Agentic RL tic-tac-toe task. The model learned from tool calls, environment feedback, and rewards: ✅rewards_mean: ~-0.5 → ~0.4 ✅response_len dropped to ~850 tokens ✅behavior became more stable after training #inclusionAI #OpenSource #ReinforcementLearning From running, to training, to adapting to your own task. Check out Areno https://t.co/mDdE10cfco!
Agentic RL post-training normally assumes cluster access. This run shows a lightweight Ling model trained on tool calls and environment rewards on a single DGX Spark, with mean reward moving from about -0.5 to 0.4.
Checking sign-in…
Loading comments…