An agentic RL post-training loop that fits on one DGX Spark
- Source
- Ant Ling
- Date
7.9B total. 1.3B active. One DGX Spark. No cluster. Ling-3.0-tiny × ASystem AReno closes the Agentic RL post-training loop—locally, on a single DGX Spark. Running locally is only step one. Ling-3.0-tiny makes iteration practical; ASystem AReno turns task feedback into a trainable, retestable loop on one node.

Tic-tac-toe is the minimal task for validating an end-to-end Agentic RL loop. Ling-3.0-tiny is post-trained with ASystem AReno on one DGX Spark. Same model. Same environment. Before vs. after post-training. Result: more stable tool calls and better move selection.
On a single DGX Spark, the ASystem AReno team post-trained Ling-3.0-tiny with GSPO for 400 steps. • rollout/rewards_mean: ~-0.5 → ~0.4 • response_len: ↓ to ~850 tokens One curve tracks stronger task feedback. The other shows shorter, more focused responses. Together, they explain the shift seen in the before/after demo.


The asset is not tic-tac-toe. It is the loop: Baseline → feedback → post-train → same-environment retest. Use it for repeated, scoreable work: tool calls, structured extraction, domain instructions, and workflow Agents. Task adaptation—not a broad capability claim.
Context
Ant's Ling team reports that it post-trained Ling-3.0-tiny, a model with 7.9 billion total and 1.3 billion active parameters, on a single DGX Spark with no cluster, using its ASystem AReno framework. The loop is a baseline run, feedback on the task, post-training with a method called GSPO, and a retest of the same model in the same environment.
Ant chose tic-tac-toe as 'the minimal task for validating an end-to-end Agentic RL loop.' Over 400 GSPO steps it reports mean reward moving from about -0.5 to about 0.4 while response length fell to about 850 . Ant says the two curves together explain the before and after demo, which it describes as more stable tool calls and better move selection.
Ant says the asset is the loop rather than tic-tac-toe, and suggests it for repeated, scoreable work such as tool calls or structured extraction. It calls the result task adaptation, not a broad capability claim. The curves are Ant's own and have not been independently replicated.
- token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Checking sign-in…
Loading comments…







