7.9B total. 1.3B active. One DGX Spark. No cluster. Ling-3.0-tiny × ASystem AReno closes the Agentic RL post-training loop—locally, on a single DGX Spark. Running locally is only step one. Ling-3.0-tiny makes iteration practical; ASystem AReno turns task feedback into a trainable, retestable loop on one node.

Tic-tac-toe is the minimal task for validating an end-to-end Agentic RL loop. Ling-3.0-tiny is post-trained with ASystem AReno on one DGX Spark. Same model. Same environment. Before vs. after post-training. Result: more stable tool calls and better move selection.
On a single DGX Spark, the ASystem AReno team post-trained Ling-3.0-tiny with GSPO for 400 steps. • rollout/rewards_mean: ~-0.5 → ~0.4 • response_len: ↓ to ~850 tokens One curve tracks stronger task feedback. The other shows shorter, more focused responses. Together, they explain the shift seen in the before/after demo.


The asset is not tic-tac-toe. It is the loop: Baseline → feedback → post-train → same-environment retest. Use it for repeated, scoreable work: tool calls, structured extraction, domain instructions, and workflow Agents. Task adaptation—not a broad capability claim.
Agentic RL usually implies a cluster. This closes the baseline, feedback, post-train and retest loop on one desktop-class box with a 7.9B model, and both the tutorial and the framework are public.
Checking sign-in…
Loading comments…