Six untouched Ling-3.0 base checkpoints released for continued pretraining
- Source
- Ant Ling
- Date
🧵 We’ve open-sourced 6 Base Model checkpoints for Ling-3.0-tiny & Ling-3.0-flash, covering pre-trained, mid-trained, and WSM-merged stages. None has undergone post-training, giving researchers flexible starting points for continued pre-training, fine-tuning, and further research. Two key highlights: - We use WSM to replace LR decay with weighted checkpoint merging, making the training process better suited for continual pre-training while enabling offline exploration of different LR decay strategies. - With one shared training recipe, the community can validate strategies on tiny-base, then scale them to flash-base.

🐣 Ling-3.0-tiny-base: 7.9B total | 1.3B active. Despite having only half as many total parameters as Ling-2.5-mini-base, Ling-3.0-tiny-base delivers comparable or superior performance on most benchmarks, with particularly strong results in coding. For code pre-training/SFT, RL post-training, teaching, model behavior & MoE studies.

⚡ Ling-3.0-flash-base: 124B total | 5.1B active. In our evaluations, Ling-3.0-flash-base achieves strong performance across coding, reasoning, and long context tasks, even when compared with models 2 to 3 times larger. This makes it well suited for continued pre-training, post training, and domain adaptation in coding, long horizon workflows, finance, healthcare, and other specialized applications.

📥 Ling-3.0-tiny-base Hugging Face: https://t.co/nSochw3F28 https://t.co/s51O7UloGZ https://t.co/rI0mv4D5sI ModelScope: https://t.co/uDtgylFqPg https://t.co/cXsHYv8rdR https://t.co/PCV2nG3Cw4 📥 Ling-3.0-flash-base Hugging Face: https://t.co/PkwesYccCV https://t.co/fUrS67UOet https://t.co/iKimzrJREu ModelScope: https://t.co/bIQJ9qdSpk https://t.co/kQ61kshYd7 https://t.co/guhrmegizt
Context
Ant's Ling team says it open-sourced six base checkpoints for Ling-3.0-tiny (7.9 billion total parameters, 1.3 billion active) and Ling-3.0-flash (124 billion total, 5.1 billion active), covering pre-trained, mid-trained, and WSM-merged stages. None has gone through post-training, the later tuning step, so Ant offers them as starting points for continued , , and research.
WSM, Ant says, replaces learning-rate decay (a gradual lowering of the training step size) with weighted merging of checkpoints. Ant says this makes training better suited to continual pretraining and allows offline exploration of different decay strategies. It also says one shared recipe lets the community validate strategies on the tiny base and then scale them to the flash base. Those are Ant's claims, and this explanation did not test whether the transfer holds.
- pretraining — The first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.
- fine-tuning — Taking a trained model and training it a bit more on your own examples so it gets better at one specific job.
Checking sign-in…
Loading comments…






