🧵 We’ve open-sourced 6 Base Model checkpoints for Ling-3.0-tiny & Ling-3.0-flash, covering pre-trained, mid-trained, and WSM-merged stages. None has undergone post-training, giving researchers flexible starting points for continued pre-training, fine-tuning, and further research. Two key highlights: - We use WSM to replace LR decay with weighted checkpoint merging, making the training process better suited for continual pre-training while enabling offline exploration of different LR decay strategies. - With one shared training recipe, the community can validate strategies on tiny-base, then scale them to flash-base.

🐣 Ling-3.0-tiny-base: 7.9B total | 1.3B active. Despite having only half as many total parameters as Ling-2.5-mini-base, Ling-3.0-tiny-base delivers comparable or superior performance on most benchmarks, with particularly strong results in coding. For code pre-training/SFT, RL post-training, teaching, model behavior & MoE studies.

⚡ Ling-3.0-flash-base: 124B total | 5.1B active. In our evaluations, Ling-3.0-flash-base achieves strong performance across coding, reasoning, and long context tasks, even when compared with models 2 to 3 times larger. This makes it well suited for continued pre-training, post training, and domain adaptation in coding, long horizon workflows, finance, healthcare, and other specialized applications.

📥 Ling-3.0-tiny-base Hugging Face: https://t.co/nSochw3F28 https://t.co/s51O7UloGZ https://t.co/rI0mv4D5sI ModelScope: https://t.co/uDtgylFqPg https://t.co/cXsHYv8rdR https://t.co/PCV2nG3Cw4 📥 Ling-3.0-flash-base Hugging Face: https://t.co/PkwesYccCV https://t.co/fUrS67UOet https://t.co/iKimzrJREu ModelScope: https://t.co/bIQJ9qdSpk https://t.co/kQ61kshYd7 https://t.co/guhrmegizt
Post-trained weights are a poor starting point for domain adaptation. These are untouched bases at three training stages, and one shared recipe lets you validate an approach on the 7.9B model before committing to the 124B one.
Checking sign-in…
Loading comments…