Today, we’re releasing the first weights from Trinity Large, our first frontier-scale model in the Trinity MoE family.
We’re releasing three variants from the same training run: - Trinity-Large-Preview: lightly post-trained, chat-ready - Trinity-Large-Base: full 17T token pretraining checkpoint - Trinity-Large-TrueBase: 10T token checkpoint with no instruct data or LR anneals
Sparsity was a core design choice. By keeping the routing fraction low (4-of-256), Trinity Large achieves significantly faster inference than dense models in its weight class, while maintaining strong performance.
Three checkpoints from a single 17T- run, including a rare base model with no instruct data or LR anneals, plus a published sparsity and stability recipe for training at this scale.
Checking sign-in…
Loading comments…