
Last week, we launched AFM-4.5B, our first foundation model. In this post by @chargoddard , you will learn how we extended the context length of AFM-4.5B from 4k to 64k context through aggressive experimentation, model merging, distillation, and a concerning amount of soup. Bon appétit 😋 Blog post:

Extending by merging and averaging checkpoints is far cheaper than retraining for long context, and this documents what the combination actually achieved on a real 4.5B model.
Checking sign-in…
Loading comments…