Clip transcript
or a 256 gig Mac Studio, things like that. When quantized down to, I think, a 4-bit quant, this should fit pretty comfortably on those listed devices. So, still something that's going to be a bit heavy to run, but at 299 billion parameters, being an MoE, this is actually a reasonable size model, especially when factoring in the touted performance here in the benchmarks. Now, before we get into it, please do feel free to subscribe, so I can get the 100k plaque. Also, I do want to specifically mention that this is licensed under Apache 2.0. Just in reading through some discussions I've seen online, it seems like the preview version wasn't, but this is now a fully, like, very solid license to have. Additionally to that, they do talk a bit about fine-tuning this model, so this may be a pretty potent option for folks who are interested in actually modifying it to a specific use case, and they do have the fine-tuning guide for that linked right here, which is just good to see. I always like to see additional things beyond models being open-source tools to actually further tailor them to your specific use case. And, being that the current point in time seems to have a little more curiosity about access to frontier state-of-the-art models, it's always good to see sizes like this, which can hypothetically be run on some hobbyist prosumer level setups. So, let's take a look at the technical specs of this model. It is a 295 billion parameter mixture of experts model with 21 billion active. It has a 3.8 B MTP layer parameters. So, basically it has multi-token prediction. So, it will run a bit quicker than the size would be indicative of thanks to that MTP. Additionally, they talk a lot about the improvements over the preview version stating that it outperforms similar size models and rivals flagship open-source models with two to five times the parameters. I did see just on X some comparisons of this to GLM 5.1 or 5.2 and folks were saying based on those benchmarks it does seem like it may trade blows with that, which would be really exciting. Context length is listed right here at 256K and there's some other more technical information about the model's architecture and things like that. Now,