Clip transcript
>> So, Alibaba has released Quen 3.8 Max. It was in preview for a couple of weeks prior to today, but now we have the official release here alongside some very exciting news for folks who are interested in openweight models because this model at 2.4 trillion parameters is going to be open weighted with the weights released next week. However, the exciting part because face it, none of us are going to be able to run this locally at home. At least 95% of us. However, they did also mention in the X announcement post here that Quen 3.827B is also going to be released in open weights probably next week as well. So, that model is pretty much the predecessor to Quen 3.827B was Quen 3.627B 627B and it is still probably best regarded as like the overall finest local model because its intelligence is fantastic and its size allows it to actually be realistically run on a wide variety of different hardware. So that is probably the most exciting bit of news for anyone interested in local and open- source AI. However, for today's video, we're going to be focused on the big bad Max Quen, which is the first time that a Quen Max model has been open weighted. So this is pretty exciting. So before we get into it, please do feel free to subscribe so I can get the 100k plaque and let's take a peek at some of the interesting things about this model. So size, it is 2.4 trillion parameters. It is a mixture of experts model with 95 billion active parameters. As we had mentioned, the open weights release is going to be next week. And they have some benchmarks right here. However, I noticed that it does not show the deep SWE benchmark score in this big benchmark JPEG. It is being compared both to its predecessor which is Quen 3.7 Max which was significantly smaller in overall footprint. I do believe that was like a trillion parameter model but don't quote me on that entirely as well as some competitors from other labs like Opus 4.8, Fable 5, Gemini 31 Pro and 56 Soul on Max effort. I think having this like around the performance level of Opus 4.8 is the general takeaway at least from these benchmark JPEGs right here which would be pretty exciting. Now, if you are interested in seeing the deep SWE benchmark, cuz I know a lot of folks put more weight onto that, at the bottom of this post, they do have a bunch of more intricate like benchmark scores for all of these models. And on that Deepswe 1.1, we see that it scores 56.6. So, not quite on par with Opus 4.8 and definitely less than Fable 5 and GPT56. However, something interesting here is the leap between its predecessor and itself. a significantly higher score from Quen 3.7 Max. So, should be interesting nonetheless to see how this performs. Now, in this introduction post, basically a lot of what I'm noticing here is just mentions of this working autonomously over long periods of time, 16 days of autonomous coding. Additionally to that, something I find pretty interesting because it somewhat relates to something I'm working on right now is they talked about this right here, which like autonomous chip design and closed loop feedback driven optimization. And then they have this little like thing right here where all these like gates, registers, things that are related to like microchips. Apparently, it designed this. Again, I don't have the pertinent knowledge to be able to look at this and be like, "Oh, wow. This is magic." Or like, "Oh, no. This wouldn't work." But I still think it's pretty interesting that one of the things they've chosen to highlight here is like a chip design thing. Cuz that is pretty cool. And like I said, it relates. So, I'm building like a custom piece of hardware right now. And for those interested, don't worry about it. But like I have um just PCB mock-ups being printed right now because it helps me to like assess how the design will be and things like that. So that was just something that definitely caught my eye cuz it is pretty darn cool. So in terms of some more specifics about this model, the price is $2 per million input tokens and $6 per million output tokens. It is multimodal and it can also accept video in which is pretty cool. Output is only text, but it can take image, text, and video in and then output text. So, it definitely will have the ability to multimodally code, which I guess is exciting. Again, it's a 2.4 trillion parameter mixture of experts model with 95 billion active. Our maximum context length is a little under a million right there. Context window is a million and the max output is 128K. And they also have like some adjustable reasoning parameters and things of the sort like you would expect. Now, in terms of how we're going to be running these tests today, I am going to do some simpler ones like the browser OS test v2.5, which I have already begun just from within the web chat interface here. However, I have also set this up using OpenAI's codecs, just the command line code editor version of Codeex because this is Linux. I have to say I've been enjoying using various models with this since the Deep Seek Light model or Flash model came out. They actually suggested I think using it through codec. So, it made me kind of eager to try other models in this as well. So, I have ensured that all cash from previous runs has been wiped just because we've been using this in the few recent videos. So, I don't want any memory of this to basically be interjecting or influencing the results. But, we can see right here it is on XHigh mode. All right. So, it