K2 Horizon 375B A23B, a new open weights model from UAE's MBZUAI, scores 47 on the Artificial Analysis Intelligence Index, with relatively strong agentic performance and a 30 point jump over its predecessor K2 Horizon 375B A23B is an open weights Mixture-of-Experts model with 375B total and 23B active parameters from @IFM_MBZUAI, MBZUAI's Institute of Foundation Models. It scores 47 on the Intelligence Index, alongside models such as MiniMax-M3 (45, also a MoE with 23B active parameters), and a large upgrade from its predecessor K2 Think V2 (17, 70B dense model). It leads nearby open weights models on agentic evals and has a low hallucination rate, but trails on knowledge and the hardest reasoning evals. K2 Think V2 ranks among the most open models on our Openness Index; MBZUAI is updating the supporting documentation and code for K2 Horizon and we expect to add it to the Openness Index soon. Key takeaways: ➤ Strong on agentic tasks, weaker on knowledge and deep reasoning. MiniMax-M3, a recent model that is close to it on the Intelligence Index, makes the cleanest comparison: K2 Horizon 375B A23B leads on GDPval-AA, our real-world knowledge work benchmark (Elo 1430 vs 1380), and on τ³-Banking (34.2% vs 15.3%), but trails on GPQA Diamond (87.3% vs 92.9%) and Humanity's Last Exam (32.0% vs 39.0%) ➤ Low hallucination rate, driven by abstention rather than knowledge. K2 Horizon 375B A23B attempts only 40% of AA-Omniscience questions, declining the remaining 60% rather than guessing. The result is a 26% hallucination rate, among the lower rates we have measured, while accuracy is 18%, essentially unchanged from K2 Think V2 ➤ A new architecture over its predecessor. K2 Horizon 375B A23B is a 375B parameter Mixture-of-Experts model with 23B active, succeeding the 70B dense K2 Think V2, and extends context from 262K to 512K tokens. Its 23B active parameters match MiniMax-M3 (428B total, 23B active) Key model details: ➤ Architecture: Mixture-of-Experts, 375B total parameters, 23B active ➤ Context window: 512K tokens ➤ Multimodality: Text input and output only ➤ Pricing and availability: Yet to be announced ➤ Licensing: Open weights (license details to be announced)

K2 Horizon 375B A23B performs strongly on agentic work for its intelligence level: its GDPval-AA Elo of 1430 leads other models at a similar intelligence level, such as MiniMax-M3 (1380), on real-world knowledge work tasks such as preparing presentations and analysis

K2 Horizon 375B A23B would rather decline a question than guess. It attempts just 40% of AA-Omniscience questions, producing a 26% hallucination rate that ranks among the lower rates we have measured, with accuracy of 18% essentially unchanged from its predecessor. The reduced hallucination is a huge improvement over K2 Think V2 (71%), taking the AA-Omniscience Index score from -40 to -3

Full breakdown of the individual evaluations in the Artificial Analysis Intelligence Index for K2 Horizon 375B A23B:

An that leads nearby peers on agentic and real-world knowledge-work while trailing on hard reasoning, so it is worth routing tool-use workloads to and keeping off deep reasoning tasks.
Checking sign-in…
Loading comments…