Qwen3 introduces hybrid reasoning (switchable thinking/non-thinking modes) and models where a 30B-A3B activates only 3B parameters yet outperforms much larger models, giving builders strong, self-hostable alternatives to closed frontier APIs.
“Our flagship model, Qwen3-235B-A22B, achieves competitive results in benchmark evaluations of coding, math, general capabilities, etc., when compared to other top-tier models such as DeepSeek-R1, o1, o3-mini, Grok-3, and Gemini-2.5-Pro.”
“Additionally, the small MoE model, Qwen3-30B-A3B, outcompetes QwQ-32B with 10 times of activated parameters, and even a tiny model like Qwen3-4B can rival the performance of Qwen2.”
Checking sign-in…
Loading comments…