
We have open-sourced HY-1.8B-2Bit, a high-efficiency 2-bit LLM built for on-device deployment. This model scales the 1.8B base down to an effective 0.3B parameter footprint, requiring only 600MB of storage, making it smaller than many mobile apps. 🔹 Ultra-Low-Bit Strategy: Uses QAT (Quantization-Aware Training) to reach a 2-bit representation (0.3B bit-equivalent size). 🔹 Dual-CoT Reasoning: Retains sophisticated Dual Chain-of-Thought capabilities despite radical precision reduction. 🔹 Performance: 3-8x faster prefill on Apple M4 and MediaTek Dimensity 9500; 2-3x faster token generation on-device. 🔹 Benchmark Gains: Achieves a 17% average accuracy lead over models of equivalent size. 🔹 Hardware Synergy: Optimized for Arm SME2 and modern consumer silicon. HY-1.8B-2Bit is available now in GGUF format for seamless integration into edge-based inference engines. Project Page: https://t.co/pFp6vgpooa Weights: https://t.co/Pvua0eec3L GGUF Version: https://t.co/wdRIsUqwFb Technical Report:
QAT to 2 bits with reasoning intact means a 600MB model that fits inside a mobile app bundle, and the GGUF release drops straight into existing edge engines.
Checking sign-in…
Loading comments…