- Category
- AI/LLM model
- Rank
- No. 1038Tools index
- Type
- APP
- GitHub
- 5.8k stars
- Added
- Aug 15, 2026
About
Needle 2 Needle 2 is an open 45M-parameter model for tool calling, device use and structured extraction. The whole model is a single 14MB binary that runs a full session in about 28MB of RAM. It is built on our Simple Attention Network findings, compressed to CQ2-bit with Cactus Quants, and baked into its own engine. On the benchmarks below, Needle 2 trades wins with other small models like FunctionGemma 270M, LFM2.5 230M and Apple FM, at 5x to 70x smaller, and 2 bits against their f16. This rep
Why it made the leaderboard
On-device tool calling at 14MB puts agent function-calling inside mobile and embedded targets that cannot host a hosted API round trip.
Tags
cactusgeminigemmallmon-device-ai
Tech Stack
Python
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.
