Vibeleaderboard
← All Intel
Intel / post

Ling-3.0-tiny launches at 7.9B total and 1.3B active parameters

Source
Ant Ling
Date
Ant Ling@AntLingAGI
Thread · 6 parts

Today, we’re releasing Ling-3.0-tiny: 7.9B total parameters, with only 1.3B active per token. A native hybrid reasoning model built for real-world tasks, math, instruction following, and resource-sensitive deployment. More intelligence with less compute. 🧵

🔒📚 Demo: Fully Local Inference & Knowledge Retrieval Ling-3.0-tiny runs entirely locally, with no cloud dependency. Integrated with obsidian-cli, it can retrieve, organize, and generate content from local text repositories—reducing data exposure and network overhead.

⚡ Demo: Fast Local Generation for Everyday Tasks Ling-3.0-tiny delivers precise text generation for local web translation and routine tasks, helping reduce compute and operational costs for high-frequency API workloads.

🛠️ Demo: Tool Use & Automated Environment Control Ling-3.0-tiny natively supports tool use for cost-efficient automation, making it well suited to mobile-device and browser UI control tasks.

Read the full thread on X

Context

Ant Ling releases Ling-3.0-tiny, a 7.9-billion-parameter hybrid-reasoning model that activates only 1.3 billion parameters per token, describing it as built for math, instruction-following and what the company calls resource-sensitive deployment. The thread shows three demonstrations: running the model fully locally with no cloud dependency to retrieve and organize notes through an Obsidian integration, using it for fast local web translation to cut per-request compute cost, and having it control a browser or mobile-device interface through native . These were shown as video walkthroughs and are described here only as the company's written summary, not independently assessed.

Ling-3.0-tiny launched on OpenRouter and Vercel's AI Gateway with a free trial through August 13, 2026, and Ant Ling said open-source weights would follow later. Three weeks after this launch, the independent evaluator Artificial Analysis benchmarked the model directly, reporting it sits on a favorable intelligence-versus-speed frontier for phone-class , though it uses more tokens per task than similarly sized peers to get there.

Terms in this piece · Glossary
  • mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
  • tool use — A model's ability to call external functions — run code, search the web, edit files — instead of only generating text.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
More from Ant Ling
Recommended reads
Comments

Checking sign-in…

Loading comments…