Why Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, Google
www.youtube.com- Category
- AI Tools
- Type
- ARTICLE
- Added
- Jul 28, 2026
About
The constraint on edge AI is not compute, it is RAM, and it is getting worse: phone makers are shipping less of it this year, and a 6GB Raspberry Pi costs 2.5 times what it did at launch. So Cormac Brick's team at Google AI Edge spends its effort making models small enough to fit. A 2 billion parameter Gemma, quantized to 2.9 bits per weight, runs on a Raspberry Pi at about 8 tokens per second and on a Qualcomm NPU fast enough for a few frames of vision a second. Below that sit tiny models, from
Why it made the leaderboard
Names the actual bottleneck for on-device models — DRAM cost, which is getting worse, not better — and shows the quantization and fine-tuning numbers that decide what fits.
Media

Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.