
2 years ago, we achieved the first big milestone with @exolabs. We clustered 2 MacBooks to run Llama 405B. It felt like magic. The consensus was running this model was only possible in a data center. We ran it on consumer hardware, on 2 M3 Max MacBook Pros. Most people thought it was a gimmick. It only ran at 2 tok/sec! But, we believed that improvements to the software, hardware, and models would all compound. So that maybe in a few years, we thought, this would improve 10x in software, 10x in hardware, 10x models = 1000x. That was the vision. We imagined a world where you would have frontier intelligence running quietly on your desk. Today is the day that vision became reality. The M5 Ultra is a 10x step-change improvement vs the M3 Max we originally clustered. That, compounded with software improvements like RDMA over Thunderbolt, MTP and better kernels, and high intelligence density models like Qwen 3.8 27B, means we now have 1,000x better Local AI than when we started. I am so grateful to the small group of people at Apple (including @doogie69 @awnihannun @angeloskath @DiganiJagrit @doogie69) who believed in this vision and had the foresight as well as the…

Unmetered local at API-class speed changes build-vs-buy math for : a desk-side Mac Studio cluster becomes a credible target for workloads you'd otherwise route to a hosted API.
Checking sign-in…
Loading comments…