
Running Kimi K3 on my desk? It will require 4 x 512GB M3 Ultra Mac Studios (2TB @ 3.2TB/s). Waiting on active parameter count, but based on the rumours I expect we can run it at 30+ tok/sec with MTP + Tensor Parallelism using RDMA over Thunderbolt 5. Prefill will be slow but 512GB M5 Ultra should make that ~5x faster (expecting in October). It will be on local dot ai with full benchmarks as soon as the weights drop. Comment below for early access - sending access codes out throughout the day.
If you are sizing hardware for local frontier-scale models, this gives ballpark memory, bandwidth, and throughput figures plus the interconnect technique (RDMA over TB5) that makes multi-Mac viable.
postLocal inference crosses over: the M5 Ultra and the 1,000x claimChecking sign-in…
Loading comments…