Apple's marketing video posted by Joz, includes the 4 x M5 Ultra clustered solution. The cluster has 2TB unified memory @ 4.8TB/s.
"Run trillion parameter frontier models locally" https://t.co/qqmXMr3H4X
Local inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → for trillion-parameter models now has a vendor-blessed hardware path: 2TB of unified memory at 4.8TB/s across four clustered Macs, changing what you can run without an API bill or a datacenter.
Terms in this piece · Glossary
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.