
TPU🚨 is working with the popular OSS inference optimization library Mooncake on integrating TPU with Mooncake Store. Similar to NVL72, KVCache DRAM P2P pooling will initially happen on the scale-out network via TENT instead of using ICI/NVLink. 🔥 Mooncake basically improves performance per TCO of production inference!

pooling across accelerators is one of the larger levers on cost, and its arrival on TPU changes the comparison between serving platforms.
postAMD code lands upstream in NVIDIA's inference transport library
postCongrats to @Zai_org on GLM-5.3. It massively beats every American open model: NSign in to comment.
Loading comments…