TPUšØ is working with the popular OSS inference optimization library Mooncake on integrating TPU with Mooncake Store. Similar to NVL72, KVCache DRAM P2P pooling will initially happen on the scale-out network via TENT instead of using ICI/NVLink. š„ Mooncake basically improves performance per TCO of production inference!
KV cacheThe memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.Full definition ā pooling across accelerators is one of the larger levers on inferenceRunning a trained model to get answers ā the phase where AI is actually used, as opposed to trained.Full definition ā cost, and its arrival on TPU changes the comparison between serving platforms.
Terms in this piece Ā· Glossary
inference ā Running a trained model to get answers ā the phase where AI is actually used, as opposed to trained.
KV cache ā The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.