How to size Provisioned Throughput capacity for MiniMax M3
Source
MiniMax (official)
Author
MiniMax (official)
Date
Terms in this piece · Glossary
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Why it matters
It supplies the conversion rates and formula to size reserved M3 capacity against your own uncached, cached and output tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → mix, and shows how cache hit rate rather than request volume drives what you have to provision.