token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Why it matters
This explains why video generation was staged, with structure-conditioned Gen-1 coming before open text-to-video, and frames next-frame prediction as the video analogue of next-tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → prediction for learning world structure.