← All IntelIntel / article
ttok 1.0
- Source
- simonwillison.net
- Date
Why it matters
If you budget with a token counter, the default should now match GPT-5/6 tokenization. Early evidence suggests GPT-6 adds no input-count change, though OpenAI has not confirmed it.
Key takeaways · AI-distilled
- Simon Willison shipped ttok 1.0 after noticing that the just-released 0.4 still counted and truncated with the GPT-4 tokenizer by default; the 1.0 release switches the default to the GPT-5/GPT-6 tokenizer.
- OpenAI has not confirmed that GPT-6 shares the GPT-5 family's tokenizer, and Willison notes an open, contested issue about it, so the new default rests on outside testing rather than an official spec.
- The evidence Willison cites is an experiment by William Liu: seven GPT models (5.5, 5.6 Sol/Terra/Luna, 6 Astra/Sol/Luna) each reported 44,794 and agreed on all 31 test fixtures, suggesting GPT-6 brings no input-count change on that corpus.
Terms in this piece · Glossary
- token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Read the source simonwillison.net
Recommended reads
Comments
Checking sign-in…
Loading comments…