Vibeleaderboard
← All Intel
Intel / article

ttok 1.0

Source
simonwillison.net
Date
Why it matters

If you budget with a token counter, the default should now match GPT-5/6 tokenization. Early evidence suggests GPT-6 adds no input-count change, though OpenAI has not confirmed it.

Key takeaways · AI-distilled
  • Simon Willison shipped ttok 1.0 after noticing that the just-released 0.4 still counted and truncated with the GPT-4 tokenizer by default; the 1.0 release switches the default to the GPT-5/GPT-6 tokenizer.
  • OpenAI has not confirmed that GPT-6 shares the GPT-5 family's tokenizer, and Willison notes an open, contested issue about it, so the new default rests on outside testing rather than an official spec.
  • The evidence Willison cites is an experiment by William Liu: seven GPT models (5.5, 5.6 Sol/Terra/Luna, 6 Astra/Sol/Luna) each reported 44,794 and agreed on all 31 test fixtures, suggesting GPT-6 brings no input-count change on that corpus.
Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Read the source simonwillison.net
Recommended reads
Comments

Checking sign-in…

Loading comments…