From Training an LLM from Scratch, Locally — Angelos Perivolaropoulos, ElevenLabs · ≈21:37
“we're going to be using like a character-level tokenizer because it's going to be easy for this project, but most big labs, they don't use character-level tokenization.”
AI Engineer
“Uh and the way it works, it will just look at common patterns. Uh so, if you have like a lot of training data that is that is code itself, it will it will see, and then it will realize, "Okay, for loops seem to be like a good candidate for for that to be a token."”
AI Engineer
“So, it will look at all these different tokens, and then create this tokenizer based on the common uh relationship between them.”
AI Engineer
“So again like there's no there's no specific way there's no like human in the loop in this process. It depends on your training data that you use to train this tokenizer.”
AI Engineer
“Like if your your if your variable names are like very strange like random characters and yeah probably that's not going to be in the tokenizer.”
AI Engineer
articleBpe Stays On Script Structured Encoding For Robust Multilingual Pretokenization 2025 05 30Cohere editorial sitemap
postWhy a token is not a standard unit of AI pricingTibo
articleOne Tokenizer To Rule Them All Emergent Language Plasticity Via Multilingual Tokenizers 2025 05 30Cohere editorial sitemapChecking sign-in…
Loading comments…