← All IntelClip / OtherLetter-counting failure isn't caused by tokenization
From Why can't ChatGPT Voice set a timer? | Voice AI expert explains · ≈15:30
“they proved that that error happens even if all of the letters are separated in different tokens.”
“So, it's a matter of just not knowing how to count as opposed to being limited by tokenization.”
What’s in it
- Debunks the popular theory that tokenization causes LLM letter-counting failures
- Cites a Spanish university study testing letters split across separate tokens
- Reveals models fail at counting even when each letter is a distinct token
Clip transcript
data is has gone into the whole thing. And yeah, like yeah, they just can't can't count letters. >> I was reading a paper from a university in Spain and they were saying that it's not tokenization that is the issue, which was a hypothesis for a long time because if the same letter is within the same token, our hypothesis was the letter would only be counted once because it just doesn't have visibility within one token. It's kind of atomic. >> Yeah. >> Um but they proved that that error happens even if all of the letters are separated in different tokens. So, it's a matter of just not knowing how to count as opposed to being limited by tokenization. I thought that was an interesting finding.
Comments
Checking sign-in…
Loading comments…