Clip transcript
>> Yeah, exactly. Like >> The tokens have time. Like every single token when it comes to sort of speech tokens, audio tokens, they are there's a hop length, you know? Every token is 80 milliseconds or 100 milliseconds or something like that. Implicitly, it's all there. However, they're LMs at the end of the day, they just don't have this concept of time. There's a joke where you know, I actually actually I don't know. Let's try it. Like, if you ask me what the essence of comedy is. >> What's the essence of comedy? >> Timing. >> [laughter] >> Okay, so software engineers are cooked, but comedians, stand-up comedians still have a job. >> Exactly. They'll have great careers on the stage. >> So, you're saying that the information is technically, physically there. It's just not exploiting it. It's just not doing the math of counting how many tokens have I seen times 80 milliseconds. >> Yeah, exactly. It would be interesting to see sort of how newer models handle things like that. Um because in addition to sort of the model itself, there's also the fact that there's network latency, and there's the latency of going from the sort of the audio or the audio tokens into the actual audio waveforms and natural speech, getting all of that transported over the internet, playing on your phone, etc. And so, there's sort of some delays there. Those are the kind of delays that I think us as humans, we've adapted to handling cuz, you know, we're on a voice call right now. We're on a video call. And we know that there is some delay. We know to anticipate that delay as well. A lot of the sort of voice models themselves don't quite know how to deal with right