Clip transcript
memory layer for video intelligence? To start, I want to be clear about what makes video different from other data type, right? So this is the first mental model that I want to highlight, which is that video is not a stack of frames. Um so in in a many of my conversations with like developers, you know, who are using our product, a lot of them still treat video as like a stack of images, maybe a transcript being attached or you know, like you know, but essentially like like a frame level, right? And that is a useful approximation for some tasks, but it throw away the thing that makes video very unique, which is continuity, right? So meaning in video derives from space, time, modalities, and sequence. So a better mental model for video is a special temporal volume. So what I mean that inside that volume, you have visual information, speech, sound, motion, OCR, camera changes, scene transition, metadata, and time, right? So the hard part here is really well, how can you preserve any relationship across this volume so that later an application can traverse it? And then you know, especially at the enterprise scale, you know, across like industry like entertainment, sport, you know, um, short-form content, then you're sitting on petabytes of of footage, right? So, finding moment is already hard. So, how can you preset meaning across millions of moments in