Most models watch every frame - T* learns to search. In our latest blog post, we show how T* rethinks long-form video understanding as temporal search, finding the “needles” in long video haystacks with just a few key frames. https://t.co/CPmM49mCpE
T* lets models locate relevant frames in long videos through search instead of scanning every frame, cutting the compute needed for long-form video understanding tasks.