
New research with @Tsinghua_Uni: Spatial-TTT. A framework for streaming visual-based spatial intelligence with test-time training (TTT). Spatial-TTT adapts fast weights to capture and organize spatial evidence from long video streams, enabling models to build structured 3D spatial memory over time. Highlights: 🔹Efficient streaming memory. Fast weights act as compact spatial memory with sublinear memory growth over 7000+ frames and more than 40% lower compute. 🔹Spatial-predictive mechanism. TTT layers with 3D spatiotemporal convolution capture geometric correspondence and temporal continuity. 🔹SOTA results on long-horizon video spatial understanding (VSI-Bench). The paper ranked #1 on @huggingface Daily Papers on March 13. Project page: https://t.co/Lw0WxIWgGr GitHub: https://t.co/95c4EomcCa Paper: https://t.co/TMW8LciRMn Model & Data:
A concrete answer to unbounded cost in video: adapt fast weights instead of growing a buffer, with numbers on both memory growth and long-horizon spatial accuracy.
Checking sign-in…
Loading comments…