We’re bringing agentic video understanding to our latest Gemini models. They can now analyze videos with better accuracy while using up to 88% fewer tokens. 🧵

Instead of scanning an entire file, Gemini reasons across the video’s transcript, audio, and frames, dynamically adjusting the frame rate to pull the exact moments needed. The efficiency gains are most significant for long-form content, from 10-minute guides to multi-hour recordings. Agentic video understanding is rolling out to 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite via API in @GoogleAIStudio and coming soon in the @GeminiApp. Find out more →


Long video analysis gets substantially cheaper. Selective frame sampling cuts use by up to 88 percent on multi-hour recordings, live now via API on 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite.
Checking sign-in…
Loading comments…