
Gemini's agentic video understanding cuts video analysis tokens up to 88%
Google shipped agentic video understanding on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite: the model picks which segments, frames, audio, or transcripts to inspect rather than ingesting video at a fixed 1 FPS, cutting tokens up to 88% and analysis cost up to 66% while raising accuracy up to 7%. Enabled with one API setting at standard token pricing and no feature fee, it removes the cost-versus-detail tradeoff on 90-minute lectures and multi-hour recordings, and puts a Flash-class model at the accuracy-to-cost frontier for video.
Source: blog.google ↗
agentic video understanding reduces analysis costs by up to 66% and token consumption by up to 88%, while improving accuracy by up to 7%
Google
Why this matters
- → Reduces video analysis tokens by 88%, cutting costs 66% while improving accuracy
- → Enables long-form video processing without forced tradeoff between cost and detail
- → Agentic selection of frames/audio/transcripts eliminates fixed-rate ingestion overhead
Smarter video parsing