415.tech
AI & tech, from the frontlines of Silicon Valley
Gemini's agentic video understanding cuts video analysis tokens up to 88%

Gemini's agentic video understanding cuts video analysis tokens up to 88%

Google shipped agentic video understanding on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite: the model picks which segments, frames, audio, or transcripts to inspect rather than ingesting video at a fixed 1 FPS, cutting tokens up to 88% and analysis cost up to 66% while raising accuracy up to 7%. Enabled with one API setting at standard token pricing and no feature fee, it removes the cost-versus-detail tradeoff on 90-minute lectures and multi-hour recordings, and puts a Flash-class model at the accuracy-to-cost frontier for video.

Source: blog.google

Post on XEmail

agentic video understanding reduces analysis costs by up to 66% and token consumption by up to 88%, while improving accuracy by up to 7%

Google

Why this matters

  • → Reduces video analysis tokens by 88%, cutting costs 66% while improving accuracy
  • → Enables long-form video processing without forced tradeoff between cost and detail
  • → Agentic selection of frames/audio/transcripts eliminates fixed-rate ingestion overhead
Smarter video parsing
Also in this edition