On September 5, 2026, Google released the video understanding proxy function in the Gemini Flash model. This function allows the model to independently determine the timing, frame rate, and format of viewing, replacing the traditional fixed processing mode. This reduces video tokens by up to 88%, lowers costs by 66%, and improves accuracy by 7% in standard benchmark tests. Currently, this function is available only as a hosted API; it does not support open-source weights and must be invoked through Google AI Studio and Gemini Enterprise Agent Platform. File uploads and public YouTube links are supported. This feature is applicable to Gemini 3.8, 3.7, 3.6 Flash, and 3.5 Flash-Lite models, aiming to improve the efficiency of long-content video analysis.
“Google introduced the proxy video understanding feature in the Gemini Flash model this week. This feature reduces video tokens by up to 88% and costs by 66%, while improving accuracy by 7% on standard benchmark tests. It replaces the traditional fixed-processing approach of processing one frame per second with the model deciding autonomously when to watch, at what frame rate, and which modality to use. Only necessary content is loaded as needed. This feature is currently available only as a hosted API and does not support open-source weight; it must be called through Google AI Studio and Gemini Enterprise Agent Platform. It supports file uploads and public YouTube links. Compared to static processing, the proxy mode focuses video analysis efficiency on longer content, such as 10-minute tutorials or multi-hour recordings, and adds processing calls and result steps in the response to support progress tracking. This feature is applicable to Gemini 3.8, 3.7, 3.6 Flash, and 3.5 Flash-Lite models…”