The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

21FEB2024replayed
one year on
model launchGoogle · Simon Willison

Gemini Pro 1.5 video input lets users extract structured data from footage

Google's new model can process up to an hour of video as a sequence of frames, returning structured data in JSON-like arrays after prompting.

Google has quietly introduced a feature that may redefine how developers interact with video content: Gemini Pro 1.5, announced last week, accepts video as direct input and can extract structured data from moving images.

Developer Simon Willison demonstrated the capability by uploading a seven-second video of a bookshelf and prompting the model for a JSON array of titles. The model returned a list of 21 books, accurately identifying even partially obscured spines. A follow-up with a 22-second cookbook shelf yielded 52 results, though a safety filter initially blocked the word “Cocktail” before repeated prompts succeeded.

Industry observers note that while GPT-4 Vision can process video by sending individual frames, Gemini Pro 1.5’s 1,000,000-token context window allows it to handle up to an hour of footage at roughly one frame per second, using approximately 258 tokens per frame. The practicality of extracting structured data from video without manual frame selection has generated significant interest, with the Hacker News thread accumulating over 1,100 points and 482 comments in its first day.

S
Simon Willison@simonw

Demonstrated extracting book titles from a 7-second video and a 22-second cookbook shelf video using Gemini Pro 1.5, achieving JSON output with minor hallucinations and safety filter issues.

H
Hacker News commenters

Debated whether video processing is simply frame-by-frame analysis, with some pointing to similar capabilities in GPT-4V but noting Gemini's superior context window and token efficiency.

One year later — open only if you can handle spoilers

Within a year, Gemini Pro 1.5's video input became a standard tool for automated document scanning and inventory management, though safety filter quirks continued to frustrate users. Google later expanded the feature to support audio alongside video in subsequent model versions.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy