All web services

Video understanding on Okou

Read a video instead of watching it: a timestamped transcript first, then the handful of frames worth actually looking at. Currently free.

Runtime · No setup · Recorded for usage analytics, currently 0 credits

An agent cannot watch an hour of video, and neither can you before lunch. So it reads it: the spoken words with timestamps, which is an index of everything that happened, and then stills pulled from the two or three moments that need to be seen rather than read.

It works from a link, from a file shared in chat, or from a file on disk, and calls currently cost nothing. That combination makes it the sensible first step for any recording that arrives with a question attached.

What Video understanding is

Recordings pile up faster than anyone reviews them: demos, walkthroughs, interviews, meetings, sessions someone captured because it seemed important at the time.

The efficient way through one is not to watch it. It is to read the transcript, find the moments that matter by their timestamps, and look only at those. That is a two-step pattern, and both steps are here.

Transcription returns the spoken words with time ranges, which doubles as a table of contents. Frame extraction takes a list of timestamps and returns the pictures, which is what you want when the answer is on screen rather than in the audio.

Because calls currently cost nothing, there is no reason to be selective about which recordings get read. The cost of a workflow that indexes every recording is the same as the cost of a workflow that indexes none of them.

What Video understanding can do

What an agent can do with it, and what comes back when it does.

What you give itA video link, a file shared in chat, or a file on disk
TranscriptThe spoken words with a time range on every line
FramesStills at any timestamps you name
The patternRead the transcript, then look only at the moments that matter
Plain text optionThe transcript without timestamps when it is going into prose

Coverage and limits

What it does not do, said up front, so nobody plans a workflow around something that is out of scope.

Words and pictures, not comprehension

It gives an agent the transcript and the frames. Understanding what happened is then the agent's own reading of those, which is usually better than a summary produced without either.

Audio drives the transcript

A silent screen recording produces nothing to read. For those, frames at intervals are the whole approach, and the timestamps have to come from somewhere other than the words.

Long recordings make long transcripts

An hour of speech is a lot of text to hand a model at once. Working section by section is more reliable than asking for everything in a single pass.

It transcribes, it does not label speakers

You get what was said and when. Who said it comes from context in the transcript, not from the service, so a workflow that needs attribution should say so and check.

What Video understanding costs

Web services are billed in credits per call, not per token. There is no separate vendor bill to reconcile.

BillingRecorded for usage analytics, currently 0 credits
Commandzero video

Calls are recorded so you can see them and currently cost 0 credits, subject to fair use. Frame extraction happens on the machine running the agent, so it costs nothing beyond the time it takes. There is no transcription plan underneath and no per-minute rate to watch.

Setup and access

Nothing to set up

Nothing to connect. No transcription account, no key, no plan. Give an agent a link to a recording and it can read it on the first run.

Who can use it

Reading a recording an agent was given is part of what an agent does with the files it is handed, so there is no separate permission to grant here. What limits it is what you give it access to.

What teams use Video understanding for

Turning a recording into notes

A demo, an interview or a session becomes a timestamped summary with the three screenshots that carry the point, produced while the meeting is still fresh.

Finding the moment in a long session

Somebody says the answer is in there somewhere. The transcript locates it in seconds, and one frame confirms it, instead of a person scrubbing a timeline.

Making video searchable

Transcribe recordings as they arrive and keep the text. A library nobody could search becomes one an agent can answer questions from.

When not to use Video understanding

Skip it when a silent recording needs describing, where frames at intervals are the actual method and the transcript is empty. Skip it when a formal record is required, because an automatic transcript is a working document rather than minutes. And skip transcribing an hour in one pass when the question only concerns one section of it.

What this is called elsewhere

The same thing goes by several names in the market. If you have shopped for one of these, this is how it maps to what you get here.

Video transcription API

Turning speech in a recording into text a program can use. That is the first half of this, with no account or key of yours.

Speech to text

The standard name for the underlying step. Here it comes back with time ranges, which is what makes it useful for navigation rather than just for reading.

Video to text

What people search for when they have a recording and need something they can read, quote and search.

Frame extraction

Pulling stills at chosen timestamps. It is the part people forget to ask for, and it is what makes the transcript actionable when the answer is on screen.

Meeting notes from a recording

The common workflow built from both halves: transcript for what was said, frames for what was shown, and the agent writing it up.

Video understanding compared

Video understanding vs a transcription service

A transcription service gives you an interface, an editor and a stored archive, and charges by the minute. This gives an agent the text and the frames inside a workflow, and currently charges nothing.

Video understanding vs a meeting notetaker

A notetaker joins the call and produces notes as a product. This works on any recording you already have, including ones nobody planned to capture properly.

Video understanding vs asking a model to watch the video

Handing over a whole recording is slow and expensive. Reading the transcript and then looking at three frames gets to the same answer for a fraction of the work.

The short version

The cheapest way to make a recording usable: transcript first as an index, frames second for the moments that need eyes. Currently free, which means there is no reason not to read every recording that arrives.

Frequently asked questions

What does transcription cost?

Calls are recorded and currently cost 0 credits, subject to fair use. Frame extraction runs locally and costs nothing beyond the time.

Where can the video come from?

A link, a file shared in the chat, or a file on the machine running the agent.

Do I get timestamps?

Yes, a time range per line by default, which is what makes the transcript work as an index. You can ask for plain text when it is going into prose.

Does it tell me who was speaking?

No. You get what was said and when. Attribution has to come from context, and a workflow that depends on it should check rather than assume.

Can it read a silent screen recording?

There is nothing to transcribe, so the approach is frames at intervals instead, with the timestamps chosen rather than found in the words.

How long a video can it handle?

Long ones work, but an hour of speech is a lot of text to reason over at once. Section by section is the more reliable pattern.

Which languages does it transcribe?

The common ones the underlying model supports. For an unusual language it is worth transcribing a short sample before building on it.

Can I get a summary instead of a transcript?

Ask for one. The service returns the transcript and the frames, and the agent that read them writes the summary, which is why it can point at the timestamp it came from.

How you ask for Video understanding

You describe what you want in plain language. The agent works out which service it needs and makes the calls.

Notes from a recording

Transcribe this session, summarize the decisions with timestamps, and pull the frames for the three moments worth screenshotting.

Find the moment

Somewhere in this hour-long call they discuss pricing. Find it, quote it, and show me what was on screen.

Make a library searchable

Transcribe every recording in this folder and tell me which ones mention onboarding.