All web services

Speech and voice on Okou

Turn written text into an audio file, in a voice you choose and a delivery you can describe. No text to speech account of your own.

Generation · No setup · Managed credits per successful call

Give it words and it gives you audio. Choose a voice, and describe how the line should land, calm and unhurried or brisk and bright, in the same plain language you would use with a person.

The audio comes back as a file the agent can use immediately: dropped into a video, attached to a message, or published on a page. Okou holds the account and charges by how much audio you actually got, so there is no voice platform plan in your name.

What Speech and voice is

Most teams need a voice track long before they need a voice strategy. A demo needs narration, a report could be listened to on a commute, a video needs a line read before it can be cut.

This is that step as a call. Text goes in, an audio file comes out in a voice you picked, with a short instruction about delivery when the default reading is not the one you want.

Because it is priced by the length of the audio rather than by the request, a workflow that reads out a paragraph costs a paragraph. That is what makes narration reasonable to add to a scheduled job rather than a project of its own.

It is the same thing people shop for as a text to speech API or an AI voiceover tool. The difference is where the account sits, and that the caller is usually an agent finishing a piece of work rather than a person exporting a file.

What Speech and voice can do

What an agent can do with it, and what comes back when it does.

What you give itThe text to speak, in any of the supported languages
What comes backAn audio file at a link the agent can use straight away
VoiceA choice of named voices
DeliveryA plain-language instruction about tone and pace
Typical useNarration for video, audio versions of written work, spoken alerts

Coverage and limits

What it does not do, said up front, so nobody plans a workflow around something that is out of scope.

Chosen voices, not cloned ones

You pick from the voices the service offers. Recreating a specific person's voice is a different product with different consent questions, and belongs on a provider connector with that person's agreement, not here.

One read at a time

Each call is one voice reading one piece of text. A conversation between two speakers is several calls and an edit, which is worth planning before writing the script.

Delivery is a request, not a control

An instruction about tone changes the reading, but it is not a timeline with markers on it. Anything that has to hit an exact beat is easier to fix in an edit than to specify in words.

It reads what you wrote

Numbers, abbreviations and names get read the way the model thinks they are said. Where it matters, spelling them out in the text is faster than arguing with the delivery.

What Speech and voice costs

Web services are billed in credits per call, not per token. There is no separate vendor bill to reconcile.

BillingManaged credits per successful call
Runs onOpenAI
Commandzero generate voice

Charged in credits by the length of the audio that comes back, which works out at roughly nineteen credits for a minute of speech. A short narration line costs very little, and an hour of audio is a decision worth making on purpose. There is no subscription underneath and no minimum.

Product names and logos belong to their owners and appear here only to say what a service runs on. They do not imply any endorsement or partnership.

Setup and access

Nothing to set up

Nothing to connect. No text to speech account, no key, no plan. If you need a specific cloned voice or a studio feature set, Okou also ships connectors for voice platforms that run on your own account.

Who can use it

Producing an audio file is a permission you grant per agent and per person, and it is the same permission that covers other generated files, so an agent that can speak can also draw.

What teams use Speech and voice for

Narration for something the agent just made

A clip, a walkthrough, a weekly summary. The agent writes the script and reads it, and the piece is finished rather than waiting for someone with a microphone.

An audio version of written work

A briefing, a report, a set of release notes. The people who never open the document listen to it on the way in, and it costs a few credits per read.

Spoken alerts and prompts

Short lines that need to be heard rather than read, generated when they are needed instead of recorded in advance for every case.

When not to use Speech and voice

Skip it when the voice has to be a particular person's, which is a consent question before it is a technical one. Skip it for long-form audio you will re-cut repeatedly, where regenerating the whole file each time is wasteful. And skip it when a person reading badly would still be better, which is usually anything where the point is that a person said it.

What this is called elsewhere

The same thing goes by several names in the market. If you have shopped for one of these, this is how it maps to what you get here.

Text to speech API

Words in, audio out, over an interface a program can call. That is exactly this, with the account and the key on our side.

AI voiceover

The narration track for a video or a demo, produced from the script rather than recorded. The usual reason teams reach for this service.

TTS without an API key

What people search for when the model is not the problem and the account is. There is no signup and nothing to paste into an environment variable.

Article to audio

Turning written work into something listenable. Priced by the length of the audio, so it scales with what you actually publish.

Voice cloning

A different thing, deliberately not offered here. Use a voice platform connector on your own account, with the consent of the person whose voice it is.

Speech and voice compared

Speech and voice vs a text to speech account of your own

Same idea, far less setup. No key to hold, no plan to choose, and access is a permission rather than a shared secret.

Speech and voice vs a voice platform connector

A connector runs on your account and gives you that platform's full range, including cloned voices and studio controls. This is the fast path for narration that just needs to exist.

Speech and voice vs recording it yourself

A person reading is still better when the point is that a person said it. For a line that has to be regenerated every week because the numbers changed, recording is the wrong tool.

The short version

The quickest way to put a voice on work an agent already produced. Named voices only, priced by the minute, and nothing to sign up for.

Frequently asked questions

Do I need an OpenAI account?

No. Okou makes the call and charges the audio in credits. There is no key of yours in the workflow.

Can I choose the voice?

Yes, from the voices the service offers, and you can describe the delivery you want in plain language alongside it.

Can it clone my voice?

No. Cloned voices are deliberately not part of this service. Use a voice platform connector on your own account, with the consent of the person involved.

What does it cost?

Credits by the length of the audio, around nineteen credits for a minute of speech. Short lines cost very little.

Which languages does it handle?

The model reads several languages. For anything that has to sound native, it is worth listening to a sample before committing a workflow to it.

Can I use it for a podcast or a video?

Yes. The output is a normal audio file, so it goes into an edit like any other track.

Can two people have a conversation?

Not in one call. Generate each part separately and assemble them, which is also how you get the pauses right.

How do I get a name pronounced correctly?

Spell it phonetically in the text you send. It is faster and more reliable than describing the pronunciation in the delivery instruction.

How you ask for Speech and voice

You describe what you want in plain language. The agent works out which service it needs and makes the calls.

Narrate a clip

Write a twenty-second script for this demo video and read it in a calm, unhurried voice.

Listen to the report

Turn this week's summary into audio I can listen to on the way in, and keep it under four minutes.

One line, read well

Read this warning message in a firm, matter-of-fact voice and give me the file.