Studio-quality speech.
A third of the price.
A fast, natural text-to-speech API with zero-shot voice cloning. Drop it into the OpenAI SDK you already use and pay $ per million characters — about 18 hours of audio.
No credit card to start. No subscription. No minimums.
Everything you need from a speech API. Nothing you don't pay for.
Built for apps, agents and content pipelines that need great audio at a price that scales.
$ per million characters
One flat price for every voice, including your clones. Prepaid credit, per-character billing, and the exact cost in the response if you want it.
Clone a voice in 10 seconds
Upload or record a short clip and get a reusable voice instantly. We keep an encrypted voice profile, never your recording.
Streams as it speaks
Turn on stream and audio starts after the first sentence. Perfect for voice agents and live narration.
Drop-in OpenAI compatible
Same endpoint, same parameters. Change the base URL and your existing audio.speech code just works.
Reads text like a human
Prices, dates, times, ordinals, acronyms, markdown and emoji are spoken correctly. Add your own pronunciations per request.
Zero data retention
Flip one switch and we keep no logs and no text. Audio is never stored unless you ask for a 24-hour download link.
Pick a voice. Or bring your own.
Every stock voice is licensed for commercial use. Tap to hear it — these clips are generated by the API.
Pay as you go from prepaid credit. Turn on auto top-up and never think about it again.
- free characters when you sign up
- Voice cloning included — no surcharge, no per-voice fee
- Failed requests are refunded automatically
- Free: voices list, text normalization preview, warmups
Price per million characters
Public list prices for comparable neural TTS. Lower is better.
Ship it in five minutes.
Use the official OpenAI SDKs or plain HTTP. Errors come back with a stable code, a human message and a request id. Retries with an Idempotency-Key are never billed twice.
Questions, answered.
How is a "character" counted?
Every Unicode character of the text you send, including spaces and punctuation, before any normalization. About 900 characters is one minute of speech.
Do you store my audio?
No. Audio streams straight back to you. The only exception is if you ask for delivery: "url", which keeps the file for 24 hours so you can download it, then deletes it automatically.
What does zero data retention do?
With ZDR on (per account or per key), we write no request logs and keep none of your text. We still record character counts and charges, because that is your billing record.
Can I clone anyone's voice?
No. You may only clone your own voice or a voice whose speaker has given you documented permission. You must accept our cloning terms before the feature unlocks, and misuse gets accounts closed.
What about cold starts?
Most requests land on a warm instance. If you need a guarantee before a critical moment, call the free POST /v1/warmup endpoint a few seconds ahead.
Which formats are supported?
mp3, opus, aac, flac, wav and raw pcm (24 kHz, 16-bit mono). Streaming works with mp3, opus, aac and pcm.
Your first characters are on us.
Sign in with Google or GitHub and make your first request in under a minute.
Get your API key