Text to speech
Send text, get a voice back. Hindi and English are open today, with five voices. Text to speech is free while it is in preview: requests are recorded against your account, and charged nothing.
Request
Section titled “Request”POST https://api.utter.cc/v1/speechAuthorization: Bearer utt_live_…Content-Type: application/json{ "text": "नमस्ते! आपका ऑर्डर कल शाम तक पहुँच जाएगा।", "language": "hi-IN", "voice": "sofia", "format": "wav"}| Field | Required | Default | Values |
|---|---|---|---|
text |
yes | Up to 1,000 characters. Punctuation paces the reading: each sentence is spoken as a unit. | |
language |
no | en-US |
hi-IN or en-US. |
voice |
no | aria |
aria, jason, john, leo, sofia. |
format |
no | wav |
wav for one file when the whole text has been spoken; pcm to stream it. |
Write times and quantities out when it matters how they are read: “ten in the morning” is read more reliably than “10 am”.
Response
Section titled “Response”format: "wav" returns one audio/wav file: 16-bit mono PCM at 22,050 Hz.
format: "pcm" returns raw 16-bit little-endian mono PCM (audio/L16), streamed one
sentence at a time. Start playing as soon as the first bytes arrive; later sentences
follow while you play. The sample rate is in the X-Sample-Rate header.
Every response carries:
| Header | Meaning |
|---|---|
X-Sample-Rate |
Samples per second of the audio, 22050. |
X-Characters |
Characters of text spoken, which is what is metered. |
X-Request-Id |
Quote this if you contact support about the request. |
curl https://api.utter.cc/v1/speech \ -H "Authorization: Bearer $UTTER_API_KEY" \ -H "Content-Type: application/json" \ -d '{"text": "Your order will arrive by tomorrow evening.", "language": "en-US", "voice": "leo"}' \ -o speech.wavPython, streaming
Section titled “Python, streaming”import osimport httpx
body = {"text": "नमस्ते! आपका ऑर्डर कल शाम तक पहुँच जाएगा।", "language": "hi-IN", "voice": "sofia", "format": "pcm"}headers = {"Authorization": f"Bearer {os.environ['UTTER_API_KEY']}"}
with httpx.stream("POST", "https://api.utter.cc/v1/speech", json=body, headers=headers, timeout=120) as r: r.raise_for_status() rate = int(r.headers["X-Sample-Rate"]) with open("speech.pcm", "wb") as f: for chunk in r.iter_bytes(): f.write(chunk) # or hand each chunk to your audio outputprint(f"raw 16-bit mono PCM at {rate} Hz")Voices
Section titled “Voices”GET /v1/speech/voices lists the languages and voices this deployment offers, with the
character limit:
{ "languages": ["en-US", "hi-IN"], "voices": [{"id": "aria", "name": "Aria"}, {"id": "jason", "name": "Jason"}, {"id": "john", "name": "John"}, {"id": "leo", "name": "Leo"}, {"id": "sofia", "name": "Sofia"}], "default_voice": "aria", "sample_rate": 22050, "max_characters": 1000}Errors
Section titled “Errors”Refusals use the usual error envelope. The ones specific to this endpoint:
| Status | type |
When |
|---|---|---|
| 400 | empty_text |
text is empty or only whitespace. |
| 400 | text_too_long |
text is over 1,000 characters. Split it and send the parts in order. |
| 400 | invalid_request |
An unsupported language or voice. |
The voices are NVIDIA’s Magpie TTS multilingual model, served from our own GPUs.