Skip to content

Text to speech

Send text, get a voice back. Hindi and English are open today, with five voices. Text to speech is free while it is in preview: requests are recorded against your account, and charged nothing.

POST https://api.utter.cc/v1/speech
Authorization: Bearer utt_live_…
Content-Type: application/json
{
"text": "नमस्ते! आपका ऑर्डर कल शाम तक पहुँच जाएगा।",
"language": "hi-IN",
"voice": "sofia",
"format": "wav"
}
Field Required Default Values
text yes Up to 1,000 characters. Punctuation paces the reading: each sentence is spoken as a unit.
language no en-US hi-IN or en-US.
voice no aria aria, jason, john, leo, sofia.
format no wav wav for one file when the whole text has been spoken; pcm to stream it.

Write times and quantities out when it matters how they are read: “ten in the morning” is read more reliably than “10 am”.

format: "wav" returns one audio/wav file: 16-bit mono PCM at 22,050 Hz.

format: "pcm" returns raw 16-bit little-endian mono PCM (audio/L16), streamed one sentence at a time. Start playing as soon as the first bytes arrive; later sentences follow while you play. The sample rate is in the X-Sample-Rate header.

Every response carries:

Header Meaning
X-Sample-Rate Samples per second of the audio, 22050.
X-Characters Characters of text spoken, which is what is metered.
X-Request-Id Quote this if you contact support about the request.
Terminal window
curl https://api.utter.cc/v1/speech \
-H "Authorization: Bearer $UTTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"text": "Your order will arrive by tomorrow evening.", "language": "en-US", "voice": "leo"}' \
-o speech.wav
speak.py
import os
import httpx
body = {"text": "नमस्ते! आपका ऑर्डर कल शाम तक पहुँच जाएगा।", "language": "hi-IN",
"voice": "sofia", "format": "pcm"}
headers = {"Authorization": f"Bearer {os.environ['UTTER_API_KEY']}"}
with httpx.stream("POST", "https://api.utter.cc/v1/speech", json=body,
headers=headers, timeout=120) as r:
r.raise_for_status()
rate = int(r.headers["X-Sample-Rate"])
with open("speech.pcm", "wb") as f:
for chunk in r.iter_bytes():
f.write(chunk) # or hand each chunk to your audio output
print(f"raw 16-bit mono PCM at {rate} Hz")

GET /v1/speech/voices lists the languages and voices this deployment offers, with the character limit:

{
"languages": ["en-US", "hi-IN"],
"voices": [{"id": "aria", "name": "Aria"}, {"id": "jason", "name": "Jason"},
{"id": "john", "name": "John"}, {"id": "leo", "name": "Leo"},
{"id": "sofia", "name": "Sofia"}],
"default_voice": "aria",
"sample_rate": 22050,
"max_characters": 1000
}

Refusals use the usual error envelope. The ones specific to this endpoint:

Status type When
400 empty_text text is empty or only whitespace.
400 text_too_long text is over 1,000 characters. Split it and send the parts in order.
400 invalid_request An unsupported language or voice.

The voices are NVIDIA’s Magpie TTS multilingual model, served from our own GPUs.