Batch transcription
If you already have a recording and do not need text while it plays, post the file and get one response.
POST /v1/transcribe HTTP/1.1Host: api.utter.ccAuthorization: Bearer utt_live_…Content-Type: multipart/form-dataRequest
Section titled “Request”Sent as multipart/form-data.
| Field | Required | Default | Notes |
|---|---|---|---|
file |
yes | The audio file. | |
tier |
no | standard |
ultra, turbo, standard, economy. Sets the rate you pay. See Latency tiers. |
language |
no | auto | Omit to auto-detect, or name one (ta-IN). See Languages. |
Unlike the WebSocket, this endpoint accepts a container: WAV, FLAC, MP3, M4A, OGG and WebM are decoded server side, and multi channel audio is mixed down to mono. Sending 16 kHz mono to begin with avoids a decode step and is the fastest path.
curl https://api.utter.cc/v1/transcribe \ -H "Authorization: Bearer $UTTER_API_KEY" \ -F "file=@meeting.wav" \ -F "tier=standard"Python
Section titled “Python”With pip install requests:
import osimport requests
with open("meeting.wav", "rb") as f: response = requests.post( "https://api.utter.cc/v1/transcribe", headers={"Authorization": f"Bearer {os.environ['UTTER_API_KEY']}"}, files={"file": ("meeting.wav", f, "audio/wav")}, data={"tier": "standard"}, timeout=300, )
if response.status_code != 200: error = response.json()["error"] raise SystemExit(f"{error['type']}: {error['message']}")
result = response.json()print(result["text"])for segment in result["segments"]: print(f" [{segment['start']:6.2f} → {segment['end']:6.2f}] {segment['text']}")Response
Section titled “Response”200 OK, application/json:
{ "text": "I finished the quarterly report and shared it with the team. The deck is almost final.", "tier": "standard", "audio_seconds": 8.52, "segments": [ { "start": 0.0, "end": 4.62, "text": "I finished the quarterly report and shared it with the team." }, { "start": 4.9, "end": 8.52, "text": "The deck is almost final." } ], "languages": [ { "text": "I finished the quarterly report and shared it with the team. The deck is almost final.", "language": "en-US" } ], "request_id": "req_01JD4Z2Q8W6M"}| Field | Meaning |
|---|---|
text |
The whole transcript, segments joined in order. |
tier |
The tier that served the request, and the rate it was metered at. |
audio_seconds |
Duration of the decoded audio, in seconds, and the amount metered. |
segments |
Each closed segment, with start and end in seconds from the start of the file. |
languages |
The transcript split wherever the language changes, each run with its language. A file in one language is one run. See The language of every word. |
request_id |
Quote this if you contact support about the request. |
Choosing batch or streaming
Section titled “Choosing batch or streaming”Batch is one request and one response, which is simpler to operate. It costs the
same per second of audio as streaming on the same tier, so for a recording the
cheapest correct choice is usually economy, whose extra latency does not matter
when nobody is waiting on a live caption.
Reach for streaming when a person is listening, when the audio is being captured now, or when you want to show text before the speaker stops.
Errors
Section titled “Errors”Failures use the same envelope as everything else, with the HTTP status carrying the category. See Errors.