Skip to content

Streaming over WebSocket

Streaming is the main way to use utter. You open one WebSocket, push audio frames as you capture them, and read text back while the speaker is still talking.

wss://api.utter.cc/v1/stream?tier=standard&sample_rate=16000

Authenticate the handshake with the same bearer header used everywhere else:

Authorization: Bearer utt_live_…
Parameter Required Default Values
tier no standard ultra, turbo, standard, economy. See Latency tiers.
language no auto Omit to auto-detect the pair, or name one (ta-IN) or a pair (ta-IN+en-US). See Languages.
sample_rate no 16000 The rate of the PCM you are about to send, in Hz.
diarize no false true labels every word with who said it. See Speakers.

sample_rate declares what you are sending. It does not ask the server to resample. Send 44.1 kHz audio while claiming 16 kHz and you get a confident, completely wrong transcript. Resample on your side.

Audio frames are binary WebSocket messages containing raw PCM:

  • float32, samples in the range -1.0 to 1.0
  • little endian
  • mono, one channel, interleaving is not supported
  • at the rate you declared in sample_rate

No WAV header, no container, no compression. A frame is a whole number of samples, 4 bytes each. Frames of 100 ms to 500 ms of audio work well. Very small frames add per message overhead; very large ones add latency the tier cannot make up for.

At 16 kHz, a 320 ms frame is 16000 * 0.320 = 5120 samples, so 20480 bytes.

Send the text message "eof" when the audio is finished. It is a text frame, not binary, and it is exactly those three characters with no JSON around it.

The server finishes decoding whatever it holds, emits any remaining final events, sends done, and closes. Closing the socket without sending "eof" ends the session, and audio still in flight may not produce a transcript.

Every message from the server is a JSON text frame with a type field.

Sent once, immediately after the handshake. Wait for it before sending audio.

{
"type": "ready",
"request_id": "req_01JD4Z2Q8W6M",
"tier": "standard",
"sample_rate": 16000
}

If the tier you asked for is not servable on this deployment, the connection is refused rather than quietly downgraded. See Latency tiers.

The languages the stream is being transcribed under. Sent once the model has heard enough to tell, normally around three seconds in, and again only if the pair changes.

{
"type": "language",
"languages": ["hi-IN", "en-US"],
"t": 3.14
}

Two languages, because a stream is decoded under a pair and the speaker may switch between them as often as they like. If you set language yourself, this event carries back exactly what you asked for and arrives immediately.

Text that arrived before this event is a partial, and may be rewritten once the pair is settled. Nothing that arrives as a final is ever revised.

The current best guess for the segment being spoken. Partials are unstable by design: the text can change, shrink or be rewritten entirely as more audio arrives. Render them, do not store them.

{
"type": "partial",
"text": "the quarterly report",
"t": 1.28
}

t is seconds of audio elapsed since the start of the session.

A segment has closed and will not change. This is the text to keep.

{
"type": "final",
"text": "I finished the quarterly report and shared it with the team.",
"start": 0.0,
"end": 4.62
}

start and end are seconds from the start of the session. Concatenating every final in order gives you the transcript.

The transcript as words, each with its language and, when you asked for diarize=true, its speaker. Sent alongside partial and final, which carry the same text as plain strings; read whichever suits you.

{
"type": "words",
"from": 12,
"words": [
{ "text": "kal", "language": "hi-IN", "speaker": 1, "start": 4.02, "end": 4.26, "final": true },
{ "text": "meeting", "language": "en-US", "speaker": 1, "start": 4.3, "end": 4.66, "final": true },
{ "text": "hai", "language": "hi-IN", "speaker": 1, "start": 4.7, "end": 4.9, "final": false }
],
"t": 5.12
}

from is the index of the first word that changed. Keep a list; on every words event, cut it at from and append words. Everything before from is exactly what you already hold, so most messages carry only the last word or two. A word marked final never changes again.

A word becomes final once nothing about it can change: the decoder has moved past it, its sentence has ended and been given a language (see The language of every word), and, with diarize=true, the speaker labelling has heard past it. Until then its language or speaker may still be revised.

speaker is null without diarize=true. start and end are seconds from the start of the session, and may be null on a stream pinned to two languages.

Sent once after "eof", just before the server closes the socket.

{
"type": "done",
"audio_seconds": 8.52
}

audio_seconds is the audio metered for this session. See Pricing and credits.

Something went wrong. The socket closes after this.

{
"type": "error",
"error": "invalid_tier",
"message": "Tier 'ultra' is not available on this deployment.",
"request_id": "req_01JD4Z2Q8W6M"
}

On the WebSocket the event is flat: error is the type string, with message and request_id alongside it. See Errors.

ready → language? → (partial | final | words | language)* → done

with error possible at any point, and terminal. There is no guarantee of a partial before a final: a short segment may close without ever producing one.

Add diarize=true and every word in the words events carries a speaker: 0 for the first person heard, 1 for the second, and so on, up to eight. The numbers stay with the same voice for the whole session.

wss://api.utter.cc/v1/stream?tier=standard&diarize=true

Speakers are told apart by voice, alongside the transcription, and a word’s label follows the word by no more than about a second. When two people talk over each other, each word is given to whoever was speaking most across it; overlapping speech is not separated into two transcripts.

Labels need word timings, which a stream pinned to two languages (language=hi-IN+en-US) does not have, so its words keep speaker: null. Detection and single-language streams are labelled.

Grouping consecutive words by speaker gives you turns:

turns = []
for w in words:
if turns and turns[-1]["speaker"] == w["speaker"]:
turns[-1]["text"] += " " + w["text"]
else:
turns.append({"speaker": w["speaker"], "text": w["text"]})

Full example, streaming a file at realtime pace so the partials arrive the way they would from a live microphone. Requires pip install websockets soundfile numpy. There is no utter package on PyPI, so this uses the ordinary websockets client.

stream.py
import asyncio, json, os, sys, time
from urllib.parse import urlencode
import numpy as np
import soundfile as sf
import websockets
SAMPLE_RATE = 16000
FRAME_MS = 320
def url(tier: str = "standard") -> str:
query = urlencode({"tier": tier, "sample_rate": SAMPLE_RATE})
return f"wss://api.utter.cc/v1/stream?{query}"
async def send_audio(ws, audio: np.ndarray, realtime: bool = True) -> None:
"""Push float32 mono PCM in FRAME_MS frames, then the eof sentinel."""
step = SAMPLE_RATE * FRAME_MS // 1000
for i in range(0, len(audio), step):
frame = audio[i : i + step]
# .tobytes() on a slice copies, so only this frame goes on the wire.
await ws.send(frame.tobytes())
if realtime:
await asyncio.sleep(len(frame) / SAMPLE_RATE)
await ws.send("eof")
async def read_events(ws) -> list[str]:
finals: list[str] = []
async for raw in ws:
event = json.loads(raw)
kind = event["type"]
if kind == "partial":
print(f" … {event['text']}", end="\r", flush=True)
elif kind == "final":
print(" " * 78, end="\r")
print(event["text"])
finals.append(event["text"])
elif kind == "done":
print(
f"\n{event['audio_seconds']}s audio"
)
break
elif kind == "error":
raise RuntimeError(f"{event['error']}: {event['message']}")
return finals
async def transcribe(path: str) -> str:
audio, sr = sf.read(path, dtype="float32")
if sr != SAMPLE_RATE:
raise SystemExit(
f"{path} is {sr} Hz. Resample it first, for example:\n"
f" ffmpeg -i {path} -ac 1 -ar {SAMPLE_RATE} out.wav"
)
if audio.ndim > 1:
audio = audio.mean(axis=1)
headers = {"Authorization": f"Bearer {os.environ['UTTER_API_KEY']}"}
async with websockets.connect(url(), additional_headers=headers) as ws:
ready = json.loads(await ws.recv())
if ready["type"] != "ready":
raise RuntimeError(f"expected ready, got {ready}")
print(f"tier {ready['tier']} at {ready['sample_rate']} Hz")
sender = asyncio.create_task(send_audio(ws, audio))
finals = await read_events(ws)
await sender
return " ".join(finals)
if __name__ == "__main__":
print(asyncio.run(transcribe(sys.argv[1])))

Two things this example is careful about, both of which are easy to get wrong:

  • Send and receive concurrently. Reading only after the last frame is sent turns a streaming API into a batch one, and the partials are wasted.
  • Send a copy of the frame. NumPy’s .tobytes() copies. If you reach for a memoryview or a buffer instead, make sure you are not handing over the whole underlying array.

Live microphone capture. The browser gives you float32 samples already, at the AudioContext rate, so the only real work is resampling to 16 kHz and keeping the frames small.

mic.html
<button id="start">Start</button>
<button id="stop" disabled>Stop</button>
<p id="final"></p>
<p id="partial" style="opacity:.6"></p>
mic.js
const SAMPLE_RATE = 16000;
const FRAME = 5120; // 320 ms at 16 kHz
let ws, ctx, node, stream;
async function start() {
// Your server holds the API key and returns a ready to use wss:// URL.
const { url } = await fetch("/api/utter-session", { method: "POST" }).then(
(r) => r.json(),
);
ws = new WebSocket(url);
ws.binaryType = "arraybuffer";
const finalEl = document.getElementById("final");
const partialEl = document.getElementById("partial");
ws.onmessage = (ev) => {
const event = JSON.parse(ev.data);
switch (event.type) {
case "ready":
console.log("ready on tier", event.tier);
break;
case "partial":
partialEl.textContent = event.text;
break;
case "final":
finalEl.textContent += " " + event.text;
partialEl.textContent = "";
break;
case "done":
console.log(`${event.audio_seconds}s of audio`);
ws.close();
break;
case "error":
console.error(event.error, event.message);
ws.close();
break;
}
};
await new Promise((res) => (ws.onopen = res));
stream = await navigator.mediaDevices.getUserMedia({ audio: true });
ctx = new AudioContext({ sampleRate: SAMPLE_RATE });
await ctx.audioWorklet.addModule("pcm-worklet.js");
node = new AudioWorkletNode(ctx, "pcm-frames", {
processorOptions: { frame: FRAME },
});
node.port.onmessage = ({ data }) => {
// data is a fresh Float32Array; .buffer is safe because the worklet
// allocated it per frame rather than handing over a view into a pool.
if (ws.readyState === WebSocket.OPEN) ws.send(data.buffer);
};
ctx.createMediaStreamSource(stream).connect(node);
document.getElementById("stop").disabled = false;
}
function stop() {
ws?.send("eof"); // text frame, the server flushes and sends `done`
node?.disconnect();
ctx?.close();
stream?.getTracks().forEach((t) => t.stop());
document.getElementById("stop").disabled = true;
}
document.getElementById("start").onclick = start;
document.getElementById("stop").onclick = stop;

The worklet just accumulates 128 sample render quanta into 320 ms frames:

pcm-worklet.js
class PcmFrames extends AudioWorkletProcessor {
constructor(options) {
super();
this.frame = options.processorOptions.frame;
this.buf = new Float32Array(this.frame);
this.n = 0;
}
process(inputs) {
const channel = inputs[0]?.[0];
if (!channel) return true;
for (let i = 0; i < channel.length; i++) {
this.buf[this.n++] = channel[i];
if (this.n === this.frame) {
// slice() copies. Posting this.buf itself would send the same
// memory every time and the transcript would be garbage.
this.port.postMessage(this.buf.slice());
this.n = 0;
}
}
return true;
}
}
registerProcessor("pcm-frames", PcmFrames);
One session one WebSocket, one language, one tier
Changing tier or language open a new connection
Silence still audio, still metered

To transcribe two speakers on separate microphones, open two sessions. Two speakers on one microphone need one session with diarize=true.