Skip to content

Languages

language is an optional hint you set on a connection or a request. Leave it off and the model works out what is being spoken, and keeps working it out as the speaker changes, which is the right default for most traffic.

Eleven Indian languages, plus English, Spanish and Arabic. Each is returned in its own script: Hindi in Devanagari, Tamil in Tamil, Urdu in Perso-Arabic, English in Latin.

Language language Language language
Hindi hi-IN Kannada kn-IN
English en-US Malayalam ml-IN
Bengali bn-IN Odia od-IN
Marathi mr-IN Punjabi pa-IN
Telugu te-IN Urdu ur-PK
Tamil ta-IN Spanish es-US
Gujarati gu-IN Arabic ar-AR

Accuracy is not the same across all of them. Punjabi, Marathi and Urdu are the strongest of the Indian languages; Telugu, Tamil, Malayalam and Kannada are the weakest and are improving with each checkpoint. A language is offered only once it transcribes at least as well as the weakest one already offered. Benchmarks carries the current numbers.

Leave language off and the stream follows whoever is speaking. A classifier reads the model’s own encoder output as it is produced and says which language is being spoken; that language decodes the next chunk. A speaker changing language is not a special case, it is the next chunk answering differently, and the sentence in progress is not restarted.

{ "type": "language", "languages": ["te-IN"], "t": 4.2 }

The event arrives when the language settles and again whenever it changes. On our own measurements the classifier is right 98.4% of the time on a second of speech and 97.2% of the time over a whole clip. The languages it confuses most are the ones people confuse too: Kannada, Tamil and Telugu in the first seconds of a clip, and Hindi with Urdu.

Two things follow from how it works. A stream that opens with silence is not forced to guess: nothing is decided until somebody speaks. And a language that only appears late in a call is picked up when it appears, rather than being missed because the opening was in something else.

Urdu is not detected, only served by name. Hindi and Urdu are one spoken language written in two scripts, and a Hindi speaker whose transcript switched to Urdu script mid-sentence would be the worst mistake detection could make. Ask for ur-PK if that is what you want.

The model marks the end of each sentence with the language it heard that sentence in. Those marks never reach your text; they become the language of each word instead, so a sentence that moves between Hindi and English comes back with every word labelled.

On a stream, ask for words by reading the words event (see Streaming). Each word carries its language:

{ "text": "meeting", "language": "en-US", "speaker": 0, "start": 3.12, "end": 3.48, "final": true }

Until its sentence ends, a word carries the language the stream was being decoded in at the time, and it can change once the sentence is marked. That is why such a word is not yet final.

A stream can be pinned to two languages instead, which decodes both and writes each phrase in whichever of the two it was spoken in:

?language=hi-IN+en-US

Set this when you know both languages in advance and want neither to be chosen for you.

Set language when you already know what the audio is. It skips detection, so the first words are settled immediately rather than after a few seconds, and it removes the chance of the model choosing wrong.

?language=ta-IN one language, no detection, no second decode
?language=ta-IN+en-US this pair, pinned from the first word

The hint is fixed for the life of a WebSocket session. To change it, open a new connection. Batch takes it per request.

A region refinement is accepted and narrowed when we do not have it, so en-IN is served as English.

If a language you need is missing, tell us at founders@utter.cc, or book a call.