Languages
language is an optional hint you set on a connection or a request. Leave it off and the
model works out what is being spoken, and keeps working it out as the speaker changes, which
is the right default for most traffic.
Supported today
Section titled “Supported today”Eleven Indian languages, plus English, Spanish and Arabic. Each is returned in its own script: Hindi in Devanagari, Tamil in Tamil, Urdu in Perso-Arabic, English in Latin.
| Language | language |
Language | language |
|---|---|---|---|
| Hindi | hi-IN |
Kannada | kn-IN |
| English | en-US |
Malayalam | ml-IN |
| Bengali | bn-IN |
Odia | od-IN |
| Marathi | mr-IN |
Punjabi | pa-IN |
| Telugu | te-IN |
Urdu | ur-PK |
| Tamil | ta-IN |
Spanish | es-US |
| Gujarati | gu-IN |
Arabic | ar-AR |
Accuracy is not the same across all of them. Punjabi, Marathi and Urdu are the strongest of the Indian languages; Telugu, Tamil, Malayalam and Kannada are the weakest and are improving with each checkpoint. A language is offered only once it transcribes at least as well as the weakest one already offered. Benchmarks carries the current numbers.
Detection
Section titled “Detection”Leave language off and the stream follows whoever is speaking. A classifier reads the
model’s own encoder output as it is produced and says which language is being spoken; that
language decodes the next chunk. A speaker changing language is not a special case, it is the
next chunk answering differently, and the sentence in progress is not restarted.
{ "type": "language", "languages": ["te-IN"], "t": 4.2 }The event arrives when the language settles and again whenever it changes. On our own measurements the classifier is right 98.4% of the time on a second of speech and 97.2% of the time over a whole clip. The languages it confuses most are the ones people confuse too: Kannada, Tamil and Telugu in the first seconds of a clip, and Hindi with Urdu.
Two things follow from how it works. A stream that opens with silence is not forced to guess: nothing is decided until somebody speaks. And a language that only appears late in a call is picked up when it appears, rather than being missed because the opening was in something else.
Urdu is not detected, only served by name. Hindi and Urdu are one spoken language written in
two scripts, and a Hindi speaker whose transcript switched to Urdu script mid-sentence would
be the worst mistake detection could make. Ask for ur-PK if that is what you want.
The language of every word
Section titled “The language of every word”The model marks the end of each sentence with the language it heard that sentence in. Those
marks never reach your text; they become the language of each word instead, so a sentence
that moves between Hindi and English comes back with every word labelled.
On a stream, ask for words by reading the words event (see
Streaming). Each word carries its language:
{ "text": "meeting", "language": "en-US", "speaker": 0, "start": 3.12, "end": 3.48, "final": true }Until its sentence ends, a word carries the language the stream was being decoded in at the
time, and it can change once the sentence is marked. That is why such a word is not yet
final.
Two languages at once
Section titled “Two languages at once”A stream can be pinned to two languages instead, which decodes both and writes each phrase in whichever of the two it was spoken in:
?language=hi-IN+en-USSet this when you know both languages in advance and want neither to be chosen for you.
When to set the hint
Section titled “When to set the hint”Set language when you already know what the audio is. It skips detection, so the first
words are settled immediately rather than after a few seconds, and it removes the chance
of the model choosing wrong.
?language=ta-IN one language, no detection, no second decode?language=ta-IN+en-US this pair, pinned from the first wordThe hint is fixed for the life of a WebSocket session. To change it, open a new connection. Batch takes it per request.
A region refinement is accepted and narrowed when we do not have it, so en-IN is served
as English.
If a language you need is missing, tell us at founders@utter.cc, or book a call.