Create a room, share the code, start speaking.
Captions only suppresses synthesised speech for that language. Audio output is the most expensive rate — roughly 1.5× text output and 4× audio input — so a language whose audience only reads should not be speaking.
For an academic lecture, Mione, Katerina, Andre and Theo Calm read most naturally at lecture pace. Use your own voice instead →
Record 10–20 seconds of clear speech and your translated lectures will be spoken in your own voice, in every language. Enroll once and reuse it — each enrollment is billed and deleting does not refund the quota.
There is no similarity setting — the only lever is the sample, and improving it means enrolling again. In order of impact:
Note that your clone is only ever heard speaking a different language from the one you record — your own language is captions-only. Cross-lingual synthesis carries timbre across less faithfully, so it will always sound less similar here than in a vendor demo.
WAV, MP3 or M4A · mono · at least 24 kHz · 10–60 seconds · max 10 MB. The 16 kHz files used for translation testing are not suitable — use the original recording.
These appear in the vendor's SDK documentation but are unverified against this enrollment API. They are only sent when you tick them, so the default stays the payload known to work. Trying them is cheap — a rejected enrollment is not billed.
One source term = target term pair per line. Names,
technical terms, anything the recogniser is likely to mangle.
Pick the wireless transmitter, not the built-in laptop microphone.
A USB receiver usually appears under a generic name such as
Microphone/Headphones USB2.0 Device — the brand name
will not be shown.
Off by default for an external transmitter. These are tuned for a built-in microphone in a laptop chassis; against a clean line-level feed gain control pumps and noise suppression can gate out a quiet passage entirely.
Level appears once the session starts.
This is the speech recogniser's view of your own words, before translation.
Each language's session is restarted and about a second of your audio is buffered across the swap, so students should not hear a gap.
Audio output is the most expensive rate — roughly 1.5× text output and 4× audio input. Setting a language to captions only removes its audio output entirely.