A dialogue block can be spoken aloud, so students hear the language instead of only reading it. You choose a voice, an accent, and a speaking style for each speaker, then generate the audio.
Voice selection and audio generation are paid features. On a free account the speaker controls show a lock.
Set up your speakers
Add a Dialogue block to your lesson, or open one you already have. A new dialogue offers Generate dialogue, where Roshi writes the script, or Write dialogue manually, where you type it yourself. Either way the voice settings work the same.
The speakers appear as a row of avatars above the lines. Use the + button to add a speaker.
Click a speaker's avatar to open Edit Speaker. Alongside Voice and Accent you can set the speaker's Name and give them a face, by upload, from the Image Library, from Templates, or with Generate image.
Choose a voice
In Edit Speaker, the Voice dropdown offers eight voices, described by who they sound like rather than by a brand name:
Warm woman, Warm man
Authoritative woman, Authoritative man
Bright young woman, Bright young man
Calm older woman, Calm older man
The dropdown starts on Select a voice, so every speaker needs one chosen. Use the play button beside the dropdown to hear a sample before you commit.
Pick voices that are easy to tell apart. Two warm adult voices of the same gender are hard for beginners to follow, even when the accents differ.
Choose an accent
The Accent dropdown starts on Select an accent and offers four:
Neutral (North American), the default
British
Australian
South Asian (Indian)
Accent belongs in this dropdown. Do not type an accent into the style box, because the style box does not control pronunciation.
Why a two-accent dialogue sounds less smooth
When one speaker has an accent, Roshi can generate the whole conversation in one pass and the turn-taking sounds natural. When two speakers have different accents, Roshi has to generate the audio line by line and stitch it together, so the pauses between turns can sound slightly flatter.
Roshi shows you a note in the editor when this applies. It is a quality trade-off, not an error, and the audio is still usable. If the smoothness matters more than the accent contrast, set one speaker back to Neutral (North American).
Set the speaking style
Base style is a free text box describing how the speaker talks across the whole dialogue: friendly, formal, casual, impatient, speaking slowly for a beginner class.
Keep it to a few words about tone and pace. Long instructions are less reliable than short ones.
Generate the audio
Close the speaker dialog and check the lines are assigned to the right speakers. Use Add line for anything missing.
Click Generate Audio.
Play it back before you share the lesson. Regenerating replaces the audio, so it is safe to try a different voice and run it again. Delete audio removes it without touching the script; Delete lines removes the script.
Generate the audio after the script is final. Editing a line after generating does not re-record it on its own.
If something goes wrong
A line is too long. Each line shows a live character count with its 4,000 character limit, and Roshi will not generate audio while a line is over it. Split that line into two turns.
The wrong voice reads a line. Check which speaker the line is assigned to. The line follows the speaker, not the position on the page.
The controls are locked. Voice selection and audio generation are on paid plans. See Pricing.
The audio did not appear. Generate it again. If it fails twice, use the Help button and tell us the lesson name.
Changing the script itself
Ask Roshi to edit this dialogue, under the block, rewrites the script for you: shorten it, simplify it, change the situation, or add a speaker. Generate the audio afterwards. See Adjust an existing lesson.
Related
Transcribe audio, if you would rather use a real recording