Skip to main content
Drama 3 is in preview. Behavior can still change.
Drama 3 is our latest text-to-speech model built for voice direction and content creation. Alongside your script, you describe how each line should be performed, add sounds and pauses, and limit a delivery to specific words.

What’s new

  • Direct a performance in plain language, including a language other than the script.
  • Add a sound or a pause where it happens, or limit a delivery to specific words.
  • Switch language for a phrase, or pin a pronunciation with phoneme tags.
  • Keep your current POST /v1/tts request and set the model header to drama-3-preview.

Voice cast

Choosing a voice that fits the role is key to getting the best out of Drama 3. These public voices are used in the examples on this page:

How to use

To use Drama 3, set the model header to drama-3-preview on your existing POST /v1/tts request.
If the model header is missing or misspelled, the request doesn’t fail: it falls back to s2.1-pro.

Direct the performance

Directions are plain text in your script, and you can mix them freely in the same line. They work the same way in API requests and in the web app.

Describe the delivery in your own words

Put a direction in square brackets before the words it applies to. Write it the way you’d brief a voice actor: emotion, intent, pacing, or who the character is talking to.

Add sounds and pauses

Put a sound or pause in brackets exactly where it should happen, such as a sigh, a laugh, a breath, or a short pause.

Write directions in any language

Directions can be in English or your own language, and they don’t have to match the language of the script.

Change only specific words

Wrap words in a paired tag, such as <whisper>…</whisper>, to change the delivery of those words only. Supported tags: <whisper>, <emphasis>, <soft>, <fast>, <slow>, <stress>, <forceful>.

Switch language mid-line (Experimental)

Wrap a phrase in a language tag to switch language for those words only. This is useful for a tutor, a conversation, or any line that mixes two languages. Examples:
  • Use a voice cloned from a speaker fluent in both languages.
  • Wrap only the words that change language.
  • Try the pair you need. Accuracy varies by language and voice.
Language tags are experimental. They are not guaranteed for every language or voice, and the model may switch language on its own when the text is already multilingual.

Pronunciation

Drama 3 works with the phoneme tags already supported on other models. To pin down how a name or term is pronounced, wrap its phonemes in <|phoneme_start|> and <|phoneme_end|>. In English (CMU Arpabet), the digit on a vowel is its stress: 1 is primary and 0 is unstressed. These two takes are the same word and the same voice, with the stress on a different syllable. The symbol set depends on the language:

Multi-speaker dialogue

Drama 3 works with multi-speaker dialogue the same way current models do. Generate a whole scene in one request: pass one voice ID per speaker as an array in reference_id, and start each turn with a speaker tag. <|speaker:0|> uses the first voice, <|speaker:1|> the second, and so on. Directions and paired tags from this page work inside each turn. See reference_id in the Text to Speech API. This scene uses three public voices: Delia narrates, Adrian is the captain, and Sheila is the engineer.
The names in the script above are for reading. The request sends the same lines as one text string, with a speaker tag instead of a name: <|speaker:0|> is the narrator, <|speaker:1|> is the captain, and <|speaker:2|> is the engineer.

Support

Need help? Check out these resources: