> ## Documentation Index
> Fetch the complete documentation index at: https://docs.fish.audio/llms.txt
> Use this file to discover all available pages before exploring further.

# Drama 3

> Text-to-speech model you can direct

export const AudioClip = ({src, title, text}) => {
  const [isPlaying, setIsPlaying] = useState(false);
  const [currentTime, setCurrentTime] = useState(0);
  const [duration, setDuration] = useState(0);
  const audioRef = useRef(null);
  useEffect(() => {
    const audio = audioRef.current;
    if (!audio) return;
    const updateTime = () => setCurrentTime(audio.currentTime);
    const updateDuration = () => setDuration(audio.duration);
    const handlePlay = () => setIsPlaying(true);
    const handlePause = () => setIsPlaying(false);
    audio.addEventListener('timeupdate', updateTime);
    audio.addEventListener('loadedmetadata', updateDuration);
    audio.addEventListener('play', handlePlay);
    audio.addEventListener('pause', handlePause);
    audio.addEventListener('ended', handlePause);
    return () => {
      audio.removeEventListener('timeupdate', updateTime);
      audio.removeEventListener('loadedmetadata', updateDuration);
      audio.removeEventListener('play', handlePlay);
      audio.removeEventListener('pause', handlePause);
      audio.removeEventListener('ended', handlePause);
    };
  }, []);
  const togglePlay = () => {
    const audio = audioRef.current;
    if (!audio) return;
    if (audio.paused) {
      document.querySelectorAll('audio').forEach(other => {
        if (other !== audio) other.pause();
      });
      audio.play().catch(() => setIsPlaying(false));
    } else {
      audio.pause();
    }
  };
  const handleProgressChange = event => {
    const next = parseFloat(event.target.value);
    audioRef.current.currentTime = next;
    setCurrentTime(next);
  };
  const formatTime = time => {
    if (!Number.isFinite(time)) return '0:00';
    const minutes = Math.floor(time / 60);
    const seconds = Math.floor(time % 60);
    return `${minutes}:${seconds.toString().padStart(2, '0')}`;
  };
  return <div className="not-prose my-3 rounded-xl border border-gray-200 px-3 py-2 dark:border-white/10">
      {title || text ? <div className="mb-2">
          {title ? <p className="m-0 text-sm font-medium">{title}</p> : null}
          {text ? <p className="m-0 mt-1 font-mono text-sm leading-relaxed text-gray-600 dark:text-gray-300">
              {text}
            </p> : null}
        </div> : null}
      <div className="flex items-center gap-3">
        <audio ref={audioRef} src={src} preload="metadata" />
        <button onClick={togglePlay} className="flex h-9 w-9 flex-shrink-0 items-center justify-center rounded-full bg-primary text-white transition-opacity hover:opacity-90 focus-visible:ring-2 focus-visible:ring-primary" aria-label={isPlaying ? `Pause ${title || 'sample'}` : `Play ${title || 'sample'}`}>
          {isPlaying ? <svg className="h-4 w-4" fill="currentColor" viewBox="0 0 24 24" aria-hidden="true">
              <path d="M6 4h4v16H6V4zm8 0h4v16h-4V4z" />
            </svg> : <svg className="ml-0.5 h-4 w-4" fill="currentColor" viewBox="0 0 24 24" aria-hidden="true">
              <path d="M8 5v14l11-7z" />
            </svg>}
        </button>
        <span className="w-10 font-mono text-xs tabular-nums text-gray-500 dark:text-gray-400">
          {formatTime(currentTime)}
        </span>
        <div className="relative h-1.5 flex-1 rounded-full bg-gray-200 dark:bg-white/10 focus-within:ring-2 focus-within:ring-primary">
          <div className="absolute top-0 left-0 h-full rounded-full bg-primary" style={{
    width: `${duration ? currentTime / duration * 100 : 0}%`
  }} />
          <input type="range" min="0" max={duration || 0} step="0.1" value={currentTime} onChange={handleProgressChange} aria-label={`Seek ${title || 'sample'}`} className="absolute inset-0 h-full w-full cursor-pointer opacity-0" />
        </div>
        <span className="w-10 text-right font-mono text-xs tabular-nums text-gray-500 dark:text-gray-400">
          {formatTime(duration)}
        </span>
      </div>
    </div>;
};

<Note>
  Drama 3 is in preview. Behavior can still change.
</Note>

Drama 3 is our latest text-to-speech model built for voice direction and content creation. Alongside your script, you describe how each line should be performed, add sounds and pauses, and limit a delivery to specific words.

## What's new

* Direct a performance in plain language, including a language other than the script.
* Add a sound or a pause where it happens, or limit a delivery to specific words.
* Switch language for a phrase, or pin a pronunciation with phoneme tags.
* Keep your current `POST /v1/tts` request and set the model header to `drama-3-preview`.

## Voice cast

Choosing a voice that fits the role is key to getting the best out of Drama 3. These public voices are used in the examples on this page:

| Name | Language | ID |
| - | - | - |
| [Adrian](https://fish.audio/app/text-to-speech/?modelId=bf322df2096a46f18c579d0baa36f41d) | EN-US, male | `bf322df2096a46f18c579d0baa36f41d` |
| [Delia](https://fish.audio/app/text-to-speech/?modelId=c70cfb98867a4bfb9ae20d2401257b92) | EN-US, female | `c70cfb98867a4bfb9ae20d2401257b92` |
| [Sheila](https://fish.audio/app/text-to-speech/?modelId=f3a2b90078d54a65af4b95f63bc7798e) | EN-US, female | `f3a2b90078d54a65af4b95f63bc7798e` |

## How to use

To use Drama 3, set the `model` header to `drama-3-preview` on your existing `POST /v1/tts` request.

| Status | Model header | Availability |
| - | - | - |
| Preview | `drama-3-preview` | Available now |

<Warning>
  If the `model` header is missing or misspelled, the request doesn't fail: it
  falls back to `s2.1-pro`.
</Warning>

## Direct the performance

Directions are plain text in your script, and you can mix them freely in the same line. They work the same way in API requests and in the web app.

### Describe the delivery in your own words

Put a direction in square brackets before the words it applies to. Write it the way you'd brief a voice actor: emotion, intent, pacing, or who the character is talking to.

<AudioClip src="/snippets/drama-3/sheila-01.mp3" title="Sheila" text="[barely holding it together, forcing a smile] I'm fine. Really. Go on without me." />

### Add sounds and pauses

Put a sound or pause in brackets exactly where it should happen, such as a sigh, a laugh, a breath, or a short pause.

<AudioClip src="/snippets/drama-3/sheila-02.mp3" title="Sheila" text="I'm fine. [shaky breath] Really. [small laugh] Go on without me." />

### Write directions in any language

Directions can be in English or your own language, and they don't have to match the language of the script.

<AudioClip src="/snippets/drama-3/sheila-03.mp3" title="Sheila" text="[像是在强忍着眼泪，却努力笑着] I'm fine. Really. Go on without me." />

### Change only specific words

Wrap words in a paired tag, such as `<whisper>…</whisper>`, to change the delivery of those words only. Supported tags: `<whisper>`, `<emphasis>`, `<soft>`, `<fast>`, `<slow>`, `<stress>`, `<forceful>`.

<AudioClip src="/snippets/drama-3/sheila-04.mp3" title="Sheila" text="I'm fine. Really. <whisper>Go on without me.</whisper>" />

### Switch language mid-line (Experimental)

Wrap a phrase in a language tag to switch language for those words only. This is useful for a tutor, a conversation, or any line that mixes two languages.

Examples:

<AudioClip src="/snippets/drama-3/sheila-french.mp3" title="Sheila" text="She looked at the menu and said <french>Je voudrais une table pour deux, s'il vous plaît.</french> The waiter smiled." />

<AudioClip src="/snippets/drama-3/delia-spanish.mp3" title="Delia" text="Welcome to Madrid. <spanish>Bienvenidos a nuestra casa.</spanish> Make yourself at home." />

<AudioClip src="/snippets/drama-3/adrian-german.mp3" title="Adrian" text="In Berlin they greet you with <german>Guten Morgen, wie geht's?</german> and they actually mean it." />

<AudioClip src="/snippets/drama-3/adrian-japanese.mp3" title="Adrian" text="He bowed and said <japanese>ありがとうございました</japanese> before he left." />

<AudioClip src="/snippets/drama-3/delia-chinese.mp3" title="Delia" text="My grandmother always told me <chinese>慢慢来</chinese>, which means take your time." />

<AudioClip src="/snippets/drama-3/sheila-korean.mp3" title="Sheila" text="The song starts with <korean>안녕하세요</korean> and then the beat drops." />

<AudioClip src="/snippets/drama-3/sheila-italian.mp3" title="Sheila" text="<italian>Che bella giornata!</italian> That's what my neighbor yells every morning." />

* Use a voice cloned from a speaker fluent in both languages.
* Wrap only the words that change language.
* Try the pair you need. Accuracy varies by language and voice.

<Warning>
  Language tags are experimental. They are not guaranteed for every language or voice, and the model may switch language on its own when the text is already multilingual.
</Warning>

## Pronunciation

Drama 3 works with the phoneme tags already supported on other models. To pin down how a name or term is pronounced, wrap its phonemes in `<|phoneme_start|>` and `<|phoneme_end|>`. In English (CMU Arpabet), the digit on a vowel is its stress: `1` is primary and `0` is unstressed. These two takes are the same word and the same voice, with the stress on a different syllable.

<AudioClip src="/snippets/drama-3/sheila-phoneme-first.mp3" title="Sheila, stress on the first syllable" text="I am an <|phoneme_start|>EH1 N JH AH0 N IH0 R<|phoneme_end|>." />

<AudioClip src="/snippets/drama-3/sheila-phoneme-last.mp3" title="Sheila, stress on the last syllable" text="I am an <|phoneme_start|>EH0 N JH AH0 N IH1 R<|phoneme_end|>." />

The symbol set depends on the language:

| Language | Phonemes |
| - | - |
| [English](/developer-guide/core-features/fine-grained-control/english) | CMU Arpabet |
| [Chinese](/developer-guide/core-features/fine-grained-control/chinese) | pinyin |
| [Japanese](/developer-guide/core-features/fine-grained-control/japanese) | romaji |

## Multi-speaker dialogue

Drama 3 works with multi-speaker dialogue the same way current models do. Generate a whole scene in one request: pass one voice ID per speaker as an array in `reference_id`, and start each turn with a speaker tag. `<|speaker:0|>` uses the first voice, `<|speaker:1|>` the second, and so on. Directions and paired tags from this page work inside each turn. See `reference_id` in the [Text to Speech API](/api-reference/endpoint/openapi-v1/text-to-speech).

This scene uses three public voices: [Delia](https://fish.audio/app/text-to-speech/?modelId=c70cfb98867a4bfb9ae20d2401257b92) narrates, [Adrian](https://fish.audio/app/text-to-speech/?modelId=bf322df2096a46f18c579d0baa36f41d) is the captain, and [Sheila](https://fish.audio/app/text-to-speech/?modelId=f3a2b90078d54a65af4b95f63bc7798e) is the engineer.

```text theme={null}
Narrator: [calm, measured storyteller voice] The ship had been silent for three days.
Captain: [low, gravelly, tired but in control] Status report.
Engineer: [fast, nervous, stumbling over words] Uh, sir, so, the thing is, the reactor is, it is not, it is not great.
Captain: [slow, dangerously quiet] Define "not great."
Engineer: [voice breaking, close to panic] It is going to <emphasis>blow</emphasis>.
```

The names in the script above are for reading. The request sends the same lines as one `text` string, with a speaker tag instead of a name: `<|speaker:0|>` is the narrator, `<|speaker:1|>` is the captain, and `<|speaker:2|>` is the engineer.

```bash theme={null}
curl --fail --show-error --request POST https://api.fish.audio/v1/tts \
  --header "Authorization: Bearer $FISH_API_KEY" \
  --header "Content-Type: application/json" \
  --header "model: drama-3-preview" \
  --data '{
    "text": "<|speaker:0|>[calm, measured storyteller voice] The ship had been silent for three days.<|speaker:1|>[low, gravelly, tired but in control] Status report.<|speaker:2|>[fast, nervous, stumbling over words] Uh, sir, so, the thing is, the reactor is, it is not, it is not great.<|speaker:1|>[slow, dangerously quiet] Define \"not great.\"<|speaker:2|>[voice breaking, close to panic] It is going to <emphasis>blow</emphasis>.",
    "reference_id": ["c70cfb98867a4bfb9ae20d2401257b92", "bf322df2096a46f18c579d0baa36f41d", "f3a2b90078d54a65af4b95f63bc7798e"],
    "format": "mp3"
  }' \
  --output dialogue.mp3
```

<AudioClip src="/snippets/drama-3/ship.mp3" title="Delia, Adrian, Sheila" />

## Support

Need help? Check out these resources:

* [API Reference](/api-reference/introduction) - Complete API documentation
* [Create a Voice Clone](/api-reference/endpoint/model/create-model) - Create a voice clone model
* [Generate Speech](/api-reference/endpoint/openapi-v1/text-to-speech) - Generate realistic speech
* [Real-time Streaming](/features/realtime-streaming) - WebSocket for real-time streaming
* [Discord Community](https://discord.com/invite/dF9Db2Tt3Y) - Get help from the community
* [Support Email](mailto:support@fish.audio) - Contact our support team


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.