Try it live in the API playground
Drop text with the markers below into the
text field and send a real request to hear the emotion.Overview
Fish Audio models support 64+ emotional expressions and voice styles that can be controlled through text markers in your input. Add natural pauses, laughter, and other human-like elements to make speech more engaging and realistic.How It Works
Add emotional or stylistic cues in square brackets within your text:Complete Emotion Reference
Basic Emotions (24 expressions)
Advanced Emotions (25 expressions)
Sound & Delivery Markers
These markers aren’t emotions — they shape how a line is delivered, add natural human sounds, or layer in ambient effects. Combine them with the emotion cues above.Tone Markers (6 expressions)
Control volume, intensity, and emphasis. Place[emphasis] right before the word or phrase you want to stress:
Audio Effects (11 expressions)
Add natural human sounds:Special Effects
Additional markers for atmosphere and context:
You can also use natural expressions like “Ha,ha,ha” for laughter without tags.
Usage Guidelines
Placement Rules
For S2:- Sentence-level emotion cues usually work best at the beginning of sentences
- Tone controls can go anywhere in the text
- Sound effects can go anywhere in the text
- Bracket cues can use natural language descriptions and are not limited to a fixed set of tags
Advanced Techniques
Combining Effects
You can layer multiple emotions for complex expressions:Emotion Transitions
Create natural emotional progressions:Background Effects
Add atmospheric sounds:Intensity Modifiers
Fine-tune emotional intensity with descriptive modifiers:Language Support
All 13 supported languages can use emotion markers. For sentence-level control, cues usually work best at the sentence start in these languages:- English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish, Portuguese
Best Practices
Do’s
- Use one primary emotion per sentence
- Test different emotion combinations
- Match emotions to context logically
- Add appropriate text after sound effects (e.g., “Ha ha” after laughing)
- Use natural expressions when possible
- Space out emotional changes for realism
Don’ts
- Don’t overuse emotion tags in short text
- Don’t mix conflicting emotions
- Don’t make bracket descriptions so long that they interrupt readability
- Don’t forget brackets
- Don’t place sentence-level emotion cues far from the sentence they control
Common Use Cases
Customer Service
Storytelling
Educational Content
Marketing & Sales
Troubleshooting
Emotion Not Working?
- Check placement - Put the cue where the emotion or effect should begin
- Keep wording clear - Use concise natural language descriptions
- Use the right syntax - S2 cues use square brackets; S1 cues must use parentheses
Unnatural Sound?
- Space out emotional changes
- Use appropriate intensity
- Test with different voices
- Add context text after sound effects
Performance Notes
- Emotion markers don’t count toward token limits
- No additional latency for emotion processing
- All emotions available on all pricing tiers
- Maximum of 3 combined emotions per sentence recommended
Quick Reference Tables
Emotion Intensity Scale
Common Combinations
S1 (legacy) syntax
The default S2-Pro model uses[bracket] cues with free-form natural language. The previous-generation S1 model uses the same emotion names but requires (parentheses) and a fixed tag set:
Basic emotions (S1)
Basic emotions (S1)
Advanced emotions (S1)
Advanced emotions (S1)
Tone markers (S1)
Tone markers (S1)
Audio effects (S1)
Audio effects (S1)
Special effects (S1)
Special effects (S1)
See Also
- API Reference - Implementation details
- Text-to-Speech Guide and Best Practices

