Authorization: Bearer header; on
WebSockets, ?api_key=<key> in the URL also works. The
Groq SDK speaks the same protocol — see
Migrate from Groq.
Text to speech
POST /v1/audio/speech
- Set
response_formatexplicitly. It defaults topcmhere, notmp3as on OpenAI. Code that omits it and writes the bytes toout.mp3produces a file that won’t play. voiceis a Fish voice ID — see Voices. Preset names (nova,echo, …) are a 400, exceptalloy, which happens to synthesize in an unrelated voice — see capabilities.instructionsis refused with a 400. UseX-Fish-*headers orprovider.optionsfor delivery control instead.stream_formatmust beaudio.stream_format: "sse"is a first-class parameter in current OpenAI SDKs and is refused with a 400 here — the response is always the raw audio stream.
speed works as on OpenAI. sample_rate (Hz) triggers a real resample, but
each codec accepts only certain rates and asking for one it doesn’t take is a
400, not a silent fallback — the per-codec rate table is in
Audio formats.
One asymmetry between the two SDKs: sample_rate is not a first-class
parameter in the OpenAI Python SDK, so passing it to create() raises
TypeError. Send it through extra_body. The Node SDK is more permissive and
accepts it inline in the request object.
tts-1, tts-1-hd, and
gpt-4o-mini-tts are accepted as aliases for fish-audio/s2.1-pro. A name
that is neither a known alias nor a Fish model is refused with a 400 — never
silently substituted.
Transcription
POST /v1/audio/transcriptions
timestamp_granularities=["word"] is what adds the top-level words array;
without it verbose_json returns segments[] alone. (Fish times every word, so
segments[] is per-word either way — the flag controls the words array, not
the precision.)
Subtitle formats (srt, vtt) are aggregated from Fish’s word-level
timestamps into phrase-length cues, so they won’t line up one-to-one with the
per-word entries verbose_json returns in segments[].
In verbose_json, the language field echoes the language you sent — it is
not a detection result, and transcribe-1 needs no hint to work. Send none and
the field comes back as an empty string, which is the expected response, not a
failure. Don’t route on it. (The value you send is still passed to Fish’s ASR;
it’s only the reported field that is an echo rather than a measurement.)
Transcription model names carry over the same way: whisper-1,
gpt-4o-transcribe, and gpt-4o-mini-transcribe are accepted as aliases for
fish-audio/transcribe-1; an unrecognized name is refused with the same 400.
In
verbose_json, the Whisper-engine internals (avg_logprob,
no_speech_prob, compression_ratio, temperature) are neutral constants —
don’t build quality filters on them.client.audio.translations has no counterpart here —
POST /v1/audio/translations is a 404. Transcribe in the source language and
translate the text downstream.
Chat-modality audio
POST /v1/chat/completions serves TTS and STT in the chat-completions shape,
for SDKs and frameworks that only speak chat — LangChain, Vercel AI SDK, and
similar.
For speech output, request the audio modality; the last user message’s text
is synthesized:
stream: true the audio arrives over SSE in base64 chunks that are safe
to concatenate as they come.
For speech input, put an input_audio content part in the last user
message; the reply’s content is the transcript.
Two refusals on this endpoint: n must be 1, and one call cannot combine audio
input with audio output.
Realtime WebSocket
wss://api.fish.audio/compat/v1/realtime?model=fish-audio/s2.1-pro
The SDK derives the WebSocket URL from the same client:
input_audio_buffer.commit ends an utterance —
server-side VAD is refused, not ignored), and transcription runs over the same
socket or in a dedicated transcription session. The full event list, format
and voice rules, and transcription-session handshakes are in the
Realtime protocol reference.
Error handling
Errors map to your SDK’s typed exceptions (AuthenticationError,
RateLimitError, APIStatusError, …). The envelope carries the HTTP status as
an integer code:
type: "provider_error" plus
"metadata": {"provider_name": "fish-audio"}. Errors from the compatibility
layer itself carry no metadata. Rate limits return 429 with a Retry-After
header.
Authentication produces both kinds, and the type tells you which problem you
have. No key, or a header that can’t be parsed, is a 401 with
type: "authentication_error" and no metadata — nothing was checked against
your account. A key that fails validation is a 401 with
type: "provider_error" and metadata.provider_name: "fish-audio" — the key
itself is invalid. Both surface as your SDK’s AuthenticationError.
GET /v1/models is unauthenticated and answers even without a key, so it is
not a way to check whether a key is valid.
Going further
Compatibility
The full contract: mappings, limits, and explicit refusals.
Fish-native parameters
latency, temperature, zero-shot cloning — through the OpenAI SDK.
