Skip to main content

Prerequisites

Sign up for a free Fish Audio account to get started with our API.
  1. Go to fish.audio/auth/signup
  2. Fill in your details to create an account, complete steps to verify your account.
  3. Log in to your account and navigate to the API section
Once you have an account, you’ll need an API key to authenticate your requests.
  1. Log in to your Fish Audio Dashboard
  2. Navigate to the API Keys section
  3. Click “Create New Key” and give it a descriptive name, set a expiration if desired
  4. Copy your key and store it securely
Keep your API key secret! Never commit it to version control or share it publicly.

Recipe

A voice agent is three stages chained together: asr.transcribe() turns the caller’s audio into text, your own LLM turns that text into a reply, and tts.stream() turns the reply back into speech. The transcript and the reply are just strings, so the only Fish Audio-specific parts are the first and last calls. Streaming the reply lets you start writing (or forwarding) audio before the whole sentence is synthesized. The recipe transcribes with transcribe-1-pro. Every tab selects it with the model request header; without that header, the request is served and billed as transcribe-1.
heard is an ASRResponse: heard.text is the full transcript and heard.duration is the clip length in seconds. With transcribe-1-pro, heard.text can contain inline speaker markers such as <|speaker:0|> and cues such as [laughter]. The recipe removes the speaker markers before it calls your LLM and keeps the cues as context; see Speaker markers. The Python calls pass include_timestamps=False because the loop needs only the text, and word timestamps add processing time. Pass language="en" as a hint when you know the caller’s language; see Language.
For the lowest latency, feed your LLM’s token stream straight into stream_websocket() instead of waiting for the full reply string. See Realtime: LLM tokens → speech.

Reply in the caller’s voice

reference_id points the reply at a saved voice. Drop it to use the default voice, or clone the caller’s voice from the same clip you just transcribed by passing references instead. See Instant voice cloning.