Skip to main content

Prerequisites

Sign up for a free Fish Audio account to get started with our API.
  1. Go to fish.audio/auth/signup
  2. Fill in your details to create an account, complete steps to verify your account.
  3. Log in to your account and navigate to the API section
Once you have an account, you’ll need an API key to authenticate your requests.
  1. Log in to your Fish Audio Dashboard
  2. Navigate to the API Keys section
  3. Click “Create New Key” and give it a descriptive name, set a expiration if desired
  4. Copy your key and store it securely
Keep your API key secret! Never commit it to version control or share it publicly.

Recipe

This recipe uses transcribe-1-pro, which handles recordings up to 60 minutes long. Both tabs select it with the model request header; without that header, the request is served and billed as transcribe-1. Call asr.transcribe() with include_timestamps=True (in JavaScript, ignore_timestamps: false) to get timed segments (ASRSegment). Segments are word-level (one character, or a few, in Chinese and Japanese) and carry no punctuation, so group consecutive segments into caption-sized cues before you format them. Segment start / end are in seconds; SRT wants HH:MM:SS,mmm (comma), WebVTT wants HH:MM:SS.mmm (dot). The recipe starts a new cue after a pause longer than 0.6 seconds, or when the cue would grow past 42 characters or 6 seconds. It joins words with a space, except between Chinese or Japanese characters, which are written without spaces. A segment can have start equal to end, so a cue with no duration gets a short one that ends no later than the next cue starts. Every cue therefore ends after it starts, as SRT and WebVTT players expect.
Both tabs send the same request and write the same captions.srt and captions.vtt. Long recordings can take several minutes to transcribe, so both tabs set a 15-minute request timeout, longer than the SDKs’ default; in Node.js, also see the note below. Tune MAX_CHARS, MAX_MS, and MAX_GAP_MS for your player and audience. segments is empty when no speech was found or word timing was temporarily unavailable, and the files then contain no cues, so check the cue count before you publish them.
In Node.js, the built-in fetch that the JavaScript SDK uses stops waiting for a response after 5 minutes, even with a longer timeoutInSeconds. To wait longer, install undici@7 and, once at startup, call its setGlobalDispatcher(new Agent({ headersTimeout: 900_000, bodyTimeout: 900_000 })). See Processing time and timeouts.
Captions built from segments have no punctuation. Because this recipe requests timestamps from transcribe-1-pro, the response also includes speaker_turns, with each turn’s punctuated text and its start and end time. To keep speakers in separate cues, start a new cue at each turn boundary. In JavaScript, read result.speaker_turns; it is not in the SDK’s TypeScript type. The Python SDK does not return speaker_turns; to read it from the API, see Speaker turns.