Skip to main content
Turn spoken audio into text with transcribe-1-pro, Fish Audio’s recommended automatic speech recognition (ASR) model. Send an audio file to POST /v1/asr with the header model: transcribe-1-pro to receive a transcript, its duration, and optional word-level timestamps. Pro handles multi-speaker conversations and recordings up to 60 minutes: it marks speakers inline in the transcript, returns a structured list of speaker turns when you request timestamps, and keeps emotion and vocal-event cues.

Use it in the web app

No code: upload audio, get a transcript.

API reference

Every parameter for POST /v1/asr.

Cookbooks

Captions, batch transcription, and more.

Choose an ASR model

Select the model with the model HTTP header, and send model: transcribe-1-pro with every request; model is not a form or body field. Write the value exactly, in lowercase. If the header is missing, or its value is not an exact match (for example Transcribe-1-Pro or transcribe-1pro), the request is served and billed as transcribe-1, and no error is returned. If you expected transcribe-1-pro but the transcript has no speaker markers, check the header. See ASR pricing for usage costs. transcribe-1-pro accepts these fields: If you use transcribe-1, select it with model: transcribe-1 and send only audio, language, and ignore_timestamps. The other fields apply to transcribe-1-pro only.

When to use it

Captions & subtitles

Group word-level timestamps into SRT/VTT cues.

Meeting & call notes

Use Pro to transcribe multi-speaker recordings for summaries and search.

Voice commands & notes

Turn short utterances into text your app can act on.

Accessibility

Make audio and video content readable.

Quick start

Create an API key and set FISH_API_KEY in your environment. For the Python examples, install the Fish Audio SDK. Every example on this page selects transcribe-1-pro. The examples below also request timestamps.
The response contains text, duration (seconds), segments, and request_id, plus speaker_turns when timestamps are requested, and language_code and language when the language is known. Do not depend on the order of keys in the JSON. If you use transcribe-1, rely on text, duration, segments, and the language fields. The Python SDK currently returns only text, duration, and segments. To read the other fields, call the API directly (see Direct API).
The Python SDK selects Pro through RequestOptions.additional_headers; asr.transcribe() has no model argument, and it cannot send the transcribe-1-pro fields such as diarize. The JavaScript example calls the REST API directly so the header is explicit. For multipart uploads, let your HTTP client set Content-Type and its boundary.

Read the timestamps

Request timestamps with ignore_timestamps=false. The default is true, which skips them. Timestamps add processing time. In multipart forms, send true or false: any other value, including 1 or an empty value, is read as false and turns timestamps on. Each segment is { "text", "start", "end" }, with times in seconds. Segments are word-level: a segment is usually one word, or in Chinese and Japanese usually one character or a few. Segment text has no punctuation and no speaker markers or cues, and can be normalized (for example 35 for 3.5), so it does not always match text character for character. A segment can have start equal to end. segments is an empty array, never omitted, when you do not request timestamps, when no speech is found, or when timing is temporarily unavailable. Segments are not speaker turns; to see who spoke when, use speaker_turns.
For captions, group consecutive segments into caption-sized cues instead of writing one cue per word. The captions cookbook shows how.
In the Python SDK, segment timestamps are on by default. Pass include_timestamps=False to skip them. That’s the inverse of the API/JavaScript flag ignore_timestamps.

Multi-speaker conversations

transcribe-1-pro supports recordings with multiple speakers, such as interviews, meetings, and calls. Send the recording as one audio file. Speaker labels are consistent within one response, not across requests, so do not split a conversation into several requests.

Speaker markers

Pro identifies speakers with inline <|speaker:N|> markers in the response’s text string, where N is a numeric speaker label. A marker assigns the following text to that speaker until the next marker. The same label can appear again when that speaker resumes talking. Markers usually have a space on each side (<|speaker:0|> 你好。 <|speaker:1|> ...); do not rely on exact spacing. Any text before the first marker belongs to the first turn. A transcript with no marker comes from a single speaker; treat it as speaker 0. For example, this illustrative response contains three turns from two speakers. It uses ignore_timestamps=true, so segments is empty:
Treat speaker labels as identifiers within that response, not as names or identities you can match across separate requests. To display turns without requesting timestamps, split text at the speaker markers, keeping any emotion cues in each turn. For example, after decoding the JSON response into result:
When you request timestamps, prefer speaker_turns over parsing text.

Speaker turns

With ignore_timestamps=false, transcribe-1-pro also returns speaker_turns, a list of who spoke when. The field is present when ignore_timestamps=false and diarize is not false; otherwise it is absent. Speaker identification is a transcribe-1-pro feature. This illustrative response is the same conversation with timestamps requested:
  • speaker_turns lists the turns in the order they occur, and is [] when no speech was found. Each turn is { "speaker", "text", "start", "end" }.
  • speaker is the string speaker:N, where N matches the <|speaker:N|> marker in text. Labels identify speakers within one response only. They are not names, and the same label in two requests is not the same person.
  • text is that turn’s speech without speaker markers. It keeps emotion and event cues unless tag_audio_events=false.
  • start and end are in seconds and come from the word timestamps. If word timing is unavailable (segments is empty), turn times are approximate and can cover the whole recording.
  • Consecutive turns can have the same speaker. Do not assume turns are contiguous or non-overlapping.

Speaker options

These transcribe-1-pro fields control speaker identification:
  • diarize: auto (default) or true returns speaker_turns when timestamps are requested. false omits speaker_turns. It does not change the transcript, and the speaker markers stay in text.
  • num_speakers: the expected number of speakers. This is a hint, applied on a best-effort basis; it guides speaker identification on longer recordings and may have no effect on short ones. It cannot be combined with min_speakers or max_speakers.
  • min_speakers, max_speakers: bounds on the number of speakers, with the same best-effort rule. min_speakers must not exceed max_speakers.
Speaker counts are integers of 1 or more. Speaker-count hints cannot be sent with diarize=false. Invalid values return 400 with the code invalid_parameter. For example, to transcribe an interview with two speakers and print each turn:

Emotion and vocal-event cues

Pro can retain cues about how speech sounds alongside the spoken words. These appear as inline bracketed text, such as [高兴] (happy) or [laughter], in the response’s text field and in the turn text of speaker_turns. For example, a transcript might contain:
These are annotations inferred from the audio, not words you need to add to the request. The API has no separate emotion field or fixed emotion enum; preserve the returned text when your application needs these cues. To get a transcript without cues, send tag_audio_events=false (transcribe-1-pro). Cues are then removed from text and speaker_turns; timestamps and billing do not change. Word timestamps ignore speaker markers and cues. Read text for the annotated transcript and segments for timed speech; the segment text may therefore differ from the full transcript.

Implementation details

Language

language is optional. The language is detected automatically whether or not you send a hint, and a hint does not force the transcript into that language. With transcribe-1-pro, the hint does not change the transcript; it matters only when the language cannot be determined, as described below. Use a lowercase ISO 639-1 code such as en, zh, or ja. Other forms, such as en-US or English, may be rejected with 400. The response reports the language in two fields, which are omitted when unknown (never null):
  • language_code: the two-letter ISO 639-1 code, such as en. Use it in application logic.
  • language: the language’s English name, such as English or Chinese. Use it for display.
If the language cannot be determined (for example, very short audio), language_code reports your hint, which is not checked against the audio, and language may be absent. Responses report one language: for recordings that switch languages, language and language_code do not list every language spoken. The Python SDK does not return these fields; call the API directly to read them.

Input audio

transcribe-1-pro accepts WAV, MP3, AAC (including M4A/MP4), FLAC, Ogg (Opus or Vorbis), WebM/Matroska, and MOV, including browser recordings, and uses the first audio track of a video file. Send the original file bytes; no conversion is needed. Not supported: AIFF, CAF, WMA, AMR, AC-3, and raw (headerless) PCM. These return 400. Send audio as multipart/form-data (a file upload, shown above) or application/msgpack (see Direct API). Base64-encoded audio in a JSON body is not supported. If you use transcribe-1, send WAV, MP3, AAC (including M4A/MP4), FLAC, or Ogg (Opus or Vorbis), and convert WebM recordings (for example, from a browser’s MediaRecorder) to Ogg/Opus, MP3, or WAV first, or use transcribe-1-pro.

Limits

These limits apply to transcribe-1-pro:
  • Recordings up to 60 minutes long. Longer audio returns 400 audio_too_long.
  • For long recordings, send compressed audio such as MP3, Opus, or AAC. An hour of 128 kbps MP3 is about 55 MiB. A request that is too large returns 413.
  • Very short clips (under about 0.08 seconds) return 400 audio_too_short.
If you use transcribe-1, send up to 50 MiB per request and keep MP3 and Opus files under 25 MiB; a request that exceeds the size limit returns 413 or 400. transcribe-1 is designed for short recordings; for recordings longer than a few minutes, use transcribe-1-pro. A long request can fail with 503 if it exceeds the processing-time limit. Retry, and if it keeps failing, use transcribe-1-pro or split the audio.

Long recordings

With transcribe-1-pro, send the whole recording, up to 60 minutes, in one request. Timestamps and speaker labels cover the whole recording. Do not split conversations: speaker labels are consistent only within one response. Long recordings can take several minutes, so raise your client timeout and retry on 5xx errors. You are billed for the full duration of the recording once. If you use transcribe-1, keep each request short. For a long recording, use transcribe-1-pro, or split the audio into shorter clips, transcribe each, and offset each clip’s start/end by where it began in the full recording. Check that each response covers the complete clip before combining transcripts.

Processing time and timeouts

Each request returns only when the whole file has been transcribed. Processing time grows with the length of the recording, and timestamps add to it. Long transcribe-1-pro recordings can take several minutes. Set your HTTP client’s timeout accordingly. The official SDKs’ default timeout can be too short for long recordings, and many HTTP libraries default to even less. To wait up to 15 minutes, as several examples on this page do:
  • Python SDK: RequestOptions(timeout=900) for one request, or FishAudio(timeout=900) for the client.
  • httpx: timeout=httpx.Timeout(900.0, connect=10.0). Without it, httpx waits only 5 seconds.
  • Node.js: the built-in fetch stops waiting for a response after 5 minutes, even if you pass a longer AbortSignal. For longer requests, use fetch from the undici package with a dispatcher that waits longer. undici 7 runs on Node.js 20.18.1 and later; undici 8 requires Node.js 22.19 or later.
These values are examples, not a server guarantee. If a long request fails with a 5xx error or the connection drops, retry it. Each request also holds one of your account’s concurrent request slots until its response is returned; see concurrent request limits.

Request IDs

transcribe-1-pro returns a unique request_id in the response body, both on success and on its own errors, and the same value in the x-request-id response header. Platform and network-edge errors do not carry it. Log it, and include it when you contact support. The Python SDK does not return it on success; read it with a direct API call, or from the error body as shown in Errors.

Errors

The error body depends on where the error comes from:
  • transcribe-1-pro errors are JSON with status, message, code, and request_id. Branch on code, not on message.
  • Platform errors, such as a missing or invalid API key (401), insufficient API credit (402), or the concurrency limit (429), are JSON with status and message.
  • Errors from the network edge in front of the API, such as some 413 and 5xx responses, may not have a JSON body.
Retry 429 and 5xx responses with exponential backoff; 429 responses have no Retry-After header. Do not retry other 4xx responses; change the request first. New code values may be added; handle unknown codes by HTTP status. See the API reference for every status and code. With the Python SDK, a failed request raises APIError (RateLimitError for 429, ServerError for 5xx). e.status is the HTTP status, and e.body is the raw response body, where transcribe-1-pro errors carry code and request_id:
If you use transcribe-1, its errors are JSON with status and message; rely only on these two fields.

Async transcription

The Python SDK ships an async client with the same surface, useful when you’re transcribing many files concurrently or already running inside an event loop. Use AsyncFishAudio and await the call:
To run several files in parallel, gather the coroutines:

Direct API (MessagePack)

POST /v1/asr also accepts a MessagePack body instead of multipart form data, the same path the API reference links to for low-overhead, server-side calls. Calling the API directly also gives you the response fields the Python SDK does not return. Pack the audio bytes and options into one payload and set Content-Type: application/msgpack:
The response is the same as for multipart: text, duration (seconds), segments, and request_id, plus speaker_turns when timestamps are requested, and language_code and language when the language is known. Model selection stays in the HTTP header for both formats. MessagePack values are typed: send booleans for ignore_timestamps and tag_audio_events, integers for speaker counts, and a boolean or "auto" for diarize.

Going further

Generate speech

The reverse direction: text to lifelike audio.

Full API parameters

Every field and the raw response schema.

Python reference

asr.transcribe options and the ASRResponse type.