transcribe-1-pro, Fish Audio’s recommended automatic speech recognition (ASR) model. Send an audio file to POST /v1/asr with the header model: transcribe-1-pro to receive a transcript, its duration, and optional word-level timestamps. Pro handles multi-speaker conversations and recordings up to 60 minutes: it marks speakers inline in the transcript, returns a structured list of speaker turns when you request timestamps, and keeps emotion and vocal-event cues.
Use it in the web app
No code: upload audio, get a transcript.
API reference
Every parameter for
POST /v1/asr.Cookbooks
Captions, batch transcription, and more.
Choose an ASR model
Select the model with the
model HTTP header, and send model: transcribe-1-pro with every request; model is not a form or body field. Write the value exactly, in lowercase. If the header is missing, or its value is not an exact match (for example Transcribe-1-Pro or transcribe-1pro), the request is served and billed as transcribe-1, and no error is returned. If you expected transcribe-1-pro but the transcript has no speaker markers, check the header. See ASR pricing for usage costs.
transcribe-1-pro accepts these fields:
If you use
transcribe-1, select it with model: transcribe-1 and send only audio, language, and ignore_timestamps. The other fields apply to transcribe-1-pro only.
When to use it
Captions & subtitles
Group word-level timestamps into SRT/VTT cues.
Meeting & call notes
Use Pro to transcribe multi-speaker recordings for summaries and search.
Voice commands & notes
Turn short utterances into text your app can act on.
Accessibility
Make audio and video content readable.
Quick start
Create an API key and setFISH_API_KEY in your environment. For the Python examples, install the Fish Audio SDK.
Every example on this page selects transcribe-1-pro. The examples below also request timestamps.
text, duration (seconds), segments, and request_id, plus speaker_turns when timestamps are requested, and language_code and language when the language is known. Do not depend on the order of keys in the JSON. If you use transcribe-1, rely on text, duration, segments, and the language fields.
The Python SDK currently returns only text, duration, and segments. To read the other fields, call the API directly (see Direct API).
The Python SDK selects Pro through
RequestOptions.additional_headers;
asr.transcribe() has no model argument, and it cannot send the
transcribe-1-pro fields such as diarize. The JavaScript example calls the
REST API directly so the header is explicit. For multipart uploads, let your
HTTP client set Content-Type and its boundary.Read the timestamps
Request timestamps withignore_timestamps=false. The default is true, which skips them. Timestamps add processing time. In multipart forms, send true or false: any other value, including 1 or an empty value, is read as false and turns timestamps on.
Each segment is { "text", "start", "end" }, with times in seconds. Segments are word-level: a segment is usually one word, or in Chinese and Japanese usually one character or a few. Segment text has no punctuation and no speaker markers or cues, and can be normalized (for example 35 for 3.5), so it does not always match text character for character. A segment can have start equal to end.
segments is an empty array, never omitted, when you do not request timestamps, when no speech is found, or when timing is temporarily unavailable. Segments are not speaker turns; to see who spoke when, use speaker_turns.
In the Python SDK, segment timestamps are on by default. Pass
include_timestamps=False to skip them. That’s the inverse of the
API/JavaScript flag ignore_timestamps.Multi-speaker conversations
transcribe-1-pro supports recordings with multiple speakers, such as interviews, meetings, and calls. Send the recording as one audio file. Speaker labels are consistent within one response, not across requests, so do not split a conversation into several requests.
Speaker markers
Pro identifies speakers with inline<|speaker:N|> markers in the response’s text string, where N is a numeric speaker label. A marker assigns the following text to that speaker until the next marker. The same label can appear again when that speaker resumes talking.
Markers usually have a space on each side (<|speaker:0|> 你好。 <|speaker:1|> ...); do not rely on exact spacing. Any text before the first marker belongs to the first turn. A transcript with no marker comes from a single speaker; treat it as speaker 0.
For example, this illustrative response contains three turns from two speakers. It uses ignore_timestamps=true, so segments is empty:
Treat speaker labels as identifiers within that response, not as names or identities you can match across separate requests. To display turns without requesting timestamps, split
text at the speaker markers, keeping any emotion cues in each turn. For example, after decoding the JSON response into result:
speaker_turns over parsing text.
Speaker turns
Withignore_timestamps=false, transcribe-1-pro also returns speaker_turns, a list of who spoke when. The field is present when ignore_timestamps=false and diarize is not false; otherwise it is absent. Speaker identification is a transcribe-1-pro feature.
This illustrative response is the same conversation with timestamps requested:
speaker_turnslists the turns in the order they occur, and is[]when no speech was found. Each turn is{ "speaker", "text", "start", "end" }.speakeris the stringspeaker:N, whereNmatches the<|speaker:N|>marker intext. Labels identify speakers within one response only. They are not names, and the same label in two requests is not the same person.textis that turn’s speech without speaker markers. It keeps emotion and event cues unlesstag_audio_events=false.startandendare in seconds and come from the word timestamps. If word timing is unavailable (segmentsis empty), turn times are approximate and can cover the whole recording.- Consecutive turns can have the same speaker. Do not assume turns are contiguous or non-overlapping.
Speaker options
Thesetranscribe-1-pro fields control speaker identification:
diarize:auto(default) ortruereturnsspeaker_turnswhen timestamps are requested.falseomitsspeaker_turns. It does not change the transcript, and the speaker markers stay intext.num_speakers: the expected number of speakers. This is a hint, applied on a best-effort basis; it guides speaker identification on longer recordings and may have no effect on short ones. It cannot be combined withmin_speakersormax_speakers.min_speakers,max_speakers: bounds on the number of speakers, with the same best-effort rule.min_speakersmust not exceedmax_speakers.
diarize=false. Invalid values return 400 with the code invalid_parameter.
For example, to transcribe an interview with two speakers and print each turn:
Emotion and vocal-event cues
Pro can retain cues about how speech sounds alongside the spoken words. These appear as inline bracketed text, such as[高兴] (happy) or [laughter], in the response’s text field and in the turn text of speaker_turns. For example, a transcript might contain:
emotion field or fixed emotion enum; preserve the returned text when your application needs these cues.
To get a transcript without cues, send tag_audio_events=false (transcribe-1-pro). Cues are then removed from text and speaker_turns; timestamps and billing do not change.
Word timestamps ignore speaker markers and cues. Read text for the annotated transcript and segments for timed speech; the segment text may therefore differ from the full transcript.
Implementation details
Language
language is optional. The language is detected automatically whether or not you send a hint, and a hint does not force the transcript into that language. With transcribe-1-pro, the hint does not change the transcript; it matters only when the language cannot be determined, as described below.
Use a lowercase ISO 639-1 code such as en, zh, or ja. Other forms, such as en-US or English, may be rejected with 400.
The response reports the language in two fields, which are omitted when unknown (never null):
language_code: the two-letter ISO 639-1 code, such asen. Use it in application logic.language: the language’s English name, such asEnglishorChinese. Use it for display.
language_code reports your hint, which is not checked against the audio, and language may be absent. Responses report one language: for recordings that switch languages, language and language_code do not list every language spoken. The Python SDK does not return these fields; call the API directly to read them.
Input audio
transcribe-1-pro accepts WAV, MP3, AAC (including M4A/MP4), FLAC, Ogg (Opus or Vorbis), WebM/Matroska, and MOV, including browser recordings, and uses the first audio track of a video file. Send the original file bytes; no conversion is needed.
Not supported: AIFF, CAF, WMA, AMR, AC-3, and raw (headerless) PCM. These return 400.
Send audio as multipart/form-data (a file upload, shown above) or application/msgpack (see Direct API). Base64-encoded audio in a JSON body is not supported.
If you use transcribe-1, send WAV, MP3, AAC (including M4A/MP4), FLAC, or Ogg (Opus or Vorbis), and convert WebM recordings (for example, from a browser’s MediaRecorder) to Ogg/Opus, MP3, or WAV first, or use transcribe-1-pro.
Limits
These limits apply totranscribe-1-pro:
- Recordings up to 60 minutes long. Longer audio returns 400
audio_too_long. - For long recordings, send compressed audio such as MP3, Opus, or AAC. An hour of 128 kbps MP3 is about 55 MiB. A request that is too large returns 413.
- Very short clips (under about 0.08 seconds) return 400
audio_too_short.
transcribe-1, send up to 50 MiB per request and keep MP3 and Opus files under 25 MiB; a request that exceeds the size limit returns 413 or 400. transcribe-1 is designed for short recordings; for recordings longer than a few minutes, use transcribe-1-pro. A long request can fail with 503 if it exceeds the processing-time limit. Retry, and if it keeps failing, use transcribe-1-pro or split the audio.
Long recordings
Withtranscribe-1-pro, send the whole recording, up to 60 minutes, in one request. Timestamps and speaker labels cover the whole recording. Do not split conversations: speaker labels are consistent only within one response. Long recordings can take several minutes, so raise your client timeout and retry on 5xx errors. You are billed for the full duration of the recording once.
If you use transcribe-1, keep each request short. For a long recording, use transcribe-1-pro, or split the audio into shorter clips, transcribe each, and offset each clip’s start/end by where it began in the full recording. Check that each response covers the complete clip before combining transcripts.
Processing time and timeouts
Each request returns only when the whole file has been transcribed. Processing time grows with the length of the recording, and timestamps add to it. Longtranscribe-1-pro recordings can take several minutes.
Set your HTTP client’s timeout accordingly. The official SDKs’ default timeout can be too short for long recordings, and many HTTP libraries default to even less. To wait up to 15 minutes, as several examples on this page do:
- Python SDK:
RequestOptions(timeout=900)for one request, orFishAudio(timeout=900)for the client. - httpx:
timeout=httpx.Timeout(900.0, connect=10.0). Without it, httpx waits only 5 seconds. - Node.js: the built-in
fetchstops waiting for a response after 5 minutes, even if you pass a longerAbortSignal. For longer requests, usefetchfrom theundicipackage with a dispatcher that waits longer.undici7 runs on Node.js 20.18.1 and later;undici8 requires Node.js 22.19 or later.
Request IDs
transcribe-1-pro returns a unique request_id in the response body, both on success and on its own errors, and the same value in the x-request-id response header. Platform and network-edge errors do not carry it. Log it, and include it when you contact support. The Python SDK does not return it on success; read it with a direct API call, or from the error body as shown in Errors.
Errors
The error body depends on where the error comes from:transcribe-1-proerrors are JSON withstatus,message,code, andrequest_id. Branch oncode, not onmessage.- Platform errors, such as a missing or invalid API key (401), insufficient API credit (402), or the concurrency limit (429), are JSON with
statusandmessage. - Errors from the network edge in front of the API, such as some 413 and 5xx responses, may not have a JSON body.
Retry 429 and 5xx responses with exponential backoff; 429 responses have no
Retry-After header. Do not retry other 4xx responses; change the request first. New code values may be added; handle unknown codes by HTTP status. See the API reference for every status and code.
With the Python SDK, a failed request raises APIError (RateLimitError for 429, ServerError for 5xx). e.status is the HTTP status, and e.body is the raw response body, where transcribe-1-pro errors carry code and request_id:
transcribe-1, its errors are JSON with status and message; rely only on these two fields.
Async transcription
The Python SDK ships an async client with the same surface, useful when you’re transcribing many files concurrently or already running inside an event loop. UseAsyncFishAudio and await the call:
Direct API (MessagePack)
POST /v1/asr also accepts a MessagePack body instead of multipart form data, the same path the API reference links to for low-overhead, server-side calls. Calling the API directly also gives you the response fields the Python SDK does not return. Pack the audio bytes and options into one payload and set Content-Type: application/msgpack:
text, duration (seconds), segments, and request_id, plus speaker_turns when timestamps are requested, and language_code and language when the language is known. Model selection stays in the HTTP header for both formats. MessagePack values are typed: send booleans for ignore_timestamps and tag_audio_events, integers for speaker counts, and a boolean or "auto" for diarize.
Going further
Generate speech
The reverse direction: text to lifelike audio.
Full API parameters
Every field and the raw response schema.
Python reference
asr.transcribe options and the ASRResponse type.
