Skip to main content
POST
Speech to Text
This BETA endpoint transcribes one audio file per request. Send the audio as multipart/form-data (a file upload) or application/msgpack. Base64-encoded audio in a JSON body is not supported.

Select a model

Send POST https://api.fish.audio/v1/asr with Authorization: Bearer <API_KEY> and a model HTTP header that selects the model. The request is synchronous: the response arrives when the whole file has been transcribed. Send model: transcribe-1-pro with every request, written exactly, in lowercase. If the header is missing, or its value is not an exact match (for example Transcribe-1-Pro or transcribe-1pro), the request is served and billed as transcribe-1, and no error is returned. If you expected transcribe-1-pro but the transcript has no speaker markers, check the header. This applies to the API playground on this page too: set its model header to transcribe-1-pro. The model belongs in the header for both multipart and MessagePack requests; a model field in the request body is ignored. See Limits for recording length and request size. To link the request to your own distributed trace, also send a W3C traceparent header. See Tracing & Performance Analysis.

Request fields

Multipart and MessagePack bodies use the same field names. Multipart values are text; MessagePack values are typed. In multipart forms, send true or false for ignore_timestamps. Any value other than true (case-insensitive), including 1 or an empty value, is read as false and turns timestamps on. In MessagePack, send a boolean; a string or nil returns 400. Speaker-count hints are applied on a best-effort basis: they guide speaker identification on longer recordings and may have no effect on short ones. They cannot be sent with diarize=false. tag_audio_events, diarize, and the speaker-count hints are validated strictly:
  • In multipart forms, tag_audio_events takes true or false, and diarize takes auto, true, or false, in any letter case. Speaker counts take digits only. Send each field at most once.
  • In MessagePack, send tag_audio_events as a boolean, diarize as a boolean or one of the lowercase strings auto, true, or false, and speaker counts as integers, not floats or strings. nil means the default.
  • Any other value, a repeated multipart field, or a conflicting combination returns 400 with code invalid_parameter.
If you use transcribe-1: send audio, language, and ignore_timestamps. The other fields apply to transcribe-1-pro; send them only with model: transcribe-1-pro.

Examples

This request returns word-level segments and speaker_turns. For a two-person interview, you can add a speaker-count hint, and remove emotion and vocal-event cues from the transcript:
Let your HTTP client set the multipart Content-Type and boundary. For MessagePack, set Content-Type: application/msgpack and encode audio as binary bytes. See the Speech to Text guide for an example.

Read the response

A successful response (HTTP 200) is a JSON object: A field that is not returned is absent from the JSON, never null. Do not depend on the order of keys in the JSON. If you use transcribe-1: rely on text, duration, segments, language, and language_code. Speaker markers, speaker_turns, and request_id are transcribe-1-pro features.
Both SDKs select transcribe-1-pro through the model header. In Python, pass request_options=RequestOptions(additional_headers={"model": "transcribe-1-pro"}) to asr.transcribe(). In JavaScript, pass { headers: { model: "transcribe-1-pro" } } as the second argument to speechToText.convert(). The Python SDK currently returns only text, duration, and segments, and neither SDK can send tag_audio_events, diarize, or the speaker-count hints. To use them, call the API directly, as in the examples above. The JavaScript SDK returns the whole response body, but its STTResponse type declares only text, duration, and segments.

Multi-speaker response format

transcribe-1-pro marks speaker changes inside the text string with <|speaker:N|> markers. The label N identifies the speaker for the text that follows, up to the next marker. A repeated label means that speaker is speaking again.
  • Markers usually have a space on each side (<|speaker:0|> 你好。 <|speaker:1|> ...). Do not depend on exact spacing.
  • Any text before the first marker belongs to the first turn.
  • A transcript with no marker comes from a single speaker; treat it as speaker 0.
  • Labels identify speakers within one response only. They are not names, and the same label in two requests is not the same person. Send a whole conversation as one file.
This illustrative response has two speakers and three turns. With ignore_timestamps=true (the default), timestamps are skipped and speaker_turns is absent:
Read this as speaker 0 saying 你好。, speaker 1 saying [高兴]很开心认识你。, and speaker 0 saying 我也是。. The [高兴] cue describes the delivery of speaker 1’s speech. There is no separate emotion field; the cues are part of the text.

Speaker turns

With ignore_timestamps=false, transcribe-1-pro also returns the turns as a structured speaker_turns array, unless you send diarize=false. When you request timestamps, prefer speaker_turns over parsing text.
  • speaker_turns lists the turns in the order they occur, and is [] when no speech was found. Each turn is { "speaker": "speaker:N", "text": "...", "start": <seconds>, "end": <seconds> }.
  • speaker is the string speaker:N, where N matches the <|speaker:N|> marker in text. Like the markers, labels identify speakers within one response only.
  • text is that turn’s speech without speaker markers. It keeps emotion and vocal-event cues unless tag_audio_events=false.
  • start and end come from the word timestamps. If word timing is unavailable (segments is empty), turn times are approximate and can cover the whole recording.
  • Consecutive turns can have the same speaker. Do not assume turns are contiguous or non-overlapping.
This illustrative response is the same conversation with ignore_timestamps=false:
See the Speech to Text guide for parsing examples.

Segments

A segment is usually one word; in Chinese and Japanese it is usually one character, or a few. Segment text has no punctuation and no speaker markers or cues, and can be normalized (for example 35 for 3.5), so it does not always match text character for character. A segment can have start equal to end. Segments are not speaker turns and carry no speaker ID.

Language

  • language is optional. The language is detected automatically whether or not you send a hint, and a hint does not force the transcript into that language.
  • Use a lowercase ISO 639-1 code such as en, zh, or ja. Other forms, such as en-US or English, may be rejected with 400.
  • If the language cannot be determined (for example, very short audio), language_code reports your hint, which is not checked against the audio, and language may be absent.
  • Responses report one language. For recordings that switch languages, language and language_code do not list every language spoken.

Limits

  • Audio length: up to 60 minutes per request. Longer audio returns 400 with code audio_too_long.
  • Request size: for long recordings, send compressed audio such as MP3, Opus, or AAC. An hour of 128 kbps MP3 is about 55 MiB. A request that is too large returns 413.
  • One conversation per request: send the whole recording in one request. Timestamps and speaker labels then cover the whole recording. Do not split a conversation, because labels are consistent only within one response.
If you use transcribe-1: it is designed for short recordings; for recordings longer than a few minutes, use transcribe-1-pro. It accepts up to 50 MiB per request; keep MP3 and Opus files under 25 MiB. A request that exceeds the size limit returns 413 or 400. A long request can fail with 503 if it exceeds the processing-time limit. Retry, and if it keeps failing, use transcribe-1-pro or split the audio.

Processing time and timeouts

  • Processing time grows with the length of the recording. Long transcribe-1-pro recordings can take several minutes.
  • Set your HTTP client’s timeout accordingly. The official SDKs’ default timeout can be too short for long recordings, and many HTTP libraries default to even less (httpx defaults to 5 seconds). For long recordings, use a generous timeout, for example 15 minutes: httpx.Timeout(900.0, connect=10.0) with httpx, RequestOptions(timeout=900) with the Python SDK, or { timeoutInSeconds: 900 } with the JavaScript SDK.
  • In Node.js, the built-in fetch, which the JavaScript SDK uses, also stops waiting for a response after 5 minutes, even if you set a longer timeout. To wait longer, call setGlobalDispatcher(new Agent({ headersTimeout: 900_000, bodyTimeout: 900_000 })) from the undici package once at startup. See Processing time and timeouts.
  • If a long request fails with a 5xx error or the connection drops, retry it.
  • Each request holds one of your account’s concurrent request slots until its response is returned.

Supported audio formats

  • transcribe-1-pro accepts WAV, MP3, AAC (including M4A/MP4), FLAC, Ogg (Opus or Vorbis), WebM/Matroska, and MOV, including browser recordings. For a video file, it uses the first audio track. Send the original file bytes; no conversion is needed.
  • Not supported: AIFF, CAF, WMA, AMR, AC-3, and raw (headerless) PCM. These return 400.
If you use transcribe-1: it accepts WAV, MP3, AAC (including M4A/MP4), FLAC, and Ogg (Opus or Vorbis). Browser WebM recordings (for example, from MediaRecorder) may not be accepted; convert them to Ogg/Opus, MP3, or WAV first.

Errors

Error responses have one of these shapes:
  • transcribe-1-pro errors are JSON with status, message, code, and request_id. Branch on code, not on message.
  • Platform errors, such as authentication, credit, and concurrency errors, are returned before the request reaches a model. They are JSON with exactly status and message.
  • Errors from the network edge in front of the API, such as some 413 and 5xx responses, may not have a JSON body. Handle them by HTTP status.
This illustrative transcribe-1-pro error body shows the shape:
  • A dash means the error has no code; a 503 without a code also means the service is temporarily unavailable.
  • New code values may be added; handle unknown codes by HTTP status.
  • Retry 429 and 5xx responses with exponential backoff. Do not retry other 4xx responses; change the request first.
  • Requests that return an error response are not billed.
  • With the Python SDK, the raw error body is in the exception’s body attribute; parse it with json.loads to read code and request_id.
If you use transcribe-1: its errors are JSON with status and message. Rely only on these two fields and on the HTTP status. See Errors for retry and SDK exception examples.

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Headers

model
enum<string>
default:transcribe-1

Specify which speech-to-text model to use.

Available options:
transcribe-1,
transcribe-1-pro

Body

audio
file
required

Audio file to be converted to text

language
string | null

Optional hint. The language is auto-detected regardless; the detected language is returned as language_code.

ignore_timestamps
boolean
default:true

Whether to return precise timestamps in the text, this will increase the latency in audio shorter than 30 seconds

Response

Request fulfilled, document follows

text
string
required
duration
number
required

Duration of the audio in seconds

segments
ASRSegment · object[]
required
language_code
string | null

Detected language as an ISO 639-1 code (e.g. en, ja). Omitted if no language is detected.

language
string | null

Detected language name (e.g. English). For display only; use language_code in code.