Skip to content

Speech to text ​

POST /v1/audio/transcriptions

OpenAI-compatible multipart upload. Billed per second of input audio.

python
with open("meeting.m4a", "rb") as audio:
    result = client.audio.transcriptions.create(model="xvoice-stt1", file=audio)

print(result.text)
print(result.usage.seconds, "seconds billed")

For audio you are still capturing, use realtime instead — this endpoint needs the whole file up front.

Parameters ​

FieldTypeDefaultNotes
filefilerequiredThe audio. At most 25 MB and 600 seconds
modelstringrequiredAn STT model id
languagestring—ISO-639-1. Omit to auto-detect
promptstring—Spellings or context, at most 1000 characters
temperaturenumber—0.0 to 1.0
response_formatstringjsonjson or text. Ignored when streaming
streamboolfalseDeliver the transcript as it is produced

Accepted containers: flac, m4a, mp3, mp4, ogg, wav, webm.

language is not only a hint. The model writes the transcript in the language you name, so passing en over German speech gets you a translation rather than a transcript. Pass the language actually spoken, or leave it out. A language the model does not support is unsupported_language, with the supported list in the message.

Responses ​

json (the default):

json
{"text": "...", "usage": {"type": "duration", "seconds": 12.5}}

text returns the transcript as text/plain with no envelope.

seconds is the duration of the audio we measured, which is what you are billed for — not the length of the transcript and not wall-clock time.

Streaming ​

stream=true sends the transcript as server-sent events while it is produced. The file still uploads in full first: this shortens the wait for the text, not the wait for the upload. Worth it for long recordings, pointless for short ones.

python
with open("long-recording.mp3", "rb") as audio:
    for event in client.audio.transcriptions.create(
        model="xvoice-stt1", file=audio, stream=True
    ):
        if event.type == "transcript.text.delta":
            print(event.delta, end="", flush=True)
        elif event.type == "transcript.text.done":
            print(f"\n[{event.usage.seconds}s billed]")

Two events:

text
data: {"type":"transcript.text.delta","delta":"Hello"}

data: {"type":"transcript.text.done","text":"Hello there.","usage":{"type":"duration","seconds":2.0}}

Deltas are appended, never revised. transcript.text.done carries the whole transcript, so you can ignore the deltas entirely if you only want the final text.

response_format is ignored when streaming — the events are the format.

A failure after the stream has begun arrives as an error event with the usual envelope, which the openai client raises as an APIError from the loop. A stream that simply stops carries no error: that is a dropped connection.

Limits ​

LimitValue
File size25 MB
Duration600 seconds (10 minutes)
prompt1000 characters

Size is checked against the declared Content-Length before the body is read, so an oversized upload is refused without being pulled into memory. Duration is checked after probing the file, which is the first moment it is knowable.

Both are per request. Splitting a long recording into 10-minute parts is the intended way to handle it; each part bills for its own audio.

What a request costs ​

Per second of audio, measured by us from the file. A request that fails before the engine is reached bills nothing. Once text has been delivered it bills, including when you disconnect mid-stream — see What is billable.

Transcription charges the rate limiter twice: once for the request, then again for the audio seconds, which only become known after the file is probed. A 429 on the second charge means you are within your request rate but over your audio rate.

Errors ​

CodeStatusCause
invalid_model400Not an STT model, or no such id
unsupported_language400The model does not write that language
file_too_large400Over 25 MB
audio_too_long400Over 600 seconds
invalid_audio400Unreadable, unsupported container, or no audio stream
model_unavailable404The model has no deployed version
inference_unavailable502We could not reach the engine
stream_interrupted—In-band, mid-stream: the engine stopped

Runnable: stt.py and stt_stream.py.

Built on the OpenAI audio API surface.