Appearance
Speech to text
POST /v1/audio/transcriptions
OpenAI-compatible multipart upload. Billed per second of input audio.
python
with open("meeting.m4a", "rb") as audio:
result = client.audio.transcriptions.create(model="xvoice-stt1", file=audio)
print(result.text)
print(result.usage.seconds, "seconds billed")For audio you are still capturing, use realtime instead — this endpoint needs the whole file up front.
Parameters
| Field | Type | Default | Notes |
|---|---|---|---|
file | file | required | The audio. At most 25 MB and 600 seconds |
model | string | required | An STT model id |
language | string | — | ISO-639-1. Omit to auto-detect |
prompt | string | — | Spellings or context, at most 1000 characters |
temperature | number | — | 0.0 to 1.0 |
response_format | string | json | json or text. Ignored when streaming |
stream | bool | false | Deliver the transcript as it is produced |
Accepted containers: flac, m4a, mp3, mp4, ogg, wav, webm.
language is not only a hint. The model writes the transcript in the language you name, so passing en over German speech gets you a translation rather than a transcript. Pass the language actually spoken, or leave it out. A language the model does not support is unsupported_language, with the supported list in the message.
Responses
json (the default):
json
{"text": "...", "usage": {"type": "duration", "seconds": 12.5}}text returns the transcript as text/plain with no envelope.
seconds is the duration of the audio we measured, which is what you are billed for — not the length of the transcript and not wall-clock time.
Streaming
stream=true sends the transcript as server-sent events while it is produced. The file still uploads in full first: this shortens the wait for the text, not the wait for the upload. Worth it for long recordings, pointless for short ones.
python
with open("long-recording.mp3", "rb") as audio:
for event in client.audio.transcriptions.create(
model="xvoice-stt1", file=audio, stream=True
):
if event.type == "transcript.text.delta":
print(event.delta, end="", flush=True)
elif event.type == "transcript.text.done":
print(f"\n[{event.usage.seconds}s billed]")Two events:
text
data: {"type":"transcript.text.delta","delta":"Hello"}
data: {"type":"transcript.text.done","text":"Hello there.","usage":{"type":"duration","seconds":2.0}}Deltas are appended, never revised. transcript.text.done carries the whole transcript, so you can ignore the deltas entirely if you only want the final text.
response_format is ignored when streaming — the events are the format.
A failure after the stream has begun arrives as an error event with the usual envelope, which the openai client raises as an APIError from the loop. A stream that simply stops carries no error: that is a dropped connection.
Limits
| Limit | Value |
|---|---|
| File size | 25 MB |
| Duration | 600 seconds (10 minutes) |
prompt | 1000 characters |
Size is checked against the declared Content-Length before the body is read, so an oversized upload is refused without being pulled into memory. Duration is checked after probing the file, which is the first moment it is knowable.
Both are per request. Splitting a long recording into 10-minute parts is the intended way to handle it; each part bills for its own audio.
What a request costs
Per second of audio, measured by us from the file. A request that fails before the engine is reached bills nothing. Once text has been delivered it bills, including when you disconnect mid-stream — see What is billable.
Transcription charges the rate limiter twice: once for the request, then again for the audio seconds, which only become known after the file is probed. A 429 on the second charge means you are within your request rate but over your audio rate.
Errors
| Code | Status | Cause |
|---|---|---|
invalid_model | 400 | Not an STT model, or no such id |
unsupported_language | 400 | The model does not write that language |
file_too_large | 400 | Over 25 MB |
audio_too_long | 400 | Over 600 seconds |
invalid_audio | 400 | Unreadable, unsupported container, or no audio stream |
model_unavailable | 404 | The model has no deployed version |
inference_unavailable | 502 | We could not reach the engine |
stream_interrupted | — | In-band, mid-stream: the engine stopped |
Runnable: stt.py and stt_stream.py.