Appearance
Realtime transcription
WS /v1/realtime?intent=transcription
A live session: you push audio as it is captured and transcripts come back while the speaker is still talking. This is the endpoint for phone calls, meetings and browser microphones. For audio you already have on disk, use transcriptions — it is simpler and costs the same.
The protocol mirrors OpenAI's realtime transcription API, so their clients work:
python
from openai import AsyncOpenAI
client = AsyncOpenAI(api_key=..., base_url="https://your-xvoice-host/v1")
async with client.realtime.connect(
extra_query={"intent": "transcription"}
) as connection:
...A complete working client is on the example scripts page.
Four things to know first
- There is no voice activity detection. The server never decides that an utterance has ended. Nothing comes back until you send
input_audio_buffer.commit, and choosing when is your client's job — a local VAD, a push-to-talk button, or a timer.turn_detectionmust benull; any other value is an error rather than a silent no-op, so you find this out at development time instead of wondering why no transcript arrives. - The audio format is fixed: 24 kHz mono PCM16, little endian, base64 inside JSON. No container, no header, nothing to negotiate. Binary WebSocket frames are rejected.
- Not every model. The model must have
realtimein itscapabilities— see Models. Realtime is a separate deployment from the file endpoint, not a mode of it. - Sessions end. 55 minutes maximum, and
session.createdtells you exactly when. Plan a reconnect rather than discovering it.
Connecting
?intent=transcriptionis required. It selects this behaviour, exactly as it does against OpenAI.- Authentication is the
Authorization: Bearerheader. There is no key-in-URL form, because a query string ends up in proxy logs and browser history. In a browser, proxy the connection through your own server rather than shipping a key. - The model is named in
session.update, not on the URL. The OpenAI SDK sends nomodelquery parameter for a transcription session, so the session opens before anyone knows what will run it. (A?model=on the URL is accepted as a fallback for hand-written clients.)
A refusal happens before the upgrade and is an ordinary HTTP response carrying the usual error envelope — 401 for a bad key, 402 for an empty balance, 429 when you are at your concurrency limit. It is not a WebSocket close, so read the body.
On success the first event is session.created.
The exchange
text
client → session.update name the model and language
server → session.updated the effective configuration, echoed back
client → input_audio_buffer.append audio, repeatedly
server → ...input_audio_transcription.delta text, while you are still talking
client → input_audio_buffer.commit this utterance is over
server → input_audio_buffer.committed
server → conversation.item.added / .done
server → ...input_audio_transcription.delta whatever was left
server → ...input_audio_transcription.completed the settled transcriptAudio before a session.update is session_not_configured.
Client events
Four, and nothing else is accepted. Unknown fields are errors, not ignored — a setting you spelled wrong tells you so.
session.update — configure the session. Only the first one attaches the model; naming a different model later is model_already_set, because one session is one engine. language and prompt can be changed later and take effect on the next utterance, never on one already being decoded.
json
{
"type": "session.update",
"session": {
"type": "transcription",
"audio": {
"input": {
"format": {"type": "audio/pcm", "rate": 24000},
"transcription": {"model": "xvoice-stt1", "language": "en", "prompt": ""},
"turn_detection": null
}
}
}
}input_audio_buffer.append — {"type": ..., "audio": "<base64 PCM16>"}. At most 15 MB per event, checked on the encoded length before decoding. A frame must be an even number of bytes; 20 ms (960 bytes) is a good size, but any even length works.
input_audio_buffer.commit — end the utterance. The input buffer is freed at once and the engine keeps decoding in the background, so you can start the next utterance immediately. Committing nothing is input_audio_buffer_commit_empty.
input_audio_buffer.clear — discard the uncommitted buffer. It does not stop an utterance already committed, and cleared audio is still billed: it has already reached the engine.
Every client event may carry an event_id of up to 512 characters, echoed on any error it causes.
Server events
session.created, session.updated, input_audio_buffer.committed, input_audio_buffer.cleared, conversation.item.added, conversation.item.done, conversation.item.input_audio_transcription.delta / .completed / .failed, and error.
Deltas arrive before the commit. Text streams while the buffer is still open, which is the point of a streaming session: you do not wait for the end of the turn to show the beginning of it. They carry the item_id of the utterance being spoken, before the input_audio_buffer.committed that announces it — so bind the id from the first delta you see rather than from the acknowledgement.
Deltas are immutable. Text is never revised once sent, so you can render it straight away, and concatenating every delta for an item gives exactly the final transcript. .completed carries the whole transcript, which is authoritative, and our measurement of the audio:
json
{
"type": "conversation.item.input_audio_transcription.completed",
"item_id": "item_01J8...",
"transcript": "Hello, this is a round trip through xVoice.",
"usage": {"type": "duration", "seconds": 3.0}
}Utterances are concurrent and finish out of order. A long turn can complete after a short one that followed it, so key everything by item_id — never assume there is a single current utterance. The ids are ours (item_01J8…), stable across the session and the ones that appear in your request logs.
The session deadline
session.created carries expires_at, a Unix timestamp 55 minutes out:
json
{
"type": "session.created",
"session": {
"id": "sess_01J8...",
"object": "realtime.transcription_session",
"type": "transcription",
"expires_at": 1758675300,
"audio": {"input": {"format": {"type": "audio/pcm", "rate": 24000},
"transcription": {"model": "xvoice-stt1", "language": "en", "prompt": ""},
"turn_detection": null,
"noise_reduction": null}},
"include": []
}
}At the deadline you get a session_expired error event and a normal close. It is published rather than left to be discovered so that a long call can open the next session before the old one ends and lose nothing.
A session that sends no audio for 5 minutes is closed with session_idle.
Billing
Per second of audio you send, from our own count of the PCM bytes: bytes / 48000. The engine's own figure is not used, because it arrives on a completion event that a client who disconnects mid-utterance never produces.
That means all of this bills:
- audio in an utterance you
clear, - audio in flight when you disconnect,
- a session where the transcript never came back.
All of it reached the engine and cost GPU time. The file endpoint works the same way.
The balance is re-checked every 60 seconds of audio rather than only at the end of an utterance, so a session cannot run up an unpayable bill inside one very long turn. Running out mid-session is an insufficient_credits error event and a close.
A session also holds one concurrency slot for its whole life, and costs one request against your per-minute rate limit when it opens. See Rate limits and billing.
Errors
An error after the upgrade is an error event on the socket, carrying the same envelope as everywhere else plus the event_id of whatever caused it:
json
{
"type": "error",
"event_id": "event_01J8...",
"error": {
"type": "invalid_request",
"code": "unsupported_parameter",
"message": "Server-side turn detection is not available. Send 'input_audio_buffer.commit' when an utterance ends.",
"param": "session.audio.input.turn_detection",
"request_id": "req_01J8...",
"client_event_id": "event_from_you"
}
}Recoverable problems leave the session open — fix the event and carry on. These close it: billing, rate limits, the deadline, the idle timeout, and anything on our side. The close code never carries the reason; the error event before it does.
| Close code | Meaning |
|---|---|
| 1000 | Normal — deadline reached, or idle |
| 1008 | Something you can fix: billing, limits, a fatal protocol error |
| 1011 | Our failure — the engine, or a misconfiguration on our side |
| Code | Cause |
|---|---|
session_not_configured | Audio before session.update |
missing_model | No model named, and none on the URL |
invalid_model | No such model, or not an STT model |
realtime_not_supported | The model has no realtime capability |
model_already_set | A second session.update naming a different model |
unsupported_event | Not one of the four client events — or a binary frame |
unsupported_parameter | turn_detection, noise_reduction or include |
unsupported_audio_format | Not audio/pcm, or not 24 kHz |
frame_too_large | An append over 15 MB |
audio_too_long | One utterance over 600 seconds — commit more often |
input_audio_buffer_commit_empty | commit with nothing buffered |
unsupported_language | The model does not write that language |
session_expired | 55 minutes reached |
session_idle | 5 minutes with no audio |
insufficient_credits | Balance ran out mid-session |
transcription_failed | The engine could not transcribe one utterance (the session survives) |
inference_unavailable | We could not open or hold a session with the engine |
Deploying behind a proxy
A WebSocket that carries audio for an hour is an unusual shape for a reverse proxy, and the default settings are wrong for it. On nginx-ingress, proxy-read-timeout and proxy-send-timeout default to 60 seconds, which would drop a session long before its own deadline. Set both above 3300 seconds on this path and allow the upgrade. These are transport settings and have nothing to do with session lifetime, which the server owns.
Not supported
Server-side VAD and turn_detection, ephemeral client keys, WebRTC, logprobs, include, noise reduction, and text-to-speech over the same socket. This endpoint transcribes; it is not a speech-to-speech API.