Skip to content

Realtime transcription ​

WS /v1/realtime?intent=transcription

A live session: you push audio as it is captured and transcripts come back while the speaker is still talking. This is the endpoint for phone calls, meetings and browser microphones. For audio you already have on disk, use transcriptions — it is simpler and costs the same.

The protocol mirrors OpenAI's realtime transcription API, so their clients work:

python
from openai import AsyncOpenAI

client = AsyncOpenAI(api_key=..., base_url="https://your-xvoice-host/v1")

async with client.realtime.connect(
    extra_query={"intent": "transcription"}
) as connection:
    ...

A complete working client is on the example scripts page.

Four things to know first ​

  1. There is no voice activity detection. The server never decides that an utterance has ended. Nothing comes back until you send input_audio_buffer.commit, and choosing when is your client's job — a local VAD, a push-to-talk button, or a timer. turn_detection must be null; any other value is an error rather than a silent no-op, so you find this out at development time instead of wondering why no transcript arrives.
  2. The audio format is fixed: 24 kHz mono PCM16, little endian, base64 inside JSON. No container, no header, nothing to negotiate. Binary WebSocket frames are rejected.
  3. Not every model. The model must have realtime in its capabilities — see Models. Realtime is a separate deployment from the file endpoint, not a mode of it.
  4. Sessions end. 55 minutes maximum, and session.created tells you exactly when. Plan a reconnect rather than discovering it.

Connecting ​

  • ?intent=transcription is required. It selects this behaviour, exactly as it does against OpenAI.
  • Authentication is the Authorization: Bearer header. There is no key-in-URL form, because a query string ends up in proxy logs and browser history. In a browser, proxy the connection through your own server rather than shipping a key.
  • The model is named in session.update, not on the URL. The OpenAI SDK sends no model query parameter for a transcription session, so the session opens before anyone knows what will run it. (A ?model= on the URL is accepted as a fallback for hand-written clients.)

A refusal happens before the upgrade and is an ordinary HTTP response carrying the usual error envelope — 401 for a bad key, 402 for an empty balance, 429 when you are at your concurrency limit. It is not a WebSocket close, so read the body.

On success the first event is session.created.

The exchange ​

text
client → session.update              name the model and language
server → session.updated             the effective configuration, echoed back
client → input_audio_buffer.append   audio, repeatedly
server → ...input_audio_transcription.delta      text, while you are still talking
client → input_audio_buffer.commit   this utterance is over
server → input_audio_buffer.committed
server → conversation.item.added / .done
server → ...input_audio_transcription.delta      whatever was left
server → ...input_audio_transcription.completed  the settled transcript

Audio before a session.update is session_not_configured.

Client events ​

Four, and nothing else is accepted. Unknown fields are errors, not ignored — a setting you spelled wrong tells you so.

session.update — configure the session. Only the first one attaches the model; naming a different model later is model_already_set, because one session is one engine. language and prompt can be changed later and take effect on the next utterance, never on one already being decoded.

json
{
  "type": "session.update",
  "session": {
    "type": "transcription",
    "audio": {
      "input": {
        "format": {"type": "audio/pcm", "rate": 24000},
        "transcription": {"model": "xvoice-stt1", "language": "en", "prompt": ""},
        "turn_detection": null
      }
    }
  }
}

input_audio_buffer.append — {"type": ..., "audio": "<base64 PCM16>"}. At most 15 MB per event, checked on the encoded length before decoding. A frame must be an even number of bytes; 20 ms (960 bytes) is a good size, but any even length works.

input_audio_buffer.commit — end the utterance. The input buffer is freed at once and the engine keeps decoding in the background, so you can start the next utterance immediately. Committing nothing is input_audio_buffer_commit_empty.

input_audio_buffer.clear — discard the uncommitted buffer. It does not stop an utterance already committed, and cleared audio is still billed: it has already reached the engine.

Every client event may carry an event_id of up to 512 characters, echoed on any error it causes.

Server events ​

session.created, session.updated, input_audio_buffer.committed, input_audio_buffer.cleared, conversation.item.added, conversation.item.done, conversation.item.input_audio_transcription.delta / .completed / .failed, and error.

Deltas arrive before the commit. Text streams while the buffer is still open, which is the point of a streaming session: you do not wait for the end of the turn to show the beginning of it. They carry the item_id of the utterance being spoken, before the input_audio_buffer.committed that announces it — so bind the id from the first delta you see rather than from the acknowledgement.

Deltas are immutable. Text is never revised once sent, so you can render it straight away, and concatenating every delta for an item gives exactly the final transcript. .completed carries the whole transcript, which is authoritative, and our measurement of the audio:

json
{
  "type": "conversation.item.input_audio_transcription.completed",
  "item_id": "item_01J8...",
  "transcript": "Hello, this is a round trip through xVoice.",
  "usage": {"type": "duration", "seconds": 3.0}
}

Utterances are concurrent and finish out of order. A long turn can complete after a short one that followed it, so key everything by item_id — never assume there is a single current utterance. The ids are ours (item_01J8…), stable across the session and the ones that appear in your request logs.

The session deadline ​

session.created carries expires_at, a Unix timestamp 55 minutes out:

json
{
  "type": "session.created",
  "session": {
    "id": "sess_01J8...",
    "object": "realtime.transcription_session",
    "type": "transcription",
    "expires_at": 1758675300,
    "audio": {"input": {"format": {"type": "audio/pcm", "rate": 24000},
                        "transcription": {"model": "xvoice-stt1", "language": "en", "prompt": ""},
                        "turn_detection": null,
                        "noise_reduction": null}},
    "include": []
  }
}

At the deadline you get a session_expired error event and a normal close. It is published rather than left to be discovered so that a long call can open the next session before the old one ends and lose nothing.

A session that sends no audio for 5 minutes is closed with session_idle.

Billing ​

Per second of audio you send, from our own count of the PCM bytes: bytes / 48000. The engine's own figure is not used, because it arrives on a completion event that a client who disconnects mid-utterance never produces.

That means all of this bills:

  • audio in an utterance you clear,
  • audio in flight when you disconnect,
  • a session where the transcript never came back.

All of it reached the engine and cost GPU time. The file endpoint works the same way.

The balance is re-checked every 60 seconds of audio rather than only at the end of an utterance, so a session cannot run up an unpayable bill inside one very long turn. Running out mid-session is an insufficient_credits error event and a close.

A session also holds one concurrency slot for its whole life, and costs one request against your per-minute rate limit when it opens. See Rate limits and billing.

Errors ​

An error after the upgrade is an error event on the socket, carrying the same envelope as everywhere else plus the event_id of whatever caused it:

json
{
  "type": "error",
  "event_id": "event_01J8...",
  "error": {
    "type": "invalid_request",
    "code": "unsupported_parameter",
    "message": "Server-side turn detection is not available. Send 'input_audio_buffer.commit' when an utterance ends.",
    "param": "session.audio.input.turn_detection",
    "request_id": "req_01J8...",
    "client_event_id": "event_from_you"
  }
}

Recoverable problems leave the session open — fix the event and carry on. These close it: billing, rate limits, the deadline, the idle timeout, and anything on our side. The close code never carries the reason; the error event before it does.

Close codeMeaning
1000Normal — deadline reached, or idle
1008Something you can fix: billing, limits, a fatal protocol error
1011Our failure — the engine, or a misconfiguration on our side
CodeCause
session_not_configuredAudio before session.update
missing_modelNo model named, and none on the URL
invalid_modelNo such model, or not an STT model
realtime_not_supportedThe model has no realtime capability
model_already_setA second session.update naming a different model
unsupported_eventNot one of the four client events — or a binary frame
unsupported_parameterturn_detection, noise_reduction or include
unsupported_audio_formatNot audio/pcm, or not 24 kHz
frame_too_largeAn append over 15 MB
audio_too_longOne utterance over 600 seconds — commit more often
input_audio_buffer_commit_emptycommit with nothing buffered
unsupported_languageThe model does not write that language
session_expired55 minutes reached
session_idle5 minutes with no audio
insufficient_creditsBalance ran out mid-session
transcription_failedThe engine could not transcribe one utterance (the session survives)
inference_unavailableWe could not open or hold a session with the engine

Deploying behind a proxy ​

A WebSocket that carries audio for an hour is an unusual shape for a reverse proxy, and the default settings are wrong for it. On nginx-ingress, proxy-read-timeout and proxy-send-timeout default to 60 seconds, which would drop a session long before its own deadline. Set both above 3300 seconds on this path and allow the upgrade. These are transport settings and have nothing to do with session lifetime, which the server owns.

Not supported ​

Server-side VAD and turn_detection, ephemeral client keys, WebRTC, logprobs, include, noise reduction, and text-to-speech over the same socket. This endpoint transcribes; it is not a speech-to-speech API.

Built on the OpenAI audio API surface.