Appearance
Errors
Every failure, on every endpoint, has the same body:
json
{
"error": {
"type": "invalid_request",
"code": "audio_too_long",
"message": "Audio must be at most 600 seconds long.",
"request_id": "req_01J8XQ2K7Z4N9P0R1S2T3V4W5X"
}
}typegroups the failure. Useful for deciding whether to retry at all.codeidentifies it exactly. Branch on this, never onmessage.messageis written for a human reading a log. It may change; it is not an API.request_idis the same value as theX-Request-Idresponse header, and it is what lets us find the exact request in our logs. Log it on every failure.
Schema validation failures add a details array holding the offending fields, in pydantic's shape. Nothing else sets it.
The OpenAI clients raise their usual exceptions (BadRequestError, AuthenticationError, RateLimitError, …) and put this body on exc.body:
python
from openai import APIStatusError
try:
...
except APIStatusError as exc:
error = (exc.body or {}).get("error", {})
log.error("xvoice %s: %s (request_id=%s)",
error.get("code"), error.get("message"), error.get("request_id"))Status codes
| Status | type | What it means |
|---|---|---|
| 400 | invalid_request | The request is wrong. Retrying it unchanged will fail again |
| 401 | authentication_error | The key is missing, unknown, revoked or expired |
| 402 | billing_error | Out of credit, or over the monthly spending limit |
| 403 | permission_error | The key or user may not do this |
| 404 | not_found | No such model, voice or resource |
| 409 | conflict | The change collides with existing state |
| 422 | invalid_request | The body failed schema validation. Code is validation_error; details names the fields |
| 429 | rate_limit_error | Too many requests, or too many at once. Honour retry-after |
| 499 | — | Log-only. You disconnected mid-stream; nothing is sent |
| 500 | api_error | Something we did not foresee. Code is internal_error — send us the request id |
| 502 | api_error | The speech or transcription engine failed |
| 503 | api_error | A dependency is unavailable |
Codes you will actually meet
Authentication — missing_api_key, invalid_api_key, api_key_revoked, api_key_expired. See Authentication.
Billing — insufficient_credits (balance is empty), spending_limit_reached (the organization's monthly cap). Both are 402 and both are fixed in the dashboard, not in code. See Rate limits and billing.
Rate limits — rate_limit_exceeded (a per-minute bucket) and concurrency_limit_exceeded (too many requests in flight at once). Both carry retry-after.
Catalog — invalid_model (no such model, or not one this endpoint can use), model_unavailable (the model exists but no version is deployed), invalid_voice, voice_not_supported (that voice does not belong to that model), unsupported_language, realtime_not_supported (the model has no realtime capability), unsupported_response_format and unsupported_speed (that model does not offer that format or speed — its speech_options on GET /v1/models says what it does).
These are 400, not 422, because the request was well formed: it asked a particular model for something that model does not do, which is only knowable once the id has been resolved against the catalog. The message names the model and lists what it accepts.
Audio in — file_too_large (over 25 MB), audio_too_long (over 600 seconds), invalid_audio (unreadable, or contains no audio stream).
Engine — inference_unavailable (502, we could not reach or open a session with the engine), inference_rejected (the engine refused the request), transcription_failed, stream_interrupted.
Errors after a 200
Streaming makes this unavoidable: once the first byte of a 200 is on the wire, the status can no longer change. So a failure partway through arrives in band rather than as an HTTP status.
- SSE (
stream=trueon transcriptions,stream_format="sse"on speech): anerrorevent carrying the same envelope. The OpenAI clients raise it as anAPIErrorout of the iteration loop, so afor event in stream:still fails loudly rather than ending early and quietly. - Realtime: an
errorevent on the socket, then a close. The close code alone never carries the reason — read the event. See Realtime transcription.
A truncated stream with no error event means the connection dropped, not that the work finished.
Retrying
| Situation | Retry? |
|---|---|
| 429 | Yes, after retry-after, with jitter |
| 502, 503 | Yes, with backoff — bounded, these are not always transient |
| 500 | Once, with backoff. If it repeats, send us the request id |
| 400, 401, 403, 404, 409, 422 | No. Fix the request or the account |
| 402 | No. Add credit or raise the limit first |
The openai client already retries 429 and 5xx with backoff; max_retries=0 turns that off when you want to own the policy.
Retries are not free: a request that reached the engine before failing has already cost GPU time, and a disconnect mid-stream still bills the audio produced. See billing.