Skip to content

xVoiceSpeech APIs that speak OpenAI

Text to speech, file transcription and live realtime sessions behind one HTTP API — with scoped keys, per-second billing and request logs.

Getting started ​

xVoice implements the OpenAI audio API, so the official clients work against it unchanged:

python
from openai import OpenAI

client = OpenAI(api_key="sk_live_...", base_url="https://your-xvoice-host/v1")

Throughout these pages the base URL is written $XVOICE_BASE_URL. It is your deployment's host with /v1 on the end.

PageWhat it covers
QuickstartSignup to a working transcript, in about five minutes
AuthenticationAPI keys, test and live environments, key hygiene
ErrorsThe error envelope, every code, and what to do about each
Models and voicesWhat is available, and what capabilities means
Rate limits and billingTiers, headers, credits, spending limits

Endpoints ​

EndpointPurpose
POST /v1/audio/speechText to speech, whole file or streamed
POST /v1/audio/transcriptionsTranscribe an audio file
WS /v1/realtimeTranscribe a live stream as it is captured
GET /v1/modelsList models, or read one
GET /v1/voicesList voices
GET /v1/meWhat a key is attached to

/docs on your deployment serves the generated OpenAPI schema for these endpoints. It cannot describe WS /v1/realtime — OpenAPI has no vocabulary for WebSockets — so that one is documented only here.

Runnable code ​

Every endpoint has a working script on the example scripts page, built on the stock openai Python client:

bash
export XVOICE_API_KEY=sk_test_...
export XVOICE_BASE_URL=https://your-xvoice-host/v1

python tts.py "Hello from xVoice."     # writes output/speech.wav
python stt.py output/speech.wav
python realtime.py output/speech.wav

Conventions ​

  • Ids are prefixed and sortable: org_…, proj_…, key_…, req_…. The prefix says what the thing is, which makes a misplaced id obvious in a log.
  • Every response carries X-Request-Id, and every error body repeats it as request_id. It is the one thing worth keeping when something goes wrong: it ties a report to the exact request.
  • Lists return {"object": "list", "data": [...]}. Where a list can grow without bound it is keyset-paginated with an opaque next_cursor.
  • Unknown fields are not silently ignored on the realtime endpoint. A setting you spelled wrong is an error there, not a no-op.

Built on the OpenAI audio API surface.