Skip to content

Models and voices ​

Models ​

python
for model in client.models.list():
    print(model.id, model.type, model.capabilities)

GET /v1/models returns a superset of OpenAI's model object, so their SDK's model page parses against us unchanged:

json
{
  "id": "xvoice-stt1",
  "object": "model",
  "created": 1758672000,
  "owned_by": "xvoice",
  "name": "xVoice STT v1",
  "type": "stt",
  "description": "Speech recognition.",
  "capabilities": ["batch"],
  "supported_languages": ["en", "de", "fr"],
  "status": "available",
  "public_version": "2026-09-17",
  "speech_options": null
}

GET /v1/models/{model_id} returns one of these, and GET /v1/models?type=stt (or tts) narrows the list.

capabilities ​

Which endpoints and transports will take this model. It is the field to filter on; once you have picked a model, speech_options below is the field that says what you may send it.

CapabilityMeaning
batchAccepts a whole request at once: POST /v1/audio/transcriptions, or a complete POST /v1/audio/speech
streamingCan produce output incrementally over HTTP — SSE, or a chunked audio body
realtimeCan hold a live session: WS /v1/realtime

It is enforced, not descriptive. A model without realtime refuses a session with realtime_not_supported, naming the file endpoint to use instead. Read the list rather than guessing from the name: a model that transcribes files well may have no streaming deployment behind it at all, and those are separate services rather than modes of one.

On a TTS model, batch and streaming are derived rather than declared — streaming appears exactly when speech_options.streamable_formats is non-empty, and batch always does, because every TTS model will return a complete clip. So the two fields cannot disagree about whether a model streams, and the detailed answer (which format streams) is only ever in one place.

type and status ​

type is stt or tts, and it decides which endpoints will take the model. Handing a TTS id to transcriptions is invalid_model, not a 404.

status is available, deprecated or retired. Deprecated models still serve — plan to move. Retired ones are gone from the list. A model that is listed but has no deployed version behind it fails with model_unavailable, which is our problem rather than yours: it means the catalog and the deployment disagree.

versions ​

public_version names the version a call resolves to today. You do not pass it; pinning is not exposed in V1. It is there so that a change in output has a label you can quote when you ask us what moved.

speech_options ​

TTS models carry a speech_options object saying which speech parameters they accept. It is null on STT models, which take none of them.

json
{
  "response_formats": ["pcm", "wav"],
  "streamable_formats": ["pcm"],
  "speed": {"min": 1.0, "max": 1.0},
  "instructions": false
}

Like capabilities, this is enforced and not merely descriptive: asking for a format or speed outside it is unsupported_response_format or unsupported_speed, a 400 naming the model and listing what it does accept, rather than a silent substitution. The API resolves these values from the same place it enforces them, so what you read and what you get cannot drift apart.

instructions: false is the one exception. That field is documented as best-effort guidance, so a model that ignores it says so here and still accepts the request — you are told rather than refused.

Build a picker from this object rather than from a hard-coded list; a model added later will then work without a change at your end.

Which model to pick ​

Both TTS models serve the same endpoint. They differ in what they trade away:

ModelReach for it whenServes
xvoice-tts1You want a choice of voices and languages, or a format other than wavwav, mp3, flac, pcm, opus; speed 0.25–4.0
xvoice-tts2Time to first audio is what matters — first bytes in about a third of a secondwav, pcm, mp3, all of which stream; speed 1.0

The table is a summary; speech_options on each model is the authority.

Voices ​

Voices are xVoice ids, not OpenAI's — alloy and nova mean nothing here. Most are lower case and hyphenated after what they sound like (neutral-female, casual-male, de-female); voices that ship with a name of their own keep it (sofia). The engine's own name for a voice is not the id you pass; ours stays the same when a voice moves to a different model.

python
import httpx

voices = httpx.get(
    f"{base_url}/voices",
    headers={"Authorization": f"Bearer {api_key}"},
).json()["data"]
json
{
  "id": "neutral-female",
  "name": "Neutral Female",
  "language": "en",
  "locale": "en",
  "gender": "female",
  "preview_url": null,
  "status": "available",
  "metadata": {}
}

?language=de and ?locale=de narrow the list. locale is the bare language tag wherever the engine exposes no region — es-ES against es-MX is not something we can claim to know — and a full tag such as en-US where it does.

Not every voice works with every model. The pairing is explicit in the catalog, and an unsupported combination fails with voice_not_supported rather than quietly substituting something that sounds wrong. ?model=xvoice-tts2 returns only the voices that model can speak with, which is the list to build a picker from:

python
voices = httpx.get(
    f"{base_url}/voices",
    params={"model": "xvoice-tts2"},
    headers={"Authorization": f"Bearer {api_key}"},
).json()["data"]

A model id nobody has bound a voice to returns an empty list rather than an error.

examples/list_models.py prints both catalogs.

Built on the OpenAI audio API surface.