Appearance
Models and voices
Models
python
for model in client.models.list():
print(model.id, model.type, model.capabilities)GET /v1/models returns a superset of OpenAI's model object, so their SDK's model page parses against us unchanged:
json
{
"id": "xvoice-stt1",
"object": "model",
"created": 1758672000,
"owned_by": "xvoice",
"name": "xVoice STT v1",
"type": "stt",
"description": "Speech recognition.",
"capabilities": ["batch"],
"supported_languages": ["en", "de", "fr"],
"status": "available",
"public_version": "2026-09-17",
"speech_options": null
}GET /v1/models/{model_id} returns one of these, and GET /v1/models?type=stt (or tts) narrows the list.
capabilities
Which endpoints and transports will take this model. It is the field to filter on; once you have picked a model, speech_options below is the field that says what you may send it.
| Capability | Meaning |
|---|---|
batch | Accepts a whole request at once: POST /v1/audio/transcriptions, or a complete POST /v1/audio/speech |
streaming | Can produce output incrementally over HTTP — SSE, or a chunked audio body |
realtime | Can hold a live session: WS /v1/realtime |
It is enforced, not descriptive. A model without realtime refuses a session with realtime_not_supported, naming the file endpoint to use instead. Read the list rather than guessing from the name: a model that transcribes files well may have no streaming deployment behind it at all, and those are separate services rather than modes of one.
On a TTS model, batch and streaming are derived rather than declared — streaming appears exactly when speech_options.streamable_formats is non-empty, and batch always does, because every TTS model will return a complete clip. So the two fields cannot disagree about whether a model streams, and the detailed answer (which format streams) is only ever in one place.
type and status
type is stt or tts, and it decides which endpoints will take the model. Handing a TTS id to transcriptions is invalid_model, not a 404.
status is available, deprecated or retired. Deprecated models still serve — plan to move. Retired ones are gone from the list. A model that is listed but has no deployed version behind it fails with model_unavailable, which is our problem rather than yours: it means the catalog and the deployment disagree.
versions
public_version names the version a call resolves to today. You do not pass it; pinning is not exposed in V1. It is there so that a change in output has a label you can quote when you ask us what moved.
speech_options
TTS models carry a speech_options object saying which speech parameters they accept. It is null on STT models, which take none of them.
json
{
"response_formats": ["pcm", "wav"],
"streamable_formats": ["pcm"],
"speed": {"min": 1.0, "max": 1.0},
"instructions": false
}Like capabilities, this is enforced and not merely descriptive: asking for a format or speed outside it is unsupported_response_format or unsupported_speed, a 400 naming the model and listing what it does accept, rather than a silent substitution. The API resolves these values from the same place it enforces them, so what you read and what you get cannot drift apart.
instructions: false is the one exception. That field is documented as best-effort guidance, so a model that ignores it says so here and still accepts the request — you are told rather than refused.
Build a picker from this object rather than from a hard-coded list; a model added later will then work without a change at your end.
Which model to pick
Both TTS models serve the same endpoint. They differ in what they trade away:
| Model | Reach for it when | Serves |
|---|---|---|
xvoice-tts1 | You want a choice of voices and languages, or a format other than wav | wav, mp3, flac, pcm, opus; speed 0.25–4.0 |
xvoice-tts2 | Time to first audio is what matters — first bytes in about a third of a second | wav, pcm, mp3, all of which stream; speed 1.0 |
The table is a summary; speech_options on each model is the authority.
Voices
Voices are xVoice ids, not OpenAI's — alloy and nova mean nothing here. Most are lower case and hyphenated after what they sound like (neutral-female, casual-male, de-female); voices that ship with a name of their own keep it (sofia). The engine's own name for a voice is not the id you pass; ours stays the same when a voice moves to a different model.
python
import httpx
voices = httpx.get(
f"{base_url}/voices",
headers={"Authorization": f"Bearer {api_key}"},
).json()["data"]json
{
"id": "neutral-female",
"name": "Neutral Female",
"language": "en",
"locale": "en",
"gender": "female",
"preview_url": null,
"status": "available",
"metadata": {}
}?language=de and ?locale=de narrow the list. locale is the bare language tag wherever the engine exposes no region — es-ES against es-MX is not something we can claim to know — and a full tag such as en-US where it does.
Not every voice works with every model. The pairing is explicit in the catalog, and an unsupported combination fails with voice_not_supported rather than quietly substituting something that sounds wrong. ?model=xvoice-tts2 returns only the voices that model can speak with, which is the list to build a picker from:
python
voices = httpx.get(
f"{base_url}/voices",
params={"model": "xvoice-tts2"},
headers={"Authorization": f"Bearer {api_key}"},
).json()["data"]A model id nobody has bound a voice to returns an empty list rather than an error.
examples/list_models.py prints both catalogs.