Appearance
Text to speech
POST /v1/audio/speech
OpenAI-compatible. Billed per character of input.
python
speech = client.audio.speech.create(
model="xvoice-tts1",
voice="neutral-female",
input="Hello, this is a round trip through xVoice.",
response_format="wav",
)
speech.write_to_file("speech.wav")The response body is the audio itself, with the media type of the format you asked for. Errors before the first byte are ordinary HTTP errors; see Errors after a 200 for what happens later.
Parameters
| Field | Type | Default | Notes |
|---|---|---|---|
model | string | required | A TTS model id — see Models |
input | string | required | The text. At most 8000 characters |
voice | string | required | An xVoice voice id, e.g. neutral-female |
response_format | string | wav | wav, mp3, flac, pcm, opus — which ones, and which of them stream, is per model: read speech_options |
speed | number | 1.0 | 0.25 to 4.0. Must be 1.0 when streaming, and on xvoice-tts2 |
stream | bool | false | Stream the audio as it is generated. Requires response_format: pcm |
stream_format | string | — | audio or sse. OpenAI's spelling; either turns streaming on |
language | string | — | Hint for the language of input |
instructions | string | — | Delivery guidance, at most 1000 characters. Passed to the engine; whether a model acts on it depends on the model |
input over the limit is rejected before any work happens, so a too-long request costs nothing.
What the model changes
response_format and speed are not equally supported by every model, and asking for one a model does not offer is refused rather than quietly substituted — generating 1x audio for a speed: 1.5 request is not something you could find out from the bytes. Read speech_options on the model rather than hard-coding a list; it is the same object the API enforces:
python
model = client.models.retrieve("xvoice-tts2")
model.speech_options.response_formats # ["pcm", "wav"]
model.speech_options.speed # {"min": 1.0, "max": 1.0}
model.speech_options.instructions # False -- accepted, but ignoredToday that reads:
xvoice-tts1 | xvoice-tts2 | |
|---|---|---|
response_format | wav, mp3, flac, pcm, opus | wav, pcm, mp3 |
| Of those, streams | pcm | wav, pcm, mp3 |
speed | 0.25 to 4.0 | 1.0 |
instructions | passed to the engine | ignored |
| Voices | 20, in 9 languages | see GET /v1/voices?model=xvoice-tts2 |
| Time to first audio, streaming | around a second | around a third of one |
xvoice-tts2 is unusual in streaming wav: it works out how long the clip will be before generating it, so the RIFF header in the first chunk is already correct. Most engines cannot do that, which is why streamable_formats is per model rather than one rule for everyone.
Call GET /v1/voices?model=… rather than assuming a voice id carries across models — they do not all exist on both.
Streaming
Streaming matters when the text is long enough that generating all of it first would be an audible wait. The first audio arrives in a fraction of the time the whole clip takes to produce.
python
with client.audio.speech.with_streaming_response.create(
model="xvoice-tts1",
voice="neutral-female",
input=long_text,
response_format="pcm",
stream_format="audio",
) as response:
for chunk in response.iter_bytes():
player.write(chunk)Two constraints, both from the engine rather than from us:
response_formatmust be one of the model'sstreamable_formats. On most engines that ispcmalone: everything else is produced by encoding the finished clip, andwavcannot stream at all because its header declares the length before a single sample exists.xvoice-tts2is the exception and streams all three of its formats. Asking for one a model cannot stream isunsupported_response_format, not a silent fallback to buffered.speedmust be1.0. Resampling needs the whole clip, so this one holds for every model whatever itsspeedrange says.
pcm is raw 16-bit little-endian samples at 24 kHz, mono, with no header — feed it straight to an audio device. To end up with a file something else can open, write a 44-byte WAV header in front of the collected samples once the stream finishes, or pipe the bytes through ffmpeg -f s16le -ar 24000 -ac 1 -i - out.wav.
As server-sent events
stream_format="sse" wraps the same audio in events, base64 in JSON, and ends with a usage event:
text
event: speech.audio.delta
data: {"type":"speech.audio.delta","audio":"<base64>","response_format":"pcm"}
event: speech.audio.done
data: {"type":"speech.audio.done","usage":{"input_tokens":12,"output_tokens":480,...}}Use it when you need the token counts in band, or when the transport between you and the player is easier to write against events than against a byte stream. It costs about a third more bytes on the wire than stream_format="audio".
What a request costs
Billing is per character of input, counted before the call, so the cost is known before any audio exists. A disconnect partway through a stream still bills: the engine did the work. See What is billable.
A request that fails before the engine produces its first chunk bills nothing — including an engine that returns no audio at all, which is a 502 (inference_unavailable) rather than an empty 200.
Errors
| Code | Status | Cause |
|---|---|---|
invalid_model | 400 | Not a TTS model, or no such id |
invalid_voice | 404 | No such voice |
voice_not_supported | 400 | That voice does not belong to that model |
model_unavailable | 404 | The model has no deployed version |
validation_error | 422 | The body failed schema validation — input too long, speed outside 0.25–4.0 |
unsupported_response_format | 400 | That model does not produce that format, or cannot stream it |
unsupported_speed | 400 | Outside that model's range, or not 1.0 while streaming |
inference_unavailable | 502 | We could not reach the engine, or it produced nothing |
Runnable: tts.py and tts_stream.py.