Skip to content

Text to speech ​

POST /v1/audio/speech

OpenAI-compatible. Billed per character of input.

python
speech = client.audio.speech.create(
    model="xvoice-tts1",
    voice="neutral-female",
    input="Hello, this is a round trip through xVoice.",
    response_format="wav",
)
speech.write_to_file("speech.wav")

The response body is the audio itself, with the media type of the format you asked for. Errors before the first byte are ordinary HTTP errors; see Errors after a 200 for what happens later.

Parameters ​

FieldTypeDefaultNotes
modelstringrequiredA TTS model id — see Models
inputstringrequiredThe text. At most 8000 characters
voicestringrequiredAn xVoice voice id, e.g. neutral-female
response_formatstringwavwav, mp3, flac, pcm, opus — which ones, and which of them stream, is per model: read speech_options
speednumber1.00.25 to 4.0. Must be 1.0 when streaming, and on xvoice-tts2
streamboolfalseStream the audio as it is generated. Requires response_format: pcm
stream_formatstring—audio or sse. OpenAI's spelling; either turns streaming on
languagestring—Hint for the language of input
instructionsstring—Delivery guidance, at most 1000 characters. Passed to the engine; whether a model acts on it depends on the model

input over the limit is rejected before any work happens, so a too-long request costs nothing.

What the model changes ​

response_format and speed are not equally supported by every model, and asking for one a model does not offer is refused rather than quietly substituted — generating 1x audio for a speed: 1.5 request is not something you could find out from the bytes. Read speech_options on the model rather than hard-coding a list; it is the same object the API enforces:

python
model = client.models.retrieve("xvoice-tts2")
model.speech_options.response_formats   # ["pcm", "wav"]
model.speech_options.speed              # {"min": 1.0, "max": 1.0}
model.speech_options.instructions       # False -- accepted, but ignored

Today that reads:

xvoice-tts1xvoice-tts2
response_formatwav, mp3, flac, pcm, opuswav, pcm, mp3
Of those, streamspcmwav, pcm, mp3
speed0.25 to 4.01.0
instructionspassed to the engineignored
Voices20, in 9 languagessee GET /v1/voices?model=xvoice-tts2
Time to first audio, streamingaround a secondaround a third of one

xvoice-tts2 is unusual in streaming wav: it works out how long the clip will be before generating it, so the RIFF header in the first chunk is already correct. Most engines cannot do that, which is why streamable_formats is per model rather than one rule for everyone.

Call GET /v1/voices?model=… rather than assuming a voice id carries across models — they do not all exist on both.

Streaming ​

Streaming matters when the text is long enough that generating all of it first would be an audible wait. The first audio arrives in a fraction of the time the whole clip takes to produce.

python
with client.audio.speech.with_streaming_response.create(
    model="xvoice-tts1",
    voice="neutral-female",
    input=long_text,
    response_format="pcm",
    stream_format="audio",
) as response:
    for chunk in response.iter_bytes():
        player.write(chunk)

Two constraints, both from the engine rather than from us:

  • response_format must be one of the model's streamable_formats. On most engines that is pcm alone: everything else is produced by encoding the finished clip, and wav cannot stream at all because its header declares the length before a single sample exists. xvoice-tts2 is the exception and streams all three of its formats. Asking for one a model cannot stream is unsupported_response_format, not a silent fallback to buffered.
  • speed must be 1.0. Resampling needs the whole clip, so this one holds for every model whatever its speed range says.

pcm is raw 16-bit little-endian samples at 24 kHz, mono, with no header — feed it straight to an audio device. To end up with a file something else can open, write a 44-byte WAV header in front of the collected samples once the stream finishes, or pipe the bytes through ffmpeg -f s16le -ar 24000 -ac 1 -i - out.wav.

As server-sent events ​

stream_format="sse" wraps the same audio in events, base64 in JSON, and ends with a usage event:

text
event: speech.audio.delta
data: {"type":"speech.audio.delta","audio":"<base64>","response_format":"pcm"}

event: speech.audio.done
data: {"type":"speech.audio.done","usage":{"input_tokens":12,"output_tokens":480,...}}

Use it when you need the token counts in band, or when the transport between you and the player is easier to write against events than against a byte stream. It costs about a third more bytes on the wire than stream_format="audio".

What a request costs ​

Billing is per character of input, counted before the call, so the cost is known before any audio exists. A disconnect partway through a stream still bills: the engine did the work. See What is billable.

A request that fails before the engine produces its first chunk bills nothing — including an engine that returns no audio at all, which is a 502 (inference_unavailable) rather than an empty 200.

Errors ​

CodeStatusCause
invalid_model400Not a TTS model, or no such id
invalid_voice404No such voice
voice_not_supported400That voice does not belong to that model
model_unavailable404The model has no deployed version
validation_error422The body failed schema validation — input too long, speed outside 0.25–4.0
unsupported_response_format400That model does not produce that format, or cannot stream it
unsupported_speed400Outside that model's range, or not 1.0 while streaming
inference_unavailable502We could not reach the engine, or it produced nothing

Runnable: tts.py and tts_stream.py.

Built on the OpenAI audio API surface.