Appearance
Rate limits and billing
Both are per organization, not per key or per project. Splitting work across projects separates the reporting, not the quota.
Rate limits
Four limits, each a separate bucket:
| Limit | Free | Paid |
|---|---|---|
| Requests per minute | 20 | 120 |
| Concurrent requests | 2 | 10 |
| TTS characters per minute | 25,000 | 200,000 |
| STT seconds per minute | 1,800 | 14,400 |
An organization moves to Paid once it has bought at least $5 of credit, cumulatively. Individual limits can be raised for an account; ask.
Concurrency is counted as work in flight, not requests started. A slot is held for as long as a response is streaming, and for the whole life of a realtime session — so ten simultaneous transcriptions and ten open realtime sessions cost the same ten slots.
Headers
The audio endpoints return the state of the buckets they charged. The catalog routes (/v1/models, /v1/voices, /v1/me) are not metered and carry no headers:
http
x-ratelimit-limit-requests: 120
x-ratelimit-remaining-requests: 118
x-ratelimit-reset-requests: 1s
x-ratelimit-limit-seconds: 14400
x-ratelimit-remaining-seconds: 14388
x-ratelimit-reset-seconds: 6m0sResources are requests, characters (TTS) and seconds (STT). reset is a duration — 20ms, 1.5s, 6m0s — not a timestamp.
A 429 adds retry-after in whole seconds. Honour it; retrying sooner just burns another request.
Two charges on transcription
POST /v1/audio/transcriptions charges the request bucket immediately, then charges the seconds bucket once the file has been probed — the duration is not knowable before that. A 429 on the second charge means you are within your request rate but over your audio rate.
Realtime sessions charge one request when the socket opens, then charge audio continuously as it arrives.
If the limiter is unavailable
Rate limiting fails open. If the counter store cannot be reached, requests are allowed rather than refused: an infrastructure problem on our side should not look like a quota problem on yours.
Credits
xVoice is prepaid. Credit is bought up front, and every priced request deducts from the balance in the same transaction that records the usage. At zero, requests get 402 insufficient_credits.
- Verifying your email grants the starting credit — $5, expiring after 30 days. It is granted once; opening the link twice does not pay twice.
- Purchased credit expires after a year.
- Top-ups are $5 to $1000 through Stripe Checkout, from Billing in the dashboard. Owners only.
- Auto-recharge keeps a service running unattended: save a card, set a threshold and a target balance, and the card is charged back up to the target whenever the balance falls below the threshold. Each recharge has to be a valid top-up on its own, so
target - thresholdmust be at least $5. - The dashboard shows a Low balance banner under $1 and an Out of credits banner at zero. There is no email for either yet, so a service left unattended wants auto-recharge rather than a person watching the banner.
Grants, not one number
Credit is held as grants, each with its own remaining amount and its own expiry, so two $5 top-ups bought on different days are different money. The balance is the sum of what is left on the unexpired ones.
Deduction order is soonest-expiring first, never-expiring last, then oldest first — so free and about-to-expire credit is spent before credit you paid for.
An expired grant stops counting immediately, not whenever a sweep next runs.
Overdraft
Requests already in flight when credit hits zero are allowed to finish and bill. Whatever no grant covers becomes overdraft, and the next top-up pays that off before anything else. Overdraft is small by construction: it is bounded by the work that was already running.
Spending limit
Separate from the balance and on top of it: a monthly cap you set, defaulting to $100. Crossing it is 402 spending_limit_reached, with credit still in the account.
It exists for the failure the balance cannot catch — a loop in your code spending real money quickly. Set it to what a bad month should cost, not to what you expect to spend.
Test keys are not free
A sk_test_ key calls the same models on the same hardware and deducts from the same balance. The environment separates usage, request logs and spending in the dashboard. It is not a sandbox, and there is no free tier of inference.
What is billable
The rule: if the engine did the work, it bills.
| Situation | Billed? |
|---|---|
| A successful request | Yes |
| Rejected before the engine — bad key, bad model, over a limit, file too large | No |
| The engine was unreachable and produced nothing | No |
| You disconnected while audio or text was streaming | Yes, for what was produced |
Realtime audio you sent and then cleared | Yes |
| Realtime audio in flight when you disconnected | Yes |
| An utterance the engine failed to transcribe, having sent no text | No |
Disconnecting does not cancel a cost that has already been incurred — the GPU time is spent whether or not you read the result. The same rule applies on every endpoint, which is why realtime bills from our own count of the bytes you sent rather than from a completion event a vanished client never triggers.
Seeing where it went
The dashboard has Usage (aggregated by day, product, model and project) and Request logs (one row per request: model, status, latency, cost, request id). Both read the same usage_events the billing deduction was made from, so the numbers reconcile by construction rather than by a nightly job.
Every request id in those logs is the request_id from the error envelope and the X-Request-Id header, which is what makes a customer report traceable to a row.