Models

Text to speech models

Prices are per million input characters. Speed is the median over real ModelRelay requests, in characters synthesized per second. Voice quality is the Artificial Analysis TTS arena Elo, from blind listener votes. Models are listed rather than ranked until there are enough of them to compare.

ModelPriceSpeed chars/s, ModelRelay medianVoice quality AA arena Elo
Gemini 3.8 Flash TTS$16.50/1M characters billed per input character28 chars/s1268 ±16
One integration

One API, with drop-in compatibility.

Use ModelRelay's Responses API directly, or point an existing OpenAI or Anthropic SDK at the compatible endpoints.

128 successful streams measured · 30m window

Call a model
ModelRelay · POST /api/v1/responses Drop-in · OpenAI /v1/responses · Anthropic /v1/messages
curl https://api.modelrelay.ai/api/v1/responses \
-H "Authorization: Bearer $MODELRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-5","input":"Hello"}'

Change the model. Keep everything else.

Build across models without rebuilding your stack.

Open to everyone. Sign in with GitHub or Google to get started.

Get started