Models

All models

Every model ModelRelay serves, grouped by what it does. Nothing here is ranked: text, speech-to-text and text-to-speech metrics share no units.

ModelPrice
Claude Fable 5.1$10.00 in · $50.00 out per 1M tokens
Claude Opus 5$5.00 in · $25.00 out per 1M tokens
Claude Sonnet 5$2.00 in · $10.00 out per 1M tokens
DeepSeek V4 Flash 0731$0.22 in · $0.66 out per 1M tokens
DeepSeek V4.1 Flash$0.22 in · $0.66 out per 1M tokens
GLM 5.3 Flash$0.15 in · $0.50 out per 1M tokens
GPT-5.6 Luna$0.20 in · $1.20 out per 1M tokens
GPT-5.6 Sol$5.00 in · $30.00 out per 1M tokens
GPT-5.6 Terra$2.00 in · $12.00 out per 1M tokens
GPT-6 Astra$10.00 in · $50.00 out per 1M tokens
Gemini 3.8 Flash$0.75 in · $3.75 out per 1M tokens
Grok 4.6$2.00 in · $6.00 out per 1M tokens
Kimi K3$3.00 in · $15.00 out per 1M tokens
Muse Spark 1.3$1.25 in · $4.25 out per 1M tokens
Muse Spark 1.3 Contributor$0.10 in · $0.20 out per 1M tokens
Nemotron Lightning 3.5 30B A3B$0.05 in · $0.20 out per 1M tokens
Qwen 3.8 Max$2.00 in · $6.00 out per 1M tokens
Qwen3.8 27B$0.99 in · $1.49 out per 1M tokens
ModelPrice
Grok Voice Transcribe 2.0$0.10/audio-hour billed per ms of input audio
Whisper Large v3 Turbo$0.0024/audio-hour billed per ms of input audio
ModelPrice
Gemini 3.8 Flash TTS$16.50/1M characters billed per input character
One integration

One API, with drop-in compatibility.

Use ModelRelay's Responses API directly, or point an existing OpenAI or Anthropic SDK at the compatible endpoints.

138 successful streams measured · 30m window

Call a model
ModelRelay · POST /api/v1/responses Drop-in · OpenAI /v1/responses · Anthropic /v1/messages
curl https://api.modelrelay.ai/api/v1/responses \
-H "Authorization: Bearer $MODELRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-5","input":"Hello"}'

Change the model. Keep everything else.

Build across models without rebuilding your stack.

Open to everyone. Sign in with GitHub or Google to get started.

Get started