Models

Speech to text models

Prices are per audio-hour of input. Speed is the median over real ModelRelay requests, in seconds of audio transcribed per second. Accuracy is not measured yet, so models are listed rather than ranked.

ModelPriceSpeed × real time, ModelRelay median
Grok Voice Transcribe 2.0$0.10/audio-hour billed per ms of input audio—
Whisper Large v3 Turbo$0.0024/audio-hour billed per ms of input audio—
One integration

One API, with drop-in compatibility.

Use ModelRelay's Responses API directly, or point an existing OpenAI or Anthropic SDK at the compatible endpoints.

135 successful streams measured · 30m window

Call a model
ModelRelay · POST /api/v1/responses Drop-in · OpenAI /v1/responses · Anthropic /v1/messages
curl https://api.modelrelay.ai/api/v1/responses \
-H "Authorization: Bearer $MODELRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-5","input":"Hello"}'

Change the model. Keep everything else.

Build across models without rebuilding your stack.

Open to everyone. Sign in with GitHub or Google to get started.

Get started