Models

Models

Claude Opus 5 and Sonnet 5 have thinking on at medium effort by default. Set reasoning_effort to none to disable thinking.

Prices are provider rates with no ModelRelay markup. Intelligence by Artificial Analysis.

Speed
Cost
Intelligence
Accepts
TTFT
2.98s — evidence: controlled · n=8
Uptime
Intelligence
53.4
Input / 1M
$10.00
Output / 1M
$50.00
Max context
1M
TTFT
-- — evidence: warming 0/3
Uptime
Intelligence
50.8
Input / 1M
$5.00
Output / 1M
$25.00
Max context
1M
TTFT
1.64s — evidence: controlled · n=8
Uptime
Intelligence
38.2
Input / 1M
$2.00
Output / 1M
$10.00
Max context
1M
TTFT
1.80s — evidence: controlled · n=8
Uptime
Intelligence
34.3
Input / 1M
$0.22
Output / 1M
$0.66
Max context
1.04M
TTFT
1.99s — evidence: controlled · n=8
Uptime
Intelligence
39.5
Input / 1M
$0.22
Output / 1M
$0.66
Max context
1.04M
TTFT
6.25s — evidence: live · 30m · n=210
Uptime
Intelligence
41.8
Input / 1M
$0.15
Output / 1M
$0.50
Max context
1.04M
TTFT
2.11s — evidence: controlled · n=8
Uptime
Intelligence
37.3
Input / 1M
$0.20 ≤272K input
Output / 1M
$1.20 ≤272K input
Max context
1.05M
TTFT
0.97s — evidence: controlled · n=8
Uptime
Intelligence
47.0
Input / 1M
$5.00 ≤272K input
Output / 1M
$30.00 ≤272K input
Max context
1.05M
TTFT
0.81s — evidence: controlled · n=8
Uptime
Intelligence
42.1
Input / 1M
$2.00 ≤272K input
Output / 1M
$12.00 ≤272K input
Max context
1.05M
TTFT
1.56s — evidence: controlled · n=8
Uptime
Intelligence
52.7
Input / 1M
$10.00 ≤272K input
Output / 1M
$50.00 ≤272K input
Max context
1.05M
TTFT
1.26s — evidence: controlled · n=8
Uptime
Intelligence
40.9
Input / 1M
$0.75
Output / 1M
$3.75
Max context
1.04M

Grok 4.6

Vision
TTFT
1.77s — evidence: controlled · n=8
Uptime
Intelligence
44.3
Input / 1M
$2.00 ≤200K input
Output / 1M
$6.00 ≤200K input
Max context
500K

Kimi K3

Vision
TTFT
4.72s — evidence: controlled · n=8
Uptime
Intelligence
43.6
Input / 1M
$3.00
Output / 1M
$15.00
Max context
1.04M
TTFT
-- — evidence: warming 0/3
Uptime
Intelligence
45.1
Input / 1M
$1.25
Output / 1M
$4.25
Max context
1.04M
TTFT
-- — evidence: warming 0/3
Uptime
Intelligence
45.1
Input / 1M
$0.10
Output / 1M
$0.20
Max context
1.04M
TTFT
0.64s — evidence: controlled · n=8
Uptime
Intelligence
12.9
Input / 1M
$0.05
Output / 1M
$0.20
Max context
262K
TTFT
1.35s — evidence: controlled · n=8
Uptime
Intelligence
40.2
Input / 1M
$2.00
Output / 1M
$6.00
Max context
1M
TTFT
0.59s — evidence: controlled · n=8
Uptime
Intelligence
33.7
Input / 1M
$0.99
Output / 1M
$1.49
Max context
131K

No model accepts every selected input.

TTFT is the median time to the first visible answer token over real ModelRelay requests, with a fixed prompt standing in for models that have no traffic yet. Speed ranks by it. Intelligence is the Artificial Analysis Intelligence Index.

One integration

One API, with drop-in compatibility.

Use ModelRelay's Responses API directly, or point an existing OpenAI or Anthropic SDK at the compatible endpoints.

210 successful streams measured · 30m window

Call a model
ModelRelay · POST /api/v1/responses Drop-in · OpenAI /v1/responses · Anthropic /v1/messages
curl https://api.modelrelay.ai/api/v1/responses \
-H "Authorization: Bearer $MODELRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-5","input":"Hello"}'

Change the model. Keep everything else.

Build across models without rebuilding your stack.

Open to everyone. Sign in with GitHub or Google to get started.

Get started