Connect

Model Benchmarks

Early-access prices are provider rates with no ModelRelay markup. Intelligence by Artificial Analysis.

Speed
Cost
Intelligence
Anthropic

Claude Fable 5

TTFT
2.63s — evidence: controlled · n=8
TPS
115.0tps
Uptime
Intelligence
59.9
Unretouched model output
Input / 1M
$10.00
Output / 1M
$50.00
Anthropic

Claude Opus 5

TTFT
1.84s — evidence: controlled · n=8
TPS
115.8tps
Uptime
Intelligence
60.7
Unretouched model output
Input / 1M
$5.00
Output / 1M
$25.00
Anthropic

Claude Sonnet 5

TTFT
1.94s — evidence: controlled · n=8
TPS
150.4tps
Uptime
Intelligence
53.4
Unretouched model output
Input / 1M
$2.00
Output / 1M
$10.00
Google AI Studio

Gemini 3.5 Flash-Lite

TTFT
0.48s — evidence: controlled · n=8
TPS
439.3tps
Uptime
Intelligence
36.5
Unretouched model output
Input / 1M
$0.30
Output / 1M
$2.50
Google AI Studio

Gemini 3.6 Flash

TTFT
0.84s — evidence: controlled · n=8
TPS
272.5tps
Uptime
Intelligence
50.1
Unretouched model output
Input / 1M
$1.50
Output / 1M
$7.50
OpenAI

GPT-5.6 Luna

TTFT
1.11s — evidence: controlled · n=8
TPS
171.3tps
Uptime
Intelligence
51.2
Unretouched model output
Input / 1M
$0.20 ≤272K input
Output / 1M
$1.20 ≤272K input
OpenAI

GPT-5.6 Sol

TTFT
1.00s — evidence: controlled · n=8
TPS
117.7tps
Uptime
Intelligence
58.9
Unretouched model output
Input / 1M
$5.00 ≤272K input
Output / 1M
$30.00 ≤272K input
OpenAI

GPT-5.6 Terra

TTFT
0.78s — evidence: controlled · n=8
TPS
155.9tps
Uptime
Intelligence
55.0
Unretouched model output
Input / 1M
$2.00 ≤272K input
Output / 1M
$12.00 ≤272K input
xAI

Grok 4.5

TTFT
1.51s — evidence: controlled · n=6
TPS
71.8tps
Uptime
Intelligence
53.8
Unretouched model output
Input / 1M
$2.00 ≤200K input
Output / 1M
$6.00 ≤200K input
Swipe to compare all columns →
AnthropicClaude Fable 52.63s — evidence: controlled · n=8115.0tps — evidence: controlled · n=8$10.00$50.0059.9
AnthropicClaude Opus 51.84s — evidence: controlled · n=8115.8tps — evidence: controlled · n=8$5.00$25.0060.7
AnthropicClaude Sonnet 51.94s — evidence: controlled · n=8150.4tps — evidence: controlled · n=8$2.00$10.0053.4
Google AI StudioGemini 3.5 Flash-Lite0.48s — evidence: controlled · n=8439.3tps — evidence: controlled · n=8$0.30$2.5036.5
Google AI StudioGemini 3.6 Flash0.84s — evidence: controlled · n=8272.5tps — evidence: controlled · n=8$1.50$7.5050.1
OpenAIGPT-5.6 Luna1.11s — evidence: controlled · n=8171.3tps — evidence: controlled · n=8$0.20 ≤272K input$1.20 ≤272K input51.2
OpenAIGPT-5.6 Sol1.00s — evidence: controlled · n=8117.7tps — evidence: controlled · n=8$5.00 ≤272K input$30.00 ≤272K input58.9
OpenAIGPT-5.6 Terra0.78s — evidence: controlled · n=8155.9tps — evidence: controlled · n=8$2.00 ≤272K input$12.00 ≤272K input55.0
xAIGrok 4.51.51s — evidence: controlled · n=671.8tps — evidence: controlled · n=6$2.00 ≤200K input$6.00 ≤200K input53.8

TTFT is time to first visible answer token; TPS is visible answer tokens per second after that token. Both use a fixed prompt at each model's minimum reasoning. Speed combines both phases into time to a 500-token answer (TTFT + 500/TPS); less time ranks higher, and neither column changes meaning. Intelligence is the Artificial Analysis Intelligence Index at each model's highest scored variant.

One integration

One API, with drop-in compatibility.

Use ModelRelay's Responses API directly, or point an existing OpenAI or Anthropic SDK at the compatible endpoints.

Call any model
ModelRelay · POST /api/v1/responses Drop-in · OpenAI /v1/responses · Anthropic /v1/messages
curl https://api.modelrelay.ai/api/v1/responses \
-H "Authorization: Bearer $MODELRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-5","input":"Hello"}'

Change the model. Keep everything else.

Build across models without rebuilding your stack.

ModelRelay is invite-only while we onboard early teams.

Request access