← All models

DeepSeek V4.1 Flash

deepseek-v4p1-flash

Multimodal open-weight MoE model for long-context reasoning and agentic work

Intelligence
39.5
Context
1.04M
Training cutoff
--
Added
2026-09-25
Inputs
Vision

Providers

One route serves DeepSeek V4.1 Flash, and every figure below is measured on it. Where the selected window has too little traffic, the figure names what it is based on instead.

Fireworks AI only route

accounts/fireworks/models/deepseek-v4p1-flash

Uptime100.0%
Input / 1M
$0.22
Output / 1M
$0.66
TTFT
1.13s fixed prompt · n=5
Output rate
120.8 tok/s fixed prompt · n=5
1.04M context 131,072 max output cache $0.007 read / $0.22 write per 1M Vision ZDRwarming · 30m

TTFT is the median time to the first visible answer token over real ModelRelay requests, with a fixed prompt standing in for models that have no traffic yet. Speed ranks by it. Intelligence is the Artificial Analysis Intelligence Index.