DeepSeek V4 Flash 0731
A 284B-parameter mixture-of-experts model with 1M-token context that activates only 13B parameters at inference, making it a fast, low-cost sibling to V4 Pro.
- Input price
- per 1M tokens
- Output price
- per 1M tokens
- Context
- 1.05M131.1K max output
- Cost / request
- 8K in · 500 out
- Writing Elo
- 1438EQ-Bench creative writing
01Routes and pricing
02Capabilities
- Input
- text
- Output
- text
- Context
- 1.05M
- Max output
- 131.1K
ReasoningTool callingStructured output
03Benchmarks
Creative writing
1438.4 EloEQ-Bench Creative Writing v3EQ-Bench Creative Writing v3Output speed
External214 tok/s1.43s to first tokenArtificial Analysis04Quick start
Call it with your RProuter key
Any OpenAI-compatible SDK works. Point it at the RProuter base URL from your API keys page and use this model ID.
Catalog source: inference.baseten.co · checked 2026-09-16
const response = await client.chat.completions.create({
model: "deepseek/deepseek-v4-flash-0731",
messages: [{ role: "user", content: "A story begins…" }],
max_tokens: 500,
stream: true,
});
for await (const chunk of response) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}