DeepSeek V4 Flash 0731 pricing and specs

open weights

Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding

16,9 ₽input / 1M$0.2
33,8 ₽output / 1M$0.4
1Mcontexttokens
384Kmax outputtokens

DeepSeek V4 Flash 0731 price per 1M tokens

Type₽ / 1M$ / 1M
Input tokens16,9 ₽$0.2
Output tokens33,8 ₽$0.4
Cache read3,38 ₽$0.04

Prices in rubles as of 18.09.2026 at the CBR rate: 1 $ = 84.51 ₽.

DeepSeek V4 Flash 0731 specs and limits

Context window
1M
Max output
384K
Input
text
Output
text
Family
deepseek-flash
Released
2026-07-31
Updated
2026-07-31
Knowledge cutoff
2025-05
Also known as
qwen deepseek v4 flash 0731, qwen3 deepseek v4 flash 0731, qwen 3 deepseek v4 flash 0731, qwen3.5 deepseek v4 flash 0731, qwen3.6 deepseek v4 flash 0731, qwen3.7 deepseek v4 flash 0731

What DeepSeek V4 Flash 0731 can do

  • Reasoning
  • Tool calling
  • Structured output
  • Attachments / files
  • Vision (images)
  • Temperature control
  • Open weights

Frequently asked questions

How much does the DeepSeek V4 Flash 0731 API cost?

Input costs $0.2 per 1M tokens (16.9 ₽) and output costs $0.4 per 1M (33.8 ₽). Alibaba bills the prompt and the completion separately, so the price of a call depends on how long both are. Ruble figures use the Bank of Russia rate of 84.51 ₽ per dollar for September 2026.

How much do 1,000 DeepSeek V4 Flash 0731 tokens cost?

List prices are quoted per 1M tokens, so 1,000 tokens cost a thousandth of that: divide $0.2 for input and $0.4 for output by 1,000. The same rates in rubles are 16.9 ₽ and 33.8 ₽ per 1M tokens. Both prompt tokens and completion tokens are billed.

How many tokens does DeepSeek V4 Flash 0731 hold?

The DeepSeek V4 Flash 0731 context window is 1M tokens (1,000,000). Everything shares that budget: the system prompt, the dialogue history, attached files and the answer the model writes. The figure comes from the Alibaba model card as of September 2026.

What are the DeepSeek V4 Flash 0731 limits?

The context window is 1M tokens (1,000,000) and a single response is capped at 384K tokens, so longer output has to be generated in parts. Rate limits are set by Alibaba per account and depend on your plan rather than on the model, so they are not listed here.

How is the DeepSeek V4 Flash 0731 price in rubles calculated?

Providers quote their tariffs in US dollars, so the ruble amounts here are a conversion at the official Bank of Russia rate of 84.51 ₽ per dollar. For DeepSeek V4 Flash 0731 that gives 16.9 ₽ per 1M input tokens and 33.8 ₽ per 1M output tokens. When the rate moves, the ruble price moves with it.

Can DeepSeek V4 Flash 0731 reason?

Yes, DeepSeek V4 Flash 0731 declares a reasoning mode: it works through intermediate steps before it answers. Those steps add output tokens, so a call with a long reasoning chain costs more and takes longer than a plain completion. Capabilities declared for the model: reasoning, tool calling, structured output.

What capabilities does DeepSeek V4 Flash 0731 support?

The declared capabilities of DeepSeek V4 Flash 0731 are: reasoning, tool calling, structured output. It accepts text as input, and every attachment consumes tokens from the shared context window of 1M tokens. The list comes from the Alibaba model card and is refreshed with the catalog as of September 2026.

Does DeepSeek V4 Flash 0731 have open weights?

Yes, DeepSeek V4 Flash 0731 ships with open weights, so it can be downloaded and served on your own hardware instead of being used only through the Alibaba API. What you may do with it is set by the license shown on the card. Catalog prices cover hosted API access.

Other Alibaba models